Optimization Method and System for Database Adaptive Data Stream Acquisition Based on Reinforcement Learning

By adopting a two-layer deep Q learning network structure based on reinforcement learning in the database, combining genetic algorithms and multi-objective evaluation system, adaptive optimization of data flow acquisition is achieved, which solves the shortcomings of acquisition efficiency and strategy adjustment in the existing technology, and improves acquisition efficiency and stability.

CN119719783BActive Publication Date: 2025-06-13北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510218162.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art is difficult to collect data streams efficiently and in real time in the database, and adaptively adjust the acquisition strategy according to the dynamic changes of the data stream, and lacks comprehensive optimization for multiple goals.

Method used

Using a two-layer deep Q learning network structure based on reinforcement learning, we use a distributed data acquisition agent network to acquire data flow information in real time, extract multimodal features, calculate initial acquisition parameters with genetic algorithms, design a hierarchical reward function and a multi-objective evaluation system to achieve dynamic optimization of the acquisition strategy.

Benefits of technology

Adaptive optimization of database data flow acquisition is realized, acquisition efficiency and stability is improved, manual intervention costs are reduced, and continuous optimization of acquisition strategies is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719783B_ABST
    Figure CN119719783B_ABST
Patent Text Reader

Abstract

The present invention provides an optimization method and system for database adaptive data stream acquisition based on reinforcement learning, which relates to the technical field of databases. It includes constructing a distributed acquisition agent network, obtaining database data stream information, and performing multi-modal feature decomposition and fusion. Using the fused feature vector, combined with a genetic algorithm to calculate the optimal initial acquisition parameter combination, and constructing a two-layer deep Q-learning network structure for hierarchical training to generate a hierarchical acquisition strategy model. This model adaptively adjusts the global acquisition strategy and parameters according to the feature drift degree and performance indicators to achieve continuous optimization of the acquisition strategy. The present invention effectively improves the data stream acquisition efficiency, reduces the acquisition delay, and optimizes the resource utilization rate through multi-modal feature fusion, reinforcement learning, and an adaptive tuning mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to database technology, and in particular to an optimized method and system for database adaptive data stream acquisition based on reinforcement learning. Background Art

[0002] With the rapid development of big data technology, data stream acquisition in databases has become an important research direction. How to efficiently and real-time acquire data streams from databases and adaptively adjust the acquisition strategy according to the dynamic changes of data streams is crucial for ensuring the accuracy and timeliness of data analysis.

[0003] Existing data stream acquisition methods mainly adopt the method of statically configuring acquisition parameters and cannot adapt to the dynamic changes of data streams. In addition, traditional acquisition methods usually only focus on a single metric, such as acquisition latency or throughput, and lack comprehensive optimization of multiple objectives. Moreover, existing methods rarely consider the changes in data stream characteristics, resulting in a decline in the effectiveness of acquisition strategies and making it difficult to meet the requirements of practical applications. Summary of the Invention

[0004] Embodiments of the present invention provide an optimized method and system for database adaptive data stream acquisition based on reinforcement learning, which can solve the problems in the prior art.

[0005] In the first aspect of the embodiments of the present invention,

[0006] An optimized method for database adaptive data stream acquisition based on reinforcement learning is provided, including:

[0007] Constructing a distributed data acquisition agent network to obtain data stream information of the database in real time and generate an initial feature set; performing multi-modal decomposition on the initial feature set to extract data stream fluctuation spectrum feature values, fine-grained type distribution feature values, multi-scale time series feature values, and graph structure association feature values to form a multi-modal feature matrix; based on the multi-modal feature matrix, calculating feature weight coefficients through an improved attention mechanism, adaptively fusing different modal features, and outputting a fused feature vector; using the fused feature vector and combining with a genetic algorithm to calculate an optimal initial acquisition parameter combination, including a concurrent acquisition threshold, a cache capacity threshold, and a batch processing size threshold, to generate an initial state space vector;

[0008] Based on the initial state space vector and the acquisition parameter combination, construct a double-layer deep Q-learning network structure. Set the input of the first network as the fusion feature vector and output the global acquisition strategy; set the input of the second network as the initial state space vector and output the specific parameter adjustment scheme. Design a hierarchical reward function, use the feature weight coefficient as the balance factor of the reward function, and establish a multi-objective evaluation system including the acquisition delay rate, throughput efficiency, and resource utilization rate. Based on the multi-objective evaluation system, perform a hierarchical training process, use the global acquisition strategy output by the first network as the constraint condition of the second network, and iteratively update the network parameters through the priority experience replay mechanism to generate a hierarchical acquisition strategy model.

[0009] Deploy the hierarchical acquisition strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fusion feature vector, and evaluate the feature drift degree. When the feature drift degree exceeds the preset feature threshold, trigger the first network to re-plan the global acquisition strategy based on the current fusion feature vector, and input the updated global acquisition strategy into the second network for parameter fine-tuning. Real-time monitor the acquisition performance indicators. When the performance indicators are lower than the preset target threshold, start the progressive optimization mechanism, input the newly collected state transition sequence into the training cache pool, and perform incremental learning based on historical optimization experience to realize the continuous optimization of the hierarchical acquisition strategy model.

[0010] Based on the multi-modal feature matrix, calculate the feature weight coefficient through an improved attention mechanism, adaptively fuse different modal features, and output the fusion feature vector. Use the fusion feature vector and combine the genetic algorithm to calculate the optimal initial acquisition parameter combination, including the concurrent acquisition threshold, cache capacity threshold, and batch size threshold, and generate the initial state space vector including:

[0011] The multi-modal feature matrix contains the data stream fluctuation spectrum feature values reflecting the dynamic changes of the data flow, the fine-grained type distribution feature values reflecting the data distribution law, and the multi-scale time series feature values reflecting the time series fluctuation law. Based on the data stream fluctuation spectrum feature values, the fine-grained type distribution feature values, and the multi-scale time series feature values, construct a data multi-dimensional representation feature matrix.

[0012] Calculate the statistical distribution characteristics of the data multi-dimensional representation feature matrix, select the corresponding normalization strategy according to the statistical distribution characteristics, use the maximum-minimum normalization for the data stream fluctuation spectrum feature values to maintain the peak-valley relationship, use the softmax normalization for the fine-grained type distribution feature values to highlight the main types, and use the z-score normalization for the multi-scale time series feature values to eliminate the dimension influence, and generate the normalized feature matrix.

[0013] Analyze the information entropy distribution of each modal feature in the standardized feature matrix, use the information entropy distribution as the modal importance index, construct a modal adaptive factor based on the information entropy distribution, and the modal adaptive factor is dynamically adjusted with the information entropy distribution to adaptively capture the information volume feature, and integrate the modal adaptive factor into the attention mechanism to calculate the dynamic attention weight;

[0014] Use the dynamic attention weight to construct a two-layer feature fusion network. Perform intra-modal self-attention operation in the first layer to extract intra-modal correlation, and perform cross-modal attention operation in the second layer to model the modal complementary relationship. Fusion the outputs of the two layers of attention through a residual structure and a normalization layer to obtain a fusion feature vector containing multi-modal complementary information;

[0015] Perform component analysis on the fusion feature vector. Use the periodic and peak characteristics of the data stream fluctuation spectrum eigenvalue to determine the concurrent acquisition threshold boundary, use the entropy value and skewness characteristics of the fine-grained type distribution eigenvalue to determine the cache capacity threshold boundary, and use the autocorrelation and stationarity characteristics of the multi-scale time series eigenvalue to determine the batch processing threshold boundary. The three threshold boundaries form a parameter search space;

[0016] Extract the feature importance based on the principal component analysis result of the fusion feature vector, map the feature importance to a weight coefficient, and construct a multi-objective fitness function including system throughput, processing delay, and resource utilization. The multi-objective fitness function dynamically weights the three threshold boundaries according to the weight coefficient;

[0017] Initialize the genetic algorithm population in the parameter search space, use the multi-objective fitness function to evaluate the quality of the population individuals, adaptively adjust the crossover probability based on the population diversity level to balance local and global searches, and adaptively adjust the mutation probability based on the optimization convergence degree to prevent premature convergence. Obtain the optimal parameter combination through iterative optimization;

[0018] Perform deep feature embedding on the optimal parameter combination and the fusion feature vector to generate a state space vector that couples parameter information and feature representation.

[0019] Based on the initial state space vector and the acquisition parameter combination, construct a two-layer deep Q-learning network structure. Set the input of the first network as the fusion feature vector and output the global acquisition strategy; set the input of the second network as the initial state space vector and output the specific parameter adjustment plan, including:

[0020] Construct a double-layer deep Q-learning network structure based on the initial state space vector and the acquisition parameter combination. Set the fused feature vector as the input of the first network and the initial state space vector as the input of the second network. The first network and the second network adopt a dual-network structure with separate target networks and evaluation networks;

[0021] Construct a two-stream feature extraction channel in the first network. The first channel extracts the static features in the fused feature vector through a multi-layer fully connected network to obtain a static feature representation. The second channel extracts the dynamic features in the fused feature vector through a bidirectional LSTM network to obtain a dynamic feature representation. The static feature representation and the dynamic feature representation are adaptively fused through a multi-head attention mechanism and then a global acquisition strategy is output. The global acquisition strategy includes an adjustment direction, an adjustment step size, and an adjustment priority. The adjustment direction is used to indicate whether to increase or decrease the parameter. The adjustment step size is used to control the magnitude of the parameter change. The adjustment priority is used to determine the order of multi-parameter adjustments;

[0022] In the second network, based on the initial state space vector, use the Dueling architecture to decompose the Q value into a state value function and an advantage function. The state value function evaluates the parameter configuration performance through a three-layer feedforward neural network to obtain a performance evaluation value. The advantage function evaluates the adjustment action advantage through a deep residual network to obtain an action advantage value. The performance evaluation value and the action advantage value are weighted and combined and then a specific parameter adjustment plan is output.

[0023] Deploy the hierarchical acquisition strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time feature vector and the fused feature vector, and evaluate the feature drift degree; when the feature drift degree exceeds the preset feature threshold, trigger the first network to re-plan the acquisition strategy based on the current fused feature vector, and input the updated global strategy into the second network for parameter fine-tuning, including:

[0024] Deploy the hierarchical acquisition strategy model to the production environment to continuously collect data stream features to obtain a real-time feature vector. Calculate the directional offset degree of the real-time feature vector and the fused feature vector on continuous features through cosine similarity to obtain a continuous feature similarity. Calculate the distribution change amplitude of the real-time feature vector and the fused feature vector on discrete features through KL divergence to obtain a discrete feature similarity;

[0025] Normalize and weight - fuse the continuous feature similarity and the discrete feature similarity to obtain a comprehensive feature drift index, and construct a multi - scale feature drift detection mechanism based on the comprehensive feature drift index; Smooth the comprehensive feature drift index by the exponentially weighted moving average method within the first time window to obtain a feature trend, identify the jump positions of the comprehensive feature drift index by a mutation point detection algorithm within the second time window to obtain feature change points, and capture the gradual change process of the comprehensive feature drift index by a trend analysis method within the third time window to obtain a feature trend;

[0026] Input the feature trend, the feature change points, and the feature trend into a decision - fusion module, compare the feature drift index at each time scale with the corresponding preset feature threshold, and generate an evaluation result of the feature drift degree;

[0027] When the evaluation result of the feature drift degree indicates that the feature drift index at any time scale exceeds the corresponding preset feature threshold, recalculate the feature importance weight matrix based on the current fused feature vector, and input the feature importance weight matrix into the first network;

[0028] The first network re - plans the acquisition strategy according to the feature importance weight matrix, updates the feature extraction scheme to obtain a feature weight update strategy, adjusts the parameter change direction and change step based on the feature weight update strategy to obtain a global strategy update scheme, and inputs the global strategy update scheme into the second network for parameter fine - tuning.

[0029] Monitor the acquisition performance indicators in real - time. When the performance indicators are lower than the preset target threshold, start a progressive tuning mechanism, input the newly acquired state transition sequence into the training cache pool, and perform incremental learning based on historical tuning experience to continuously optimize the hierarchical acquisition strategy model, including:

[0030] Construct a multi - dimensional performance monitoring model, collect efficiency index data such as data throughput, data latency, and resource utilization rate in real - time, collect quality index data such as data integrity, data accuracy, and data timeliness, and collect cost index data such as computing overhead, storage overhead, and bandwidth overhead; calculate the index weights of the efficiency index data, the quality index data, and the cost index data respectively, and perform weighted combination of the index weights and the corresponding index data to obtain a comprehensive performance score; compare the comprehensive performance score with the preset target threshold, and trigger a parameter optimization process when the comprehensive performance score is lower than the preset target threshold;

[0031] In the parameter optimization process, a parameter sampling set is obtained by randomly sampling the parameter space based on the Thompson sampling strategy, and the uncertainty of each parameter interval in the parameter sampling set is calculated to obtain the parameter exploration priority; according to the parameter exploration priority, the parameter sampling set is sorted by priority to obtain a parameter tuning sequence, and the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence are calculated; the state transition data generated during the optimization process is collected, and the state transition data is stored in the first cache layer, the second cache layer, and the third cache layer respectively according to the data timestamp to construct a hierarchical training cache pool;

[0032] Calculate the timeliness weight and value weight of the experience data in the hierarchical training cache pool, and multiply the timeliness weight and the value weight to obtain the experience sampling weight; sample a training data set from the hierarchical training cache pool according to the experience sampling weight, analyze the common features in the training data set to obtain a transferable knowledge representation; use the transferable knowledge representation to initialize the parameters of the policy model to obtain an initial policy model, and compress the initial policy model into a lightweight inference model through the knowledge distillation method; deploy the lightweight inference model and start the performance evaluation, calculate the performance difference before and after optimization to obtain the optimization improvement amplitude, and when the optimization improvement amplitude in multiple consecutive rounds is lower than the preset improvement threshold, save the current tuning state to obtain a tuning snapshot and stop the optimization process.

[0033] Using the transferable knowledge representation to initialize the parameters of the policy model to obtain an initial policy model, and compressing the initial policy model into a lightweight inference model through the knowledge distillation method; deploying the lightweight inference model and starting the performance evaluation, calculating the performance difference before and after optimization to obtain the optimization improvement amplitude, and when the optimization improvement amplitude in multiple consecutive rounds is lower than the preset improvement threshold, saving the current tuning state to obtain a tuning snapshot and stopping the optimization process includes:

[0034] Use the transferable knowledge representation to initialize the parameters of the corresponding network layer in the policy model to obtain an initial policy model; use the initial policy model as the teacher model, extract the policy distribution output and value function prediction of the teacher model as soft label information, and construct a student model with fewer layers and a parameter sharing mechanism;

[0035] Based on the soft label information and the output of the student model, construct a knowledge distillation loss function, where the knowledge distillation loss function is linearly combined by three terms: the KL divergence of the policy distribution, the mean squared error of the value function prediction, and the task objective function through weight coefficients; use the knowledge distillation loss function to optimize the parameters of the student model to obtain a lightweight inference model, and deploy the lightweight inference model and start the performance evaluation;

[0036] During the performance evaluation process, the matching degree between the output actions of the lightweight inference model and the optimal actions is calculated respectively to obtain the policy accuracy rate, and the deviation between the predicted value function and the true value function is calculated to obtain the value function error; the policy accuracy rate and the value function error are calculated according to the preset weights to obtain the comprehensive performance index, and the optimization improvement amplitude is calculated based on the comprehensive performance index of the current round and the previous round.

[0037] A sliding time window is set. When the optimization improvement amplitudes in multiple consecutive rounds within the sliding time window are all lower than the preset improvement threshold, the network structure parameters, optimizer state, and performance indicators of the lightweight inference model are saved as a tuning snapshot, and the optimization process is stopped and the tuning snapshot is output.

[0038] According to the parameter exploration priority, the parameter sampling set is sorted by priority to obtain a parameter tuning sequence, and the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence are calculated; a confidence boundary adjustment strategy is constructed based on the estimated mean and the estimated standard deviation, and the parameter optimization direction is determined by adaptively adjusting the learning rate and the confidence coefficient to obtain an optimized parameter scheme, including:

[0039] According to the parameter exploration priority, the parameter sampling set is sorted by priority to obtain a parameter tuning sequence, and the correlation degree between adjacent parameters in the parameter tuning sequence is calculated. The parameters with a correlation degree greater than the preset correlation threshold are combined into a parameter group.

[0040] A historical data sampling window of the parameter tuning sequence is established, and time decay weights are applied to the data in the historical data sampling window to obtain a weighted historical sample set; in the weighted historical sample set, the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence are calculated, and the joint distribution characteristics of the parameters within the parameter group are calculated at the same time. The estimated mean and the estimated standard deviation reflect the distribution characteristics of the parameters.

[0041] A confidence boundary adjustment strategy is constructed based on the estimated mean and the estimated standard deviation, and the upper and lower boundaries of the parameter confidence interval are determined in combination with the joint distribution characteristics. The confidence boundary adjustment strategy includes a single-parameter confidence interval and a parameter-group confidence interval; the deviation degree between the current parameter value and the parameter confidence interval is calculated, and the combined deviation between the current value of the parameter group and the parameter-group confidence interval is calculated at the same time. The parameter adjustment reference direction is determined based on the deviation degree and the combined deviation.

[0042] Analyze the performance change trend during the historical optimization of parameters, calculate the performance improvement sequence of several recent optimizations, adaptively adjust the learning rate according to the fluctuation characteristics of the performance improvement sequence, and use the ratio of the learning rate to the performance improvement amplitude as the confidence coefficient; extract several optimization directions with the most significant performance improvement in the historical optimization path of parameters, perform weighted combination on the several optimization directions and the parameter adjustment reference direction, correct the combined direction according to the confidence coefficient to obtain the parameter optimization direction, and multiply the parameter optimization direction by the learning rate to obtain the optimization parameter scheme.

[0043] In the second aspect of the embodiments of the present invention,

[0044] Provide a database adaptive data stream acquisition optimization system based on reinforcement learning, including:

[0045] The first unit is used to construct a distributed data acquisition agent network, obtain the data stream information of the database in real time, and generate an initial feature set; perform multi-modal decomposition on the initial feature set, extract the data stream fluctuation spectrum feature value, fine-grained type distribution feature value, multi-scale time series feature value and graph structure association feature value, and form a multi-modal feature matrix; based on the multi-modal feature matrix, calculate the feature weight coefficient through an improved attention mechanism, perform adaptive fusion on different modal features, and output a fused feature vector; use the fused feature vector, combine with a genetic algorithm to calculate the optimal initial acquisition parameter combination, including the concurrent acquisition threshold, cache capacity threshold and batch processing size threshold, and generate an initial state space vector;

[0046] The second unit is used to construct a two-layer deep Q-learning network structure based on the initial state space vector and the acquisition parameter combination, set the input of the high-level network as the fused feature vector, and output a global acquisition strategy; set the input of the low-level network as the initial state space vector, and output a specific parameter adjustment scheme; design a hierarchical reward function, use the feature weight coefficient as the balance factor of the reward function, and establish a multi-objective evaluation system including acquisition delay rate, data consistency, throughput efficiency and resource utilization rate; perform a hierarchical training process based on the multi-objective evaluation system, use the global acquisition strategy output by the high-level network as the constraint condition of the low-level network, and iteratively update the network parameters through a priority experience replay mechanism to generate a hierarchical acquisition strategy model;

[0047] The third unit is used to deploy the hierarchical acquisition strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fusion feature vector, and evaluate the degree of feature drift. When the degree of feature drift exceeds the preset feature threshold, the high-level network is triggered to re-plan the acquisition strategy based on the current fusion feature vector, and the updated global strategy is input into the low-level network for parameter fine-tuning. The acquisition performance metrics are monitored in real time. When the performance metrics are lower than the preset target threshold, a progressive optimization mechanism is started, the newly collected state transition sequence is input into the training cache pool, and incremental learning is performed based on historical optimization experience to continuously optimize the hierarchical acquisition strategy model.

[0048] In the third aspect of the embodiments of the present invention,

[0049] there is provided an electronic device, including:

[0050] a processor;

[0051] a memory for storing instructions executable by the processor;

[0052] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0053] In the fourth aspect of the embodiments of the present invention,

[0054] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0055] The beneficial effects of this application are as follows:

[0056] 1. Adaptive optimization of database data stream acquisition is realized. Through multi-modal feature extraction and adaptive fusion, as well as a double-layer Q-learning network based on reinforcement learning, the acquisition parameters can be dynamically adjusted according to the changes in data stream features, improving the acquisition efficiency.

[0057] 2. The efficiency and stability of data stream acquisition are improved. The design of the hierarchical reward function and the multi-objective evaluation system, combined with the priority experience replay mechanism, can effectively balance the acquisition delay rate, throughput efficiency, and resource utilization rate, ensuring the stability of the acquisition performance.

[0058] 3. The cost of manual intervention is reduced, and continuous optimization of the acquisition strategy is realized. The feature drift detection and progressive optimization mechanism can automatically identify the changes in data stream features and make corresponding parameter adjustments, reducing the need for manual intervention and realizing the continuous optimization of the acquisition strategy to adapt to the changing data stream environment. Description of the Drawings

[0059] Figure 1Schematic flowchart of the database adaptive data stream acquisition optimization method based on reinforcement learning according to an embodiment of the present invention;

[0060] Figure 2 Schematic structural diagram of the database adaptive data stream acquisition optimization system based on reinforcement learning according to an embodiment of the present invention. Detailed implementation manners

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0063] Figure 1 Schematic flowchart of the database adaptive data stream acquisition optimization method based on reinforcement learning according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0064] S11. Construct a distributed data collection agent network to obtain the data stream information of the database in real time and generate an initial feature set; perform multimodal decomposition on the initial feature set to extract the data stream fluctuation spectrum feature value, fine-grained type distribution feature value, multi-scale time series feature value, and graph structure association feature value, and form a multimodal feature matrix; based on the multimodal feature matrix, calculate the feature weight coefficient through an improved attention mechanism, adaptively fuse different modal features, and output a fused feature vector; use the fused feature vector and combine with a genetic algorithm to calculate the optimal initial acquisition parameter combination, including the concurrent acquisition threshold, cache capacity threshold, and batch processing size threshold, and generate an initial state space vector;

[0065] S12. Based on the initial state space vector and the acquisition parameter combination, construct a double-layer deep Q-learning network structure. Set the input of the first network as the fusion feature vector and output the global acquisition strategy. Set the input of the second network as the initial state space vector and output the specific parameter adjustment scheme. Design a hierarchical reward function, use the feature weight coefficient as the balance factor of the reward function, and establish a multi-objective evaluation system including the acquisition delay rate, throughput efficiency, and resource utilization rate. Based on the multi-objective evaluation system, perform a hierarchical training process. Use the global acquisition strategy output by the first network as the constraint condition of the second network, and iteratively update the network parameters through the prioritized experience replay mechanism to generate a hierarchical acquisition strategy model.

[0066] S13. Deploy the hierarchical acquisition strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fusion feature vector, and evaluate the feature drift degree. When the feature drift degree exceeds the preset feature threshold, trigger the first network to re-plan the global acquisition strategy based on the current fusion feature vector, and input the updated global acquisition strategy into the second network for parameter fine-tuning. Monitor the acquisition performance indicators in real time. When the performance indicators are lower than the preset target threshold, start the progressive optimization mechanism, input the newly collected state transition sequence into the training cache pool, and perform incremental learning based on historical optimization experience to continuously optimize the hierarchical acquisition strategy model.

[0067] In an alternative implementation, based on the multi-modal feature matrix, calculate the feature weight coefficient through an improved attention mechanism, adaptively fuse different modal features, and output the fusion feature vector. Use the fusion feature vector and combine the genetic algorithm to calculate the optimal initial acquisition parameter combination, including the concurrent acquisition threshold, cache capacity threshold, and batch size threshold, and generate the initial state space vector including:

[0068] The multi-modal feature matrix contains the data stream fluctuation spectrum feature values reflecting the dynamic changes of the data flow, the fine-grained type distribution feature values reflecting the data distribution law, and the multi-scale time series feature values reflecting the time series fluctuation law. Based on the data stream fluctuation spectrum feature values, the fine-grained type distribution feature values, and the multi-scale time series feature values, construct a data multi-dimensional representation feature matrix.

[0069] Calculate the statistical distribution characteristics of the data multi-dimensional representation feature matrix, select the corresponding normalization strategy according to the statistical distribution characteristics, use the maximum-minimum normalization for the data stream fluctuation spectrum feature values to maintain the peak-valley relationship, use the softmax normalization for the fine-grained type distribution feature values to highlight the main types, and use the z-score standardization for the multi-scale time series feature values to eliminate the dimension influence, and generate the standardized feature matrix.

[0070] Analyze the information entropy distribution of each modal feature in the standardized feature matrix, use the information entropy distribution as the modal importance index, construct a modal adaptive factor based on the information entropy distribution, and the modal adaptive factor is dynamically adjusted with the information entropy distribution to adaptively capture the information volume feature, and integrate the modal adaptive factor into the attention mechanism to calculate the dynamic attention weight;

[0071] Use the dynamic attention weight to construct a two-layer feature fusion network. In the first layer, perform intra-modal self-attention operation to extract intra-modal correlation, and in the second layer, perform inter-modal cross-attention operation to model the modal complementary relationship. Fusion the outputs of the two layers of attention through a residual structure and a normalization layer to obtain a fusion feature vector containing multi-modal complementary information;

[0072] Perform component analysis on the fusion feature vector. Use the period and peak characteristics of the data stream fluctuation spectrum eigenvalue to determine the concurrent acquisition threshold boundary, use the entropy value and skewness characteristics of the fine-grained type distribution eigenvalue to determine the cache capacity threshold boundary, and use the autocorrelation and stationarity characteristics of the multi-scale time series eigenvalue to determine the batch processing threshold boundary. The three threshold boundaries form a parameter search space;

[0073] Extract the feature importance based on the principal component analysis result of the fusion feature vector, map the feature importance to a weight coefficient, and construct a multi-objective fitness function including system throughput, processing delay, and resource utilization. The multi-objective fitness function dynamically weights the three threshold boundaries according to the weight coefficient;

[0074] Initialize the genetic algorithm population in the parameter search space, use the multi-objective fitness function to evaluate the quality of the population individuals, adaptively adjust the crossover probability based on the population diversity level to balance local and global search, and adaptively adjust the mutation probability based on the optimization convergence degree to prevent premature convergence. Obtain the optimal parameter combination through iterative optimization;

[0075] Perform deep feature embedding on the optimal parameter combination and the fusion feature vector to generate a state space vector that couples parameter information and feature representation.

[0076] First, construct a multi-dimensional representation feature matrix for the data. This matrix contains three key features: the spectral eigenvalue of the data flow fluctuation, which is used to reflect the frequency and amplitude information of the dynamic changes in the data flow; the fine-grained type distribution eigenvalue, which is used to reflect the distribution of data types; and the multi-scale time series eigenvalue, which is used to reflect the variation law of the data at different time scales. For example, for a network traffic data flow, the number of packets per second can be extracted as the spectral eigenvalue of the data flow fluctuation, the proportion of the number of packets of different protocol types as the fine-grained type distribution eigenvalue, and the change trend of the number of packets per minute and per hour as the multi-scale time series eigenvalue.

[0077] Next, perform normalization processing on the constructed multi-dimensional representation feature matrix. Different normalization strategies are adopted for different feature types: the maximum-minimum normalization is used for the spectral eigenvalue of the data flow fluctuation. For example, the number of packets per second is scaled to between 0 and 1 to maintain the relative relationship between the peaks and valleys; the softmax normalization is used for the fine-grained type distribution eigenvalue. For example, the proportion of various protocol types is converted into a probability distribution to highlight the main data types; the z-score normalization is used for the multi-scale time series eigenvalue. For example, the number of packets per minute and per hour is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1 to eliminate the dimensionality impact brought by different time scales. Suppose the spectral eigenvalue of the data flow fluctuation is [100, 200, 150], the maximum value is 200, and the minimum value is 100, then the normalized result is [0, 1, 0.5].

[0078] Then, analyze the information entropy distribution of each modal feature in the normalized feature matrix and use it as the modal importance index. The higher the information entropy, the richer the information contained in the modal feature. Based on the information entropy distribution, construct a modal adaptive factor, which will be dynamically adjusted according to the information entropy distribution to adaptively capture the information volume feature. For example, if the information entropy of the spectral eigenvalue of the data flow fluctuation is high, its corresponding modal adaptive factor will also be high. Integrate the modal adaptive factor into the attention mechanism to calculate the dynamic attention weight.

[0079] Use the dynamic attention weight to construct a two-layer feature fusion network. The first layer performs intra-modal self-attention operations to extract the correlations within the modality. For example, analyze the relationship between different frequency components within the spectral eigenvalue of the data flow fluctuation. The second layer performs inter-modal cross-attention operations to model the complementary relationships between modalities. For example, analyze the mutual influence between the spectral eigenvalue of the data flow fluctuation and the fine-grained type distribution eigenvalue. Fuse the outputs of the two layers of attention through a residual structure and a normalization layer to obtain a fusion feature vector containing multi-modal complementary information.

[0080] Perform component analysis on the fused feature vector to determine the search space of the acquisition parameters. Determine the concurrent acquisition threshold boundary using the periodic and peak characteristics of the data stream fluctuation spectrum eigenvalues. For example, determine the frequency range of concurrent acquisition based on the time interval when the peak of the number of data packets appears. Determine the cache capacity threshold boundary using the entropy and skewness characteristics of the fine-grained type distribution eigenvalues. For example, determine the size range of the cache capacity based on the distribution of different data types. Determine the batch processing threshold boundary using the autocorrelation and stationarity characteristics of the multi-scale time series eigenvalues. For example, determine the size range of batch processing based on the variation law of the data stream at different time scales. These three threshold boundaries together constitute the parameter search space.

[0081] Extract the feature importance based on the principal component analysis result of the fused feature vector and map it to a weight coefficient. Construct a multi-objective fitness function that includes system throughput, processing delay, and resource utilization. This function dynamically weights the three threshold boundaries according to the weight coefficient. For example, if the importance of system throughput is higher, a higher weight will be assigned to it in the fitness function.

[0082] Initialize the genetic algorithm population within the parameter search space and evaluate the quality of the population individuals using the multi-objective fitness function. Adaptively adjust the crossover probability based on the population diversity level to balance local and global searches. Adaptively adjust the mutation probability based on the optimization convergence degree to prevent premature convergence. Through iterative optimization, finally obtain the optimal parameter combination. For example, the concurrent acquisition threshold is set to 1000, the cache capacity threshold is set to 10MB, and the batch processing size threshold is set to 1024.

[0083] Finally, perform deep feature embedding on the optimal parameter combination and the fused feature vector to generate a state space vector that couples parameter information and feature representations. This vector can be used for subsequent data acquisition control and performance optimization.

[0084] The solution of this application can:

[0085] Improve data acquisition efficiency: By dynamically adjusting the acquisition parameters, data can be adaptively acquired according to the real-time characteristics of the data stream, avoiding redundant data acquisition, thereby improving data acquisition efficiency. Enhance system performance: By optimizing the acquisition parameters, the system throughput, processing delay, and resource utilization can be effectively balanced, thereby enhancing the overall system performance. Enhance system adaptability: This method can dynamically adjust the acquisition parameters according to the characteristics of different data streams, enhancing the system's adaptability to different data streams.

[0086] In an alternative embodiment, based on the initial state space vector and the acquisition parameter combination, a two-layer deep Q-learning network structure is constructed. The input of the first network is set as the fusion feature vector, and the global acquisition strategy is output. The input of the second network is set as the initial state space vector, and the specific parameter adjustment scheme is output, including:

[0087] Based on the initial state space vector and the acquisition parameter combination, a two-layer deep Q-learning network structure is constructed. The fusion feature vector is set as the input of the first network, and the initial state space vector is set as the input of the second network. The first network and the second network adopt a dual-network structure with separation of the target network and the evaluation network.

[0088] In the first network, a two-stream feature extraction channel is constructed. The first channel extracts the static features in the fusion feature vector through a multi-layer fully connected network to obtain a static feature representation. The second channel extracts the dynamic features in the fusion feature vector through a bidirectional LSTM network to obtain a dynamic feature representation. The static feature representation and the dynamic feature representation are adaptively fused through a multi-head attention mechanism and then the global acquisition strategy is output. The global acquisition strategy includes an adjustment direction, an adjustment step size, and an adjustment priority. The adjustment direction is used to indicate parameter increase or decrease. The adjustment step size is used to control the parameter change range. The adjustment priority is used to determine the multi-parameter adjustment order.

[0089] In the second network, based on the initial state space vector, the Q value is decomposed into a state value function and an advantage function by using a Dueling architecture. The state value function evaluates the parameter configuration performance through a three-layer feedforward neural network to obtain a performance evaluation value. The advantage function evaluates the adjustment action advantage through a deep residual network to obtain an action advantage value. The performance evaluation value and the action advantage value are weighted and combined and then the specific parameter adjustment scheme is output.

[0090] First, the initial state space vector of the system and the acquisition parameter combination are obtained. The initial state space vector is used to describe the initial state of the system. The acquisition parameter combination includes the parameters to be adjusted and their initial values. For example, the initial state space vector of a control system may include temperature, pressure, flow rate, etc. The acquisition parameter combination may include the proportional coefficient, integral coefficient, and differential coefficient of a PID controller. Assume the initial temperature is 25 °C, the pressure is 1 standard atmosphere, the flow rate is 10 m / s, the proportional coefficient is 1, the integral coefficient is 0.1, and the differential coefficient is 0.01.

[0091] Then, construct a fused feature vector. The fused feature vector integrates the static features and dynamic features of the system. Static features describe the inherent properties of the system, such as device model, parameter range, etc. Dynamic features describe the real-time state of the system, such as current performance metrics, historical adjustment records, etc. Suppose the device model is type A, the parameter range is from 0 to 10, the current performance metric is 80%, and the past 5 parameter adjustment records are an increase of 0.1, a decrease of 0.05, an increase of 0.2, no change, and a decrease of 0.1. Integrate this information into a fused feature vector.

[0092] Next, construct a two-layer deep Q-learning network structure. The input of the first network is the fused feature vector, and the output is the global acquisition strategy, including the adjustment direction (increase or decrease), adjustment step size, and adjustment priority. The input of the second network is the initial state space vector, and the output is the specific parameter adjustment plan. Both networks adopt a dual-network structure with separate target networks and evaluation networks to improve training stability.

[0093] In the first network, construct two-stream feature extraction channels. The first channel uses a multi-layer fully connected network to extract the static features in the fused feature vector to obtain a static feature representation. The second channel uses a bidirectional LSTM network to extract the dynamic features in the fused feature vector to obtain a dynamic feature representation. Then, use the multi-head attention mechanism to adaptively fuse the static feature representation and the dynamic feature representation, and finally output the global acquisition strategy. For example, the global acquisition strategy that the first network may output is: first adjust the proportionality coefficient, increase by 0.2; second, adjust the integral coefficient, decrease by 0.01; finally, adjust the differential coefficient, increase by 0.005.

[0094] In the second network, based on the initial state space vector, use the Dueling architecture to decompose the Q value into a state value function and an advantage function. The state value function uses a three-layer feedforward neural network to evaluate the performance of the parameter configuration to obtain a performance evaluation value. The advantage function uses a deep residual network to evaluate the advantage of the adjustment action to obtain an action advantage value. Then, perform a weighted combination of the performance evaluation value and the action advantage value, and output the specific parameter adjustment plan. For example, the specific parameter adjustment plan that the second network may output is: adjust the proportionality coefficient to 1.2, the integral coefficient to 0.09, and the differential coefficient to 0.015.

[0095] Finally, according to the global acquisition strategy and the specific parameter adjustment plan, adjust the system parameters and observe the change in system performance. Feed the new state space vector and the adjustment result back to the network for the next round of learning and adjustment. Repeat this process until the system performance reaches the preset target or the maximum number of iterations is reached.

[0096] The solution of this application can:

[0097] Improve the efficiency of parameter adjustment: Through the joint optimization of the global acquisition strategy and the specific parameter adjustment plan, ineffective adjustment attempts can be reduced, and the speed of parameter optimization can be accelerated. Enhance the accuracy of parameter adjustment: The application of the double-layer network structure and deep learning algorithms can more accurately capture the complex relationship between the system state and parameters, thus achieving more refined parameter adjustment. Strengthen the robustness of parameter adjustment: By integrating static features and dynamic features, and combining the target network and the evaluation network, the stability and generalization ability of the algorithm can be improved, enabling it to maintain good performance in different environments and working conditions.

[0098] In an optional implementation manner, the hierarchical acquisition strategy model is deployed to the production environment to continuously acquire data stream features, calculate the similarity between the real-time feature vector and the fusion feature vector, and evaluate the degree of feature drift; when the degree of feature drift exceeds the preset feature threshold, based on the current fusion feature vector, trigger the first network to re-plan the acquisition strategy, and input the updated global strategy into the second network for parameter fine-tuning, including:

[0099] Deploy the hierarchical acquisition strategy model to the production environment to continuously acquire data stream features to obtain a real-time feature vector, calculate the directional offset degree of the real-time feature vector and the fusion feature vector on continuous features through cosine similarity to obtain the continuous feature similarity, and calculate the distribution change amplitude of the real-time feature vector and the fusion feature vector on discrete features through KL divergence to obtain the discrete feature similarity;

[0100] Normalize and weighted-fuse the continuous feature similarity and the discrete feature similarity to obtain a comprehensive feature drift index, and construct a multi-scale feature drift detection mechanism based on the comprehensive feature drift index; Smooth the comprehensive feature drift index through the exponential weighted moving average method within the first time window to obtain the feature trend, identify the jump position of the comprehensive feature drift index through the mutation point detection algorithm within the second time window to obtain the feature change point, and capture the gradual change process of the comprehensive feature drift index through the trend analysis method within the third time window to obtain the feature trend;

[0101] Input the feature trend, the feature change point, and the feature trend into the decision fusion module, compare the feature drift index at each time scale with the corresponding preset feature threshold, and generate an evaluation result of the degree of feature drift;

[0102] When the evaluation result of the degree of feature drift indicates that the feature drift index at any time scale exceeds the corresponding preset feature threshold, recalculate the feature importance weight matrix based on the current fusion feature vector, and input the feature importance weight matrix into the first network;

[0103] The first network re-plans the acquisition strategy according to the feature importance weight matrix, updates the feature extraction scheme to obtain a feature weight update strategy, adjusts the parameter change direction and change step size based on the feature weight update strategy to obtain a global strategy update scheme, and inputs the global strategy update scheme into the second network for parameter fine-tuning.

[0104] First, deploy the trained hierarchical acquisition strategy model to the production environment. The model consists of two sub-networks: the first network is responsible for planning the acquisition strategy according to feature importance, and the second network is responsible for feature extraction and parameter fine-tuning. The initial acquisition strategy and feature extraction scheme are obtained through offline training.

[0105] Next, the model starts to continuously collect data stream features. Assume that the data stream contains features such as user age, purchase times, and product categories. The collected data undergoes preprocessing and feature engineering and is transformed into real-time feature vectors. For example, the user feature vector collected at a certain moment is: [25, 3, 'electronic products'].

[0106] Meanwhile, the system maintains a fused feature vector, which represents the overall features of the data distribution over a past period of time. The fused feature vector is updated through a regular update mechanism, such as once per hour, by incorporating the newly collected data features. Methods such as moving window average or exponentially weighted moving average can be used for the update of the fused feature vector. Assume that the current fused feature vector is: [28, 2, 'clothing'].

[0107] To evaluate the degree of feature drift, it is necessary to calculate the similarity between the real-time feature vector and the fused feature vector. For continuous features, such as user age and purchase times, cosine similarity is used to measure the degree of direction offset. For example, the cosine similarity calculation result for user age is 0.95. For discrete features, such as product categories, KL divergence is used to measure the degree of distribution change. For example, the KL divergence calculation result for product categories is 0.2.

[0108] Normalize and weighted fuse the continuous feature similarity and the discrete feature similarity to obtain a comprehensive feature drift index. For example, assume that the continuous feature weight is 0.7 and the discrete feature weight is 0.3, then the comprehensive feature drift index is 0.7 * 0.95 + 0.3 * 0.2 = 0.725.

[0109] To detect feature drift more comprehensively, a multi-scale feature drift detection mechanism is constructed. In the first shorter time window, such as the past 1 hour, the exponential weighted moving average method is used to smooth the comprehensive index of feature drift to obtain the feature trend, so as to filter out short-term fluctuations. In the second medium time window, such as the past 1 day, the change point detection algorithm is used to identify the jump positions of the comprehensive index of feature drift to obtain the feature change points, so as to capture sudden drifts. In the third longer time window, such as the past 1 week, the trend analysis method is used to capture the gradual change process of the comprehensive index of feature drift to obtain the feature trend, so as to discover long-term slow drifts.

[0110] The feature trend, feature change points, and feature trend are input into the decision fusion module, and the feature drift index at each time scale is compared with the corresponding preset feature threshold. For example, assume that the short-term feature trend threshold is 0.1, the medium-term feature change point threshold is 0.2, and the long-term feature trend threshold is 0.3. If the feature drift index at any time scale exceeds the corresponding preset threshold, it is considered that a significant feature drift has occurred.

[0111] When the evaluation result of the feature drift degree indicates that the feature drift index at any time scale exceeds the corresponding preset feature threshold, the system will recalculate the feature importance weight matrix based on the current fused feature vector. For example, if the drift of the user's age is large, its importance weight will be increased accordingly. Then, the feature importance weight matrix is input into the first network.

[0112] The first network re-plans the acquisition strategy according to the feature importance weight matrix and updates the feature extraction scheme to obtain the feature weight update strategy. For example, if the importance weight of the user's age is increased, the acquisition strategy may be more inclined to collect age-related data, and the feature extraction scheme may pay more attention to the processing of age features. Based on the feature weight update strategy, the change direction and change step size of the model parameters are adjusted to obtain the global strategy update scheme.

[0113] Finally, the global strategy update scheme is input into the second network for parameter fine-tuning to adapt to the new data distribution.

[0114] The solution of this application can:

[0115] Improve data collection efficiency: By dynamically adjusting the collection strategy according to feature importance, redundant or useless data collection can be avoided, thereby improving data collection efficiency and reducing storage and computing costs. Enhance model robustness: Through real-time monitoring and adaptive adjustment, the model can respond to changes in data distribution in a timely manner, maintain the prediction performance and stability of the model, and enhance the robustness of the model. Achieve automated operation and maintenance: This method realizes the automation of feature drift detection and model parameter adjustment, reduces the need for manual intervention, lowers operation and maintenance costs, and improves system efficiency.

[0116] In an alternative embodiment, the collection performance metrics are monitored in real time. When the performance metrics are lower than a preset target threshold, a progressive tuning mechanism is activated, and the newly collected state transition sequence is input into the training cache pool. Incremental learning is performed based on historical tuning experience, and continuous optimization of the hierarchical collection strategy model includes:

[0117] Construct a multi-dimensional performance monitoring model, and collect efficiency metric data such as data throughput, data latency, and resource utilization in real time, collect quality metric data such as data integrity, data accuracy, and data timeliness, and collect cost metric data such as computing overhead, storage overhead, and bandwidth overhead; calculate the respective metric weights for the efficiency metric data, the quality metric data, and the cost metric data, and perform weighted combination of the metric weights and the corresponding metric data to obtain a comprehensive performance score; compare the comprehensive performance score with a preset target threshold, and trigger a parameter optimization process when the comprehensive performance score is lower than the preset target threshold;

[0118] In the parameter optimization process, randomly sample the parameter space based on the Thompson sampling strategy to obtain a parameter sampling set, calculate the uncertainty of each parameter interval in the parameter sampling set to obtain a parameter exploration priority; sort the parameter sampling set according to the parameter exploration priority to obtain a parameter tuning sequence, and calculate the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence; collect the state transition data generated during the optimization process, and store the state transition data into the first cache layer, the second cache layer, and the third cache layer respectively according to the data timestamp to construct a hierarchical training cache pool;

[0119] Calculate the timeliness weight and value weight of the experience data in the hierarchical training cache pool, multiply the timeliness weight and the value weight to obtain the experience sampling weight; sample a training data set from the hierarchical training cache pool according to the experience sampling weight, analyze the common features in the training data set to obtain a transferable knowledge representation; use the transferable knowledge representation to initialize the policy model parameters to obtain an initial policy model, and compress the initial policy model into a lightweight inference model through knowledge distillation; deploy the lightweight inference model and start performance evaluation, calculate the performance difference before and after optimization to obtain the optimization improvement amplitude, and when the optimization improvement amplitude in multiple consecutive rounds is lower than the preset improvement threshold, save the current tuning state to obtain a tuning snapshot and stop the optimization process.

[0120] First, construct a multi-dimensional performance monitoring model. This model collects index data in three dimensions in real time: Efficiency index data includes data throughput, data latency, and resource utilization rate; Quality index data includes data integrity, data accuracy, and data timeliness; Cost index data includes computing overhead, storage overhead, and bandwidth overhead. For example, set the collection period to 1 minute and record the data of these nine indicators once every minute. Suppose the data throughput collected in a certain minute is 1000 pieces / second, the data latency is 0.5 seconds, the resource utilization rate is 60%, the data integrity is 99%, the data accuracy is 98%, the data timeliness is 5 minutes, the computing overhead is 10 CPU units, the storage overhead is 1 GB, and the bandwidth overhead is 10 Mbps.

[0121] Next, calculate the comprehensive performance score. Calculate the index weights of the efficiency index, quality index, and cost index respectively. The determination of the weights can be based on methods such as expert experience and analytic hierarchy process. For example, set the efficiency index weight to 0.4, the quality index weight to 0.5, and the cost index weight to 0.1. Combine the index weights with the corresponding index data through weighted combination to obtain the comprehensive performance score. Taking the above case as an example, the comprehensive performance score = 0.4 * (throughput normalized value + latency normalized value + resource utilization rate normalized value) + 0.5 * (integrity normalized value + accuracy normalized value + timeliness normalized value) + 0.1 * (computing overhead normalized value + storage overhead normalized value + bandwidth overhead normalized value). Among them, each index needs to be normalized, such as min-max normalization, to map the index value to between 0 and 1.

[0122] Then, compare the comprehensive performance score with the preset target threshold. If the comprehensive performance score is lower than the preset target threshold, trigger the parameter optimization process. For example, the preset target threshold is 0.8. If the calculated comprehensive performance score is 0.7, trigger the parameter optimization process.

[0123] In the parameter optimization process, first, random sampling is performed on the parameter space based on the Thompson sampling strategy to obtain a parameter sampling set. For example, assuming the parameters to be adjusted are the acquisition frequency and cache size, random sampling is carried out within the preset parameter range to obtain multiple combinations of acquisition frequencies and cache sizes. Then, the uncertainty of each parameter interval in the parameter sampling set is calculated to obtain the parameter exploration priority. Uncertainty can be represented by the variance of the parameter interval or the width of the confidence interval. The parameter sampling set is sorted according to the parameter exploration priority to obtain a parameter tuning sequence. For example, parameter combinations with high uncertainty are ranked first and tested preferentially. Next, the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence are calculated. The estimated mean and estimated standard deviation can be calculated based on historical tuning data or Bayesian methods.

[0124] State transition data generated during the acquisition optimization process, such as acquisition parameters, performance metrics, environmental information, etc. According to the data timestamps, the state transition data are respectively stored in the first cache layer, the second cache layer, and the third cache layer to construct a hierarchical training cache pool. For example, data within the last 1 hour are stored in the first cache layer, data within the last 1 day are stored in the second cache layer, and data within the last 1 month are stored in the third cache layer.

[0125] Calculate the timeliness weight and value weight of the empirical data in the hierarchical training cache pool. The timeliness weight is calculated based on the data timestamp, and the closer the time, the higher the weight. The value weight is calculated based on the performance metric corresponding to the data, and the better the performance metric, the higher the weight. For example, an exponential decay function can be used to calculate the timeliness weight, and the normalized value of the performance metric can be used to calculate the value weight. Multiply the timeliness weight and the value weight to obtain the empirical sampling weight.

[0126] A training data set is sampled from the hierarchical training cache pool according to the empirical sampling weight. For example, a weighted random sampling method can be used to sample data from different cache layers according to the empirical sampling weight. Analyze the common features in the training data set to obtain a transferable knowledge representation. For example, data distribution features, parameter change rules, etc. can be extracted as transferable knowledge representations.

[0127] The transferable knowledge representation is used to initialize the parameters of the policy model to obtain an initial policy model. For example, transfer learning methods can be used to integrate the transferable knowledge representation into the initialization process of the policy model. The initial policy model is compressed into a lightweight inference model through knowledge distillation. For example, a student-teacher network structure can be used for knowledge distillation to transfer the knowledge of a large complex model to a small lightweight model.

[0128] Deploy the lightweight inference model and start the performance evaluation. Calculate the performance difference before and after optimization to obtain the optimization improvement amplitude. For example, compare the difference in the comprehensive performance scores before and after optimization. When the optimization improvement amplitude in multiple consecutive rounds is lower than the preset improvement threshold, save the current tuning state, obtain the tuning snapshot, and stop the optimization process. For example, set the preset improvement threshold to 0.05. If the optimization improvement amplitude in three consecutive rounds is less than 0.05, stop the optimization.

[0129] The solution of this application can:

[0130] Improve the acquisition efficiency and quality: Through real-time monitoring and progressive tuning, the acquisition strategy can be dynamically adjusted to make the data acquisition process more efficient while ensuring data quality. Reduce resource consumption: Through parameter optimization and knowledge distillation, the computational overhead, storage overhead, and bandwidth overhead can be reduced, and the resource consumption of data acquisition can be lowered. Enhance the system adaptability: Through continuous learning and knowledge transfer, the system can automatically adapt to different data environments and application scenarios, improving the robustness and generalization ability of the system.

[0131] In an optional implementation manner, use the transferable knowledge representation to initialize the parameters of the policy model to obtain the initial policy model, and compress the initial policy model into a lightweight inference model through the knowledge distillation method; deploy the lightweight inference model and start the performance evaluation, calculate the performance difference before and after optimization to obtain the optimization improvement amplitude. When the optimization improvement amplitude in multiple consecutive rounds is lower than the preset improvement threshold, saving the current tuning state to obtain the tuning snapshot and stopping the optimization process includes:

[0132] Use the transferable knowledge representation to initialize the parameters of the corresponding network layer in the policy model to obtain the initial policy model; use the initial policy model as the teacher model, extract the policy distribution output and value function prediction of the teacher model as soft label information, and construct a student model with fewer layers and a parameter sharing mechanism;

[0133] Based on the soft label information and the output of the student model, construct a knowledge distillation loss function, where the knowledge distillation loss function is obtained by linearly combining the policy distribution KL divergence, the mean squared error of the value function prediction, and the task objective function through weight coefficients; use the knowledge distillation loss function to optimize the parameters of the student model to obtain a lightweight inference model, and deploy the lightweight inference model and start the performance evaluation;

[0134] During the performance evaluation process, calculate the matching degree between the output action of the lightweight inference model and the optimal action to obtain the policy accuracy rate, and calculate the deviation between the predicted value function and the true value function to obtain the value function error; calculate the comprehensive performance index according to the preset weight of the policy accuracy rate and the value function error, and calculate the optimization improvement amplitude based on the comprehensive performance index of the current round and the previous round;

[0135] Set a sliding time window. When the improvement amplitudes in multiple consecutive rounds within the sliding time window are all lower than a preset improvement threshold, save the network structure parameters, optimizer state, and performance metrics of the lightweight inference model as a tuning snapshot, stop the optimization process, and output the tuning snapshot.

[0136] First, obtain transferable knowledge representations. Suppose in a simulated robot navigation task, a large policy model has been trained to successfully navigate in a complex environment. The network layer parameters (such as the weights and biases of the convolutional layer) extracted from this large model can be used as transferable knowledge representations. Specifically, the parameters of the convolutional layer responsible for feature extraction in the model can be extracted. These parameters encode a general understanding of the environment, such as the recognition of walls, obstacles, and targets. These extracted parameters will be used to initialize the corresponding network layers of the student model.

[0137] Then, initialize the policy model using the transferable knowledge representations. Directly load the extracted convolutional layer parameters into the corresponding layers of the student model (a smaller convolutional neural network). Assume that the convolutional layer structure of the student model is similar to the feature extraction layer structure of the large model, and the parameter values can be directly copied. This enables the student model to have a certain environmental perception ability at the beginning of training.

[0138] Next, construct a lightweight inference model. Use the initialized policy model as the teacher model, and extract its policy distribution output and value function prediction as soft label information. For example, in the robot navigation task, the policy distribution output can be represented as the probability distribution of selecting different actions (forward, backward, left turn, right turn) at each time step, and the value function prediction can be represented as the expected cumulative reward of the current state. Construct a student model with fewer layers and a parameter sharing mechanism. For example, the student model can use smaller convolutional kernels or fewer convolutional layers, and share some convolutional kernels between different layers to reduce the number of parameters.

[0139] After that, conduct knowledge distillation training. Based on the soft label information and the output of the student model, construct a knowledge distillation loss function. This loss function consists of three terms: the difference in policy distributions, the difference in value function predictions, and the task objective function (e.g., the reward for reaching the target point in the navigation task), and is obtained by linearly combining with weight coefficients. For example, weights of 0.5, 0.3, and 0.2 can be assigned to the difference in policy distributions, the difference in value function predictions, and the task objective function respectively. Use this loss function to optimize the parameters of the student model to obtain the lightweight inference model.

[0140] Then, deploy the lightweight inference model and start the performance evaluation. During the performance evaluation, calculate the matching degree between the output actions of the lightweight inference model and the optimal actions to obtain the policy accuracy rate. For example, in 100 test scenarios, it is the proportion of the number of times the actions selected by the model are consistent with the optimal actions. At the same time, calculate the deviation between the predicted value function and the true value function to obtain the value function error. For example, it is the average difference between the predicted value and the true value. Calculate the comprehensive performance index according to the preset weights (for example, 0.7 and 0.3) of the policy accuracy rate and the value function error.

[0141] Finally, monitor the optimization improvement amplitude and save the tuning snapshot. Calculate the optimization improvement amplitude based on the comprehensive performance indexes of the current round and the previous round. For example, (Comprehensive performance index of the current round - Comprehensive performance index of the previous round) / Comprehensive performance index of the previous round. Set a sliding time window (for example, 5 rounds). When the optimization improvement amplitudes of multiple consecutive rounds within the sliding time window are all lower than the preset improvement threshold (for example, 0.01), save the network structure parameters, optimizer state, and performance indexes of the lightweight inference model as the tuning snapshot, stop the optimization process, and output the tuning snapshot.

[0142] The solution of this application can:

[0143] Reduce the model complexity: Through knowledge distillation, compress the large-scale policy model into a lightweight inference model, significantly reducing the number of model parameters and the amount of computation, enabling the model to be deployed and run on resource-constrained devices. Improve the inference speed: The lightweight model has a faster inference speed and can meet the requirements of real-time applications. For example, in fields such as robot control and autonomous driving, fast response is crucial. Maintain the model performance: The knowledge distillation process uses the soft label information of the teacher model to guide the learning of the student model, enabling the student model to inherit the knowledge and experience of the teacher model, so as to maintain a high performance level while reducing the model size. For example, maintain a high navigation success rate and a low path length in the navigation task.

[0144] In an optional implementation manner, perform priority sorting on the parameter sampling set according to the parameter exploration priority to obtain a parameter tuning sequence, and calculate the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence; construct a confidence boundary adjustment strategy based on the estimated mean and the estimated standard deviation, and determine the parameter optimization direction by adaptively adjusting the learning rate and the confidence coefficient to obtain an optimized parameter solution, including:

[0145] Perform priority sorting on the parameter sampling set according to the parameter exploration priority to obtain a parameter tuning sequence, calculate the correlation degree between adjacent parameters in the parameter tuning sequence, and combine the parameters with a correlation degree greater than a preset correlation threshold into a parameter group;

[0146] Establish a historical data sampling window for the parameter tuning sequence, apply time-decaying weights to the data within the historical data sampling window to obtain a weighted historical sample set; in the weighted historical sample set, calculate the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence, and at the same time calculate the joint distribution characteristics of the parameters within the parameter group, where the estimated mean and the estimated standard deviation reflect the distribution characteristics of the parameters;

[0147] Based on the estimated mean and the estimated standard deviation, construct a confidence boundary adjustment strategy, and determine the upper and lower boundaries of the parameter confidence interval in combination with the joint distribution characteristics. The confidence boundary adjustment strategy includes a single-parameter confidence interval and a parameter-group confidence interval; calculate the deviation degree between the current parameter value and the parameter confidence interval, and at the same time calculate the combined deviation between the current value of the parameter group and the parameter-group confidence interval, and determine the parameter adjustment reference direction based on the deviation degree and the combined deviation;

[0148] Analyze the performance change trend in the parameter historical optimization process, calculate the performance improvement sequence of the recent several optimizations, adaptively adjust the learning rate according to the fluctuation characteristics of the performance improvement sequence, and use the ratio of the learning rate to the performance improvement amplitude as the confidence coefficient; extract several optimization directions with the most significant performance improvement in the parameter historical optimization path, perform a weighted combination of the several optimization directions and the parameter adjustment reference direction, correct the combined direction according to the confidence coefficient to obtain the parameter optimization direction, and multiply the parameter optimization direction by the learning rate to obtain the optimization parameter scheme.

[0149] First, determine the priority of parameter exploration. For example, sort according to the influence degree of the parameter on the model performance or the sensitivity of the parameter. Suppose there are five parameters A, B, C, D, and E, and their priority is judged to be A > B > C > D > E according to experience, thus obtaining the parameter tuning sequence {A, B, C, D, E}.

[0150] Next, calculate the correlation degree between adjacent parameters in the parameter tuning sequence. For example, the correlation degree between parameters A and B can be measured by calculating the Pearson correlation coefficient between their historical adjustment data. Suppose the correlation degree between A and B is 0.8, the correlation degree between B and C is 0.2, the correlation degree between C and D is 0.9, and the correlation degree between D and E is 0.1. Set the preset correlation threshold to 0.7, then combine the parameters with a correlation degree greater than 0.7 into parameter groups to obtain parameter groups {A, B} and {C, D}.

[0151] Then, establish a historical data sampling window for the parameter tuning sequence. For example, the most recent 100 model training data can be taken as the historical data sampling window. Apply time-decaying weights to the data within the window, such as exponential decay, so that the most recent data has a higher weight. Suppose the parameter A has taken the values 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 in the most recent 10 times. After time-decaying weighting, a weighted historical sample set is obtained.

[0152] In the weighted historical sample set, calculate the estimated mean and estimated standard deviation of parameter A. For example, suppose the weighted sample set is {0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5}, then the estimated mean is 3 and the estimated standard deviation is 1.58. At the same time, calculate the joint distribution characteristics of the parameters within the parameter group {A, B}. For example, the joint probability distribution of A and B can be calculated. Suppose the joint probability distribution of A and B is P(A=x, B=y), which can be estimated by counting the number of value combinations of A and B in the historical data.

[0153] Construct a confidence boundary adjustment strategy based on the estimated mean and estimated standard deviation. For parameter A, assuming a confidence level of 95%, the upper and lower boundaries of the confidence interval can be calculated according to the normal distribution as the mean plus or minus twice the standard deviation, that is, 3 ± 2 * 1.58 = [-0.16, 6.16]. For the parameter group {A, B}, the upper and lower boundaries of the confidence interval can be determined according to the joint distribution characteristics. Suppose the joint confidence interval of A and B is [A_lower, A_upper] x [B_lower, B_upper].

[0154] Calculate the deviation degree between the current parameter value and the confidence interval. Suppose the current value of parameter A is 4, then its deviation degree from the confidence interval is |4 - 3| / (6.16 - (-0.16)) = 0.16. At the same time, calculate the combined deviation of the current values of the parameter group {A, B} from the confidence interval. Suppose the current value of parameter B is 2 and the confidence interval of the parameter group {A, B} is [2, 6] x [1, 3], then the combined deviation can be defined as the weighted sum of the deviations of A and B from their respective confidence intervals.

[0155] Analyze the performance change trend during the historical optimization of parameters. For example, calculate the performance improvement sequence of the last 10 optimizations. Assume the performance improvement sequence is {0.1, 0.2, 0.3, 0.2, 0.1, 0, -0.1, 0.1, 0.2, 0.3}. Adaptively adjust the learning rate according to the fluctuation characteristics of the performance improvement sequence. For example, if the performance improvement sequence fluctuates greatly, decrease the learning rate; if the performance improvement sequence is relatively stable, increase the learning rate. Assume the current learning rate is 0.1 and the performance improvement amplitude is 0.3, then the confidence coefficient is 0.1 / 0.3 = 0.33.

[0156] Extract several optimization directions with the most significant performance improvement from the historical optimization path of parameters. For example, the top three optimization directions with the largest performance improvement can be extracted. Combine these three optimization directions with the parameter adjustment reference direction through weighted combination. Correct the combined direction according to the confidence coefficient to obtain the parameter optimization direction. Multiply the parameter optimization direction by the learning rate to obtain the optimized parameter scheme. For example, assume the optimization direction is [0.1, 0.2, 0.3, 0.4, 0.5] and the learning rate is 0.1, then the optimized parameter scheme is [0.01, 0.02, 0.03, 0.04, 0.05].

[0157] The solution of this application can:

[0158] Improve the parameter optimization efficiency: Through priority sorting and correlation analysis, parameters and parameter groups that have a greater impact on the model performance can be concentrated for optimization, thereby improving the parameter optimization efficiency. Enhance the stability of parameter optimization: The confidence boundary adjustment strategy can effectively avoid excessive parameter adjustment and improve the stability of parameter optimization. Achieve adaptive parameter optimization: The adaptive learning rate and confidence coefficient can dynamically adjust the parameter optimization strategy according to the model performance change trend, achieving adaptive parameter optimization.

[0159] Figure 2 This is a schematic structural diagram of the database adaptive data stream acquisition optimization system based on reinforcement learning according to an embodiment of the present invention, as Figure 2 shown, the system includes:

[0160] The first unit is used to construct a distributed data collection agent network, obtain the data stream information of the database in real time, and generate an initial feature set; perform multimodal decomposition on the initial feature set, extract the data stream fluctuation spectrum feature value, fine-grained type distribution feature value, multi-scale time series feature value, and graph structure association feature value to form a multimodal feature matrix; based on the multimodal feature matrix, calculate the feature weight coefficient through an improved attention mechanism, adaptively fuse different modal features, and output a fused feature vector; use the fused feature vector, combine with a genetic algorithm to calculate the optimal initial collection parameter combination, including the concurrent collection threshold, cache capacity threshold, and batch processing size threshold, and generate an initial state space vector;

[0161] The second unit is used to construct a two-layer deep Q-learning network structure based on the initial state space vector and the collection parameter combination, set the input of the high-level network as the fused feature vector, and output a global collection strategy; set the input of the low-level network as the initial state space vector, and output a specific parameter adjustment plan; design a hierarchical reward function, use the feature weight coefficient as the balance factor of the reward function, and establish a multi-objective evaluation system including collection delay rate, data consistency, throughput efficiency, and resource utilization rate; perform a hierarchical training process based on the multi-objective evaluation system, use the global collection strategy output by the high-level network as the constraint condition of the low-level network, and iteratively update the network parameters through a priority experience replay mechanism to generate a hierarchical collection strategy model;

[0162] The third unit is used to deploy the hierarchical collection strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fused feature vector, and evaluate the feature drift degree; when the feature drift degree exceeds the preset feature threshold, trigger the high-level network to re-plan the collection strategy based on the current fused feature vector, and input the updated global strategy into the low-level network for parameter fine-tuning; monitor the collection performance indicators in real time, when the performance indicators are lower than the preset target threshold, start a progressive tuning mechanism, input the newly collected state transition sequence into the training cache pool, and perform incremental learning based on historical tuning experience to realize the continuous optimization of the hierarchical collection strategy model.

[0163] In the third aspect of the embodiments of the present invention,

[0164] There is provided an electronic device, including:

[0165] A processor;

[0166] A memory for storing instructions executable by the processor;

[0167] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0168] In a fourth aspect of the embodiments of the present invention,

[0169] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0170] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A database adaptive data stream acquisition optimization method based on reinforcement learning, characterized in that: include: Build a distributed data collection agent network to obtain data flow information from the database in real time and generate an initial feature set; Performing multimodal decomposition on the initial feature set, extracting data stream fluctuation spectrum eigenvalues, fine-grained type distribution eigenvalues, multi-scale time series eigenvalues ​​and graph structure association eigenvalues, and forming a multimodal feature matrix; Based on the multimodal feature matrix, feature weight coefficients are calculated through an improved attention mechanism, different modal features are adaptively fused, and a fused feature vector is output; Using the fused feature vector and combining it with a genetic algorithm, an optimal initial acquisition parameter combination is calculated, including a concurrent acquisition threshold, a cache capacity threshold, and a batch size threshold, to generate an initial state space vector; Based on the initial state space vector and the acquisition parameter combination, a two-layer deep Q learning network structure is constructed, the input of the high-level network is set to the fused feature vector, and a global acquisition strategy is output; the input of the low-level network is set to the initial state space vector, and a specific parameter adjustment plan is output; a hierarchical reward function is designed, the feature weight coefficient is used as a balancing factor of the reward function, and a multi-objective evaluation system including acquisition delay rate, data consistency, throughput efficiency and resource utilization is established; a hierarchical training process is performed based on the multi-objective evaluation system, the global acquisition strategy output by the high-level network is used as a constraint condition of the low-level network, the network parameters are iteratively updated through a priority experience replay mechanism, and a hierarchical acquisition strategy model is generated; Deploy the hierarchical collection strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fused feature vector, and evaluate the degree of feature drift; when the degree of feature drift exceeds the preset feature threshold, trigger the high-level network to re-plan the collection strategy based on the current fused feature vector, and input the updated global strategy into the low-level network for parameter fine-tuning; The acquisition performance indicators are monitored in real time. When the performance indicators are lower than the preset target threshold, the progressive tuning mechanism is started, the newly acquired state transition sequence is input into the training cache pool, and incremental learning is performed based on historical tuning experience to achieve continuous optimization of the hierarchical acquisition strategy model.

2. The method according to claim 1, characterized in that Based on the multimodal feature matrix, feature weight coefficients are calculated through an improved attention mechanism, different modal features are adaptively fused, and a fused feature vector is output; The fused feature vector is used in combination with a genetic algorithm to calculate the optimal initial acquisition parameter combination, including a concurrent acquisition threshold, a cache capacity threshold, and a batch size threshold, to generate an initial state space vector including: The multimodal feature matrix includes fluctuation spectrum features reflecting dynamic changes in data streams, fine-grained type features reflecting data distribution laws, multi-scale sequence features reflecting temporal fluctuation laws, and graph association features reflecting topological structures. A data multi-dimensional representation feature matrix is ​​constructed based on the fluctuation spectrum features, the fine-grained type features, the multi-scale sequence features, and the graph association features. Calculate the statistical distribution characteristics of the multidimensional characterization feature matrix of the data, select the corresponding normalization strategy according to the statistical distribution characteristics, use the maximum and minimum value normalization for the fluctuation spectrum characteristics to maintain the peak and trough relationship, use softmax normalization for the fine-grained type characteristics to highlight the main types, use z-score normalization for the multi-scale sequence characteristics to eliminate the dimension effect, and use degree centrality normalization for the graph association characteristics to balance the importance of nodes, and generate a standardized feature matrix; Analyze the information entropy distribution of each modal feature in the standardized feature matrix, use the information entropy distribution as a modal importance indicator, construct a modal adaptive factor based on the information entropy, dynamically adjust the modal adaptive factor with the information entropy to adaptively capture high-information features, and integrate the modal adaptive factor into the improved attention mechanism to calculate the dynamic attention weight; A two-layer feature fusion network is constructed using the dynamic attention weights. In the first layer, a self-attention operation is performed within the modality to extract the internal correlation of the modality. In the second layer, a cross-attention operation between modalities is performed to model the complementary relationship of the modalities. The outputs of the two layers of attention are fused through a residual structure and a normalization layer to obtain a fused feature vector containing multi-modal complementary information. Component analysis is performed on the fused feature vector, and the concurrent acquisition threshold boundary is determined by using the period and peak characteristics of the fluctuation spectrum feature, the cache capacity threshold boundary is determined by using the entropy value and skewness characteristics of the fine-grained type feature, and the batch processing threshold boundary is determined by using the autocorrelation and stationary characteristics of the multi-scale sequence feature, and the three threshold boundaries constitute a parameter search space; Extracting feature importance based on the principal component analysis result of the fused feature vector, mapping the feature importance into a weight coefficient, constructing a multi-objective fitness function including system throughput, processing delay and resource utilization, and dynamically weighting the three performance indicators according to the weight coefficient; Initializing a genetic algorithm population in the parameter search space, using the multi-objective fitness function to evaluate the quality of individuals in the population, adaptively adjusting the crossover probability based on the population diversity level to balance local and global searches, adaptively adjusting the mutation probability based on the optimization convergence degree to prevent premature convergence, and obtaining the optimal parameter combination through iterative optimization; The optimal parameter combination and the fused feature vector are deeply embedded to generate a state space vector of coupling parameter information and feature representation.

3. The method according to claim 1, characterized in that Based on the initial state space vector and the acquisition parameter combination, a double-layer deep Q learning network structure is constructed, the input of the high-level network is set as the fused feature vector, and a global acquisition strategy is output; The input of the low-level network is set to the initial state space vector, and the specific parameter adjustment scheme is outputted including: Based on the initial state space vector and the acquisition parameter combination, a double-layer deep Q learning network structure is constructed, the fused feature vector is set as the input of the high-level network, and the initial state space vector is set as the input of the low-level network, and the high-level network and the low-level network adopt a double network structure in which the target network and the evaluation network are separated; A dual-stream feature extraction channel is constructed in the high-level network, wherein the first channel extracts static features in the fused feature vector through a multi-layer fully connected network to obtain a static feature representation, and the second channel extracts dynamic features in the fused feature vector through a bidirectional LSTM network to obtain a dynamic feature representation. The static feature representation and the dynamic feature representation are adaptively fused through a multi-head attention mechanism to output a global acquisition strategy, wherein the global acquisition strategy includes a parameter adjustment direction, an adjustment step, and an adjustment priority. The adjustment direction is used to indicate a parameter increase or decrease, the adjustment step is used to control a parameter change amplitude, and the adjustment priority is used to determine a multi-parameter adjustment order. In the low-level network, based on the initial state space vector, the Dueling architecture is used to decompose the Q value into a state value function and an advantage function. The state value function evaluates the parameter configuration performance through a three-layer feedforward neural network to obtain a performance evaluation value. The advantage function evaluates and adjusts the action advantage through a deep residual network to obtain an action advantage value. The performance evaluation value and the action advantage value are weightedly combined to output a specific parameter adjustment plan.

4. The method according to claim 1, characterized in that: Deploy the hierarchical collection strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fused feature vector, and evaluate the degree of feature drift; when the degree of feature drift exceeds the preset feature threshold, trigger the high-level network to re-plan the collection strategy based on the current fused feature vector, and input the updated global strategy into the low-level network for parameter fine-tuning, including: Deploy the hierarchical acquisition strategy model to the production environment, establish a feature similarity hybrid measurement model, continuously collect data stream features to obtain real-time feature vectors, calculate the directional deviation between the real-time feature vector and the fused feature vector on continuous features by cosine similarity to obtain continuous feature similarity, and calculate the distribution change amplitude between the real-time feature vector and the fused feature vector on discrete features by KL divergence to obtain discrete feature similarity; The continuous feature similarity and the discrete feature similarity are normalized and weighted to obtain a feature drift comprehensive index, and a multi-scale feature drift detection mechanism is constructed based on the feature drift comprehensive index; within a first time window, the feature drift comprehensive index is smoothed by an exponentially weighted moving average method to obtain a feature trend, within a second time window, a mutation point detection algorithm is used to identify the jump position of the feature drift comprehensive index to obtain a feature change point, and within a third time window, a trend analysis method is used to capture the gradual change process of the feature drift comprehensive index to obtain a feature trend; Input the feature trend, the feature change point and the feature trend into a decision fusion module, compare the feature drift index at each time scale with the corresponding preset feature threshold, and generate a feature drift degree assessment result; When the feature drift degree evaluation result indicates that the feature drift index at any time scale exceeds the corresponding preset feature threshold, recalculating the feature importance weight matrix based on the current fused feature vector, and inputting the feature importance weight matrix into the high-level network; The high-level network replans the acquisition strategy according to the feature importance weight matrix, updates the feature extraction scheme to obtain the feature weight update strategy, adjusts the parameter change direction and change step size based on the feature weight update strategy to obtain the global strategy update scheme, and inputs the global strategy update scheme into the low-level network for parameter fine-tuning.

5. The method according to claim 1, characterized in that Real-time monitoring of acquisition performance indicators. When the performance indicators are lower than the preset target threshold, the progressive tuning mechanism is started, the newly acquired state transition sequence is input into the training cache pool, and incremental learning is performed based on historical tuning experience to achieve continuous optimization of the hierarchical acquisition strategy model, including: Construct a multi-dimensional performance monitoring model, collect data throughput, data delay time, and resource utilization in real time to obtain efficiency index data, collect data integrity, data accuracy, and data timeliness to obtain quality index data, and collect computing overhead, storage overhead, and bandwidth overhead to obtain cost index data; calculate the respective index weights for the efficiency index data, the quality index data, and the cost index data, and perform a weighted combination of the index weights and the corresponding index data to obtain a comprehensive performance score; compare the comprehensive performance score with a preset target threshold, and trigger a parameter optimization process when the comprehensive performance score is lower than the preset target threshold; In the parameter optimization process, the parameter space is randomly sampled based on the Thompson sampling strategy to obtain a parameter sampling set, and the uncertainty of each parameter interval in the parameter sampling set is calculated to obtain the parameter exploration priority; the parameter sampling set is prioritized according to the parameter exploration priority to obtain a parameter tuning sequence, and the estimated mean and estimated standard deviation of each parameter in the parameter tuning sequence are calculated; a confidence boundary adjustment strategy is constructed based on the estimated mean and the estimated standard deviation, and the parameter optimization direction is determined by adaptively adjusting the learning rate and the confidence coefficient to obtain an optimized parameter solution; the state transition data generated during the optimization process is collected, and the state transition data is stored in the first cache layer, the second cache layer, and the third cache layer respectively according to the data timestamp to construct a hierarchical training cache pool; Calculate the timeliness weight and value weight of the experience data in the hierarchical training cache pool, multiply the timeliness weight and the value weight to obtain the experience sampling weight; sample the training data set from the hierarchical training cache pool according to the experience sampling weight, analyze the common features in the training data set to obtain the transferable knowledge representation; use the transferable knowledge representation to initialize the policy model parameters to obtain the initial policy model, and compress the initial policy model into a lightweight reasoning model through the knowledge distillation method; deploy the lightweight reasoning model and start performance evaluation, calculate the performance difference before and after optimization to obtain the optimization improvement, and when the optimization improvement is lower than the preset improvement threshold for multiple consecutive rounds, save the current tuning state to obtain a tuning snapshot and stop the optimization process.

6. The method according to claim 5, characterized in that Prioritizing the parameter sampling set according to the parameter exploration priority to obtain a parameter tuning sequence, and calculating an estimated mean and an estimated standard deviation of each parameter in the parameter tuning sequence; Constructing a confidence boundary adjustment strategy based on the estimated mean and the estimated standard deviation, and determining the parameter optimization direction by adaptively adjusting the learning rate and the confidence coefficient to obtain an optimization parameter solution includes: Prioritizing the parameter sampling set according to the parameter exploration priority to obtain a parameter tuning sequence, calculating the correlation between adjacent parameters in the parameter tuning sequence, and combining parameters with correlations greater than a preset correlation threshold into a parameter group; Establishing a historical data sampling window of the parameter tuning sequence, applying a time decay weight to the data in the historical data sampling window, and obtaining a weighted historical sample set; in the weighted historical sample set, calculating an estimated mean and an estimated standard deviation of each parameter in the parameter tuning sequence, and calculating the joint distribution characteristics of the parameters in the parameter group, wherein the estimated mean and the estimated standard deviation reflect the distribution characteristics of the parameters; A confidence boundary adjustment strategy is constructed based on the estimated mean and the estimated standard deviation, and the upper and lower boundaries of the parameter confidence interval are determined in combination with the joint distribution characteristics, wherein the confidence boundary adjustment strategy includes a single parameter confidence interval and a parameter group confidence interval; the degree of deviation between the current parameter value and the parameter confidence interval is calculated, and the combined deviation between the current value of the parameter group and the parameter group confidence interval is calculated, and a reference direction for parameter adjustment is determined based on the degree of deviation and the combined deviation; Analyze the performance change trend in the parameter history optimization process, calculate the performance improvement sequence of the most recent optimizations, adaptively adjust the learning rate according to the fluctuation characteristics of the performance improvement sequence, and use the ratio of the learning rate to the performance improvement as the confidence coefficient; extract several optimization directions with the most significant performance improvement in the parameter history optimization path, perform weighted combination of the several optimization directions and the parameter adjustment reference direction, correct the combined direction according to the confidence coefficient to obtain the parameter optimization direction, and multiply the parameter optimization direction by the learning rate to obtain the optimization parameter solution.

7. A database adaptive data stream acquisition optimization system based on reinforcement learning, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to build a distributed data collection agent network, obtain the data flow information of the database in real time, and generate an initial feature set; Performing multimodal decomposition on the initial feature set, extracting data stream fluctuation spectrum eigenvalues, fine-grained type distribution eigenvalues, multi-scale time series eigenvalues ​​and graph structure association eigenvalues, and forming a multimodal feature matrix; Based on the multimodal feature matrix, feature weight coefficients are calculated through an improved attention mechanism, different modal features are adaptively fused, and a fused feature vector is output; Using the fused feature vector and combining it with a genetic algorithm, an optimal initial acquisition parameter combination is calculated, including a concurrent acquisition threshold, a cache capacity threshold, and a batch size threshold, to generate an initial state space vector; The second unit is used to construct a two-layer deep Q learning network structure based on the initial state space vector and the acquisition parameter combination, set the input of the high-level network to the fused feature vector, and output a global acquisition strategy; set the input of the low-level network to the initial state space vector, and output a specific parameter adjustment plan; design a hierarchical reward function, use the feature weight coefficient as a balancing factor of the reward function, and establish a multi-objective evaluation system including acquisition delay rate, data consistency, throughput efficiency and resource utilization; perform a hierarchical training process based on the multi-objective evaluation system, use the global acquisition strategy output by the high-level network as a constraint condition of the low-level network, iteratively update the network parameters through a priority experience replay mechanism, and generate a hierarchical acquisition strategy model; The third unit is used to deploy the hierarchical collection strategy model to the production environment, continuously collect data stream features, calculate the similarity between the real-time features and the fused feature vector, and evaluate the degree of feature drift; when the degree of feature drift exceeds a preset feature threshold, trigger the high-level network to re-plan the collection strategy based on the current fused feature vector, and input the updated global strategy into the low-level network for parameter fine-tuning; The acquisition performance indicators are monitored in real time. When the performance indicators are lower than the preset target threshold, the progressive tuning mechanism is started, the newly acquired state transition sequence is input into the training cache pool, and incremental learning is performed based on historical tuning experience to achieve continuous optimization of the hierarchical acquisition strategy model.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Distributed data analysis method based on mass data

    CN114201486A

  • Big data analysis-based plan scheduling optimization method

    CN117076077A