A method and system for dynamic adjustment of multi-source heterogeneous data resources

CN122711352APending Publication Date: 2026-09-08SUZHOU DIANJINGZHIBI DECORATION ENGINEERING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610616110.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0004]针对现有技术的不足,本发明提供了一种针对多源异构数据资源动态调整的方法及系统,解决了现有技术难以精准捕捉多源异构数据异构特征与资源需求的动态关联、导致资源供给与实际处理需求偏差的问题

Benefits of technology

本发明通过先初始化异构数据特征库与异构数据处理行为档案,建立统一的异构数据特征描述体系并关联历史资源消耗数据,再精准获取多源异构数据的实时状态信息,基于该信息分析数据异构特征以识别计算、存储、传输资源需求,随后利用机器学习算法构建异构特征与资源需求的动态关联关系模型并按周期优化,结合当前数据处理平台资源总量与实时负载,通过多目标优化算法生成资源优化分配方案,最后执行资源动态调整并持续监测反馈以迭代优化模型与方案,有效解决了现有技术难以精准捕捉多源异构数据异构特征与资源需求动态关联、依赖固定规则或单一指标导致资源供给与实际需求存在偏差的问题,既提升了数据处理的实时性与准确性,又减少了资源浪费,同时增强了系统稳定性,确保技术方案可被本领域技术人员实现且能适配多源异构数据的动态变化,适用于各类多源异构数据处理场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122711352A_ABST
    Figure CN122711352A_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods and systems for the dynamic adjustment of multi-source heterogeneous data resources, it is related to data processing and resource scheduling technical field.The application includes: S1, the real-time state information of multi-source heterogeneous data is acquired, real-time state information includes the structural difference of multi-source heterogeneous data, data format, data update frequency and data processing dependency relationship;S2, based on real-time state information, the heterogeneous characteristics of multi-source heterogeneous data are analyzed.The application is initialized by first heterogeneous data feature library and heterogeneous data processing behavior file, establishes unified heterogeneous data feature description system and is associated with historical resource consumption data, then the real-time state information of multi-source heterogeneous data is accurately acquired, the data heterogeneous characteristics are analyzed based on the information to identify calculation, storage, transmission resource demand, then the dynamic correlation relationship model of heterogeneous characteristics and resource demand is constructed using machine learning algorithm and is optimized according to period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing and resource scheduling technology, specifically to a method and system for dynamically adjusting multi-source heterogeneous data resources. Background Technology

[0002] In the field of data processing, multi-source heterogeneous data refers to data sets originating from various acquisition devices, systems, databases, and other channels, with differences in data format, structure, and storage methods. This type of data is widely present in various information systems, and its efficient utilization requires dynamic adjustment of relevant resources based on the real-time status of the data and processing needs. By rationally allocating computing, storage, and transmission resources, the smooth and efficient data processing process can be ensured, which is also one of the important research directions in the current field of data management.

[0003] Existing dynamic data resource adjustment technologies struggle to accurately capture the dynamic relationship between heterogeneous data characteristics and resource requirements when dealing with multi-source, heterogeneous data. This results in adjustment strategies lacking deep adaptation to changes in the data's inherent attributes. These technologies often trigger adjustments based on fixed resource allocation rules or single-dimensional monitoring indicators, failing to optimize resource allocation schemes in real time according to heterogeneous attributes such as structural differences, update frequency variations, and processing dependencies among different data types. Consequently, a discrepancy exists between resource supply and actual data processing needs, impacting both the real-time performance and accuracy of data processing and leading to resource waste. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for dynamically adjusting multi-source heterogeneous data resources, solving the problem that existing technologies struggle to accurately capture the dynamic correlation between the heterogeneous characteristics of multi-source heterogeneous data and resource demands, leading to discrepancies between resource supply and actual processing needs.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for dynamically adjusting multi-source heterogeneous data resources, comprising: S1. Obtain real-time status information of multi-source heterogeneous data. The real-time status information includes the structural differences, data format, data update frequency, and data processing dependencies of multi-source heterogeneous data. S2. Based on real-time status information, analyze the heterogeneous characteristics of multi-source heterogeneous data, and identify the computing resource requirements, storage resource requirements, and transmission resource requirements of data processing tasks according to the heterogeneous characteristics. S3. Construct a dynamic correlation model between the heterogeneous characteristics of multi-source heterogeneous data and the requirements for computing resources, storage resources, and transmission resources. S4. Based on the dynamic correlation model and the real-time status information of multi-source heterogeneous data, generate a resource optimization allocation scheme. The scheme is used to adjust computing resources, storage resources and transmission resources. S5. Based on the resource optimization allocation plan, dynamically adjust computing resources, storage resources, and transmission resources.

[0006] Preferably, before acquiring the real-time status information of the multi-source heterogeneous data, the method further includes: Initialize the heterogeneous data feature library, and based on the preset data acquisition strategy, acquire metadata, structured information, semi-structured information and unstructured information of multi-source heterogeneous data from multiple data sources according to the preset period, and perform standardized processing on the information to establish a unified heterogeneous data feature description system. Continuously record historical resource consumption data of multi-source heterogeneous data under different processing tasks. The historical resource consumption data includes computing resource utilization, storage space usage, network transmission bandwidth usage, and task processing completion time. The historical resource consumption data is associated with the corresponding data heterogeneity characteristics to construct a heterogeneous data processing behavior profile.

[0007] Preferably, step S1 includes: S11. When the structure of multi-source heterogeneous data changes, the structure of the newly added or modified data fields is parsed and their structural difference descriptions are updated. S12. When the update frequency of multi-source heterogeneous data exceeds the preset threshold, the data flow change rate and data volume growth trend are monitored in real time, and the update frequency information is updated. S13. When multi-source heterogeneous data is accessed or processed by a specific processing task, analyze its input-output relationship with other data, identify and update its processing dependencies.

[0008] Preferably, step S2 includes: S21. Based on the structural differences and data formats of multi-source heterogeneous data, analyze the computational resource overhead required for data cleaning, transformation and integration; S22. Based on the update frequency and data volume of multi-source heterogeneous data, predict the growth of data storage capacity and the storage space requirements for data caching; S23. Based on the processing dependencies of multi-source heterogeneous data, determine the network bandwidth requirements and data transmission latency requirements for data transmission between different processing nodes.

[0009] Preferably, step S3 includes: S31. Obtain historical heterogeneous data characteristics and corresponding historical resource consumption data from the heterogeneous data processing behavior archive; S32. Using machine learning algorithms, historical heterogeneous data features are taken as input and corresponding historical resource consumption data are taken as output to train a resource demand prediction model. The prediction model is used to map the resource demand of multi-source heterogeneous data under different heterogeneous features. S33. Based on the resource demand prediction model, make real-time predictions on newly incoming heterogeneous data, and update and optimize the prediction model according to a preset cycle using the latest collected heterogeneous data characteristics and resource consumption data.

[0010] Preferably, step S32 includes: S321. Represent the characteristics of historical heterogeneous data as feature vectors. The feature vectors include data structure complexity, data volume, update rate, access frequency, processing chain length, and data sensitivity. S322. Represent historical resource consumption data as a resource consumption vector. The resource consumption vector includes computing core usage, memory usage, storage throughput, network inbound traffic, network outbound traffic, and average task response time. S323. Based on feature vectors and resource consumption vectors, establish a nonlinear mapping relationship between heterogeneous features of multi-source heterogeneous data and resource demand through regression analysis or neural network models.

[0011] Preferably, step S4 includes: S41. Based on the predicted resource requirements output by the dynamic correlation model, combined with the total amount and real-time load of the available computing resources, storage resources and transmission resources of the current data processing platform. S42. Employ a multi-objective optimization algorithm to comprehensively consider the real-time requirements of data processing, resource utilization, processing costs, and system stability constraints, and generate detailed resource allocation decisions. S43. The resource allocation decision clearly stipulates the amount of computing resources, storage resources, and transmission resources that each data processing task should receive within a specified time.

[0012] Preferably, step S5 includes: S51. Send the resource adjustment instruction contained in the resource optimization allocation plan to the resource scheduling service. The resource adjustment instruction includes the resource type, the adjustment target value, and the effective time. S52, The resource scheduling service performs elastic scaling of computing resources, capacity expansion or data migration of storage resources, and bandwidth QoS adjustment or routing optimization of transmission resources. S53. Continuously monitor the processing performance and actual resource utilization of the adjusted multi-source heterogeneous data, and use it as feedback information for iterative optimization of the dynamic correlation model and adjustment of the resource optimization allocation scheme.

[0013] Preferably, the continuous monitoring of the processing performance and actual resource utilization of the adjusted multi-source heterogeneous data includes: Real-time acquisition of adjusted data processing task completion time, data processing throughput, error rate, and resource performance metrics including CPU utilization, memory utilization, disk IOPS, and network latency; Compare and analyze the performance indicators with the expected effects of the resource optimization allocation plan to assess the suitability of the current adjustment; When a deviation is found between actual performance and expected results, the dynamic correlation model is retrained or the resource optimization allocation scheme is partially modified.

[0014] This invention also provides a system for dynamically adjusting multi-source heterogeneous data resources, comprising: The data status acquisition module is used to acquire real-time status information of multi-source heterogeneous data. The real-time status information includes the structural differences, data format, data update frequency, and data processing dependencies of the multi-source heterogeneous data. The heterogeneous feature analysis module is used to analyze the heterogeneous features of multi-source heterogeneous data based on real-time status information, and to identify the computing resource requirements, storage resource requirements and transmission resource requirements of data processing tasks based on the heterogeneous features. The dynamic correlation modeling module is used to construct a dynamic correlation model between the heterogeneous characteristics of multi-source heterogeneous data and the requirements for computing resources, storage resources, and transmission resources. The scheme generation module is used to generate resource optimization allocation schemes based on the dynamic correlation model and the real-time status information of multi-source heterogeneous data. The schemes are used to adjust computing resources, storage resources and transmission resources. The resource adjustment execution module is used to dynamically adjust computing resources, storage resources, and transmission resources according to the resource optimization allocation plan.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention first initializes a heterogeneous data feature library and a heterogeneous data processing behavior archive, establishing a unified heterogeneous data feature description system and associating it with historical resource consumption data. Then, it accurately acquires real-time status information of multi-source heterogeneous data. Based on this information, it analyzes the heterogeneous characteristics of the data to identify computing, storage, and transmission resource requirements. Subsequently, it uses machine learning algorithms to construct a dynamic correlation model between heterogeneous characteristics and resource requirements and optimizes it periodically. Combining the current total resources and real-time load of the data processing platform, it generates a resource optimization allocation scheme through a multi-objective optimization algorithm. Finally, it performs dynamic resource adjustments and continuously monitors feedback to iteratively optimize the model and scheme. This effectively solves the problems of existing technologies, such as difficulty in accurately capturing the dynamic correlation between heterogeneous characteristics and resource requirements of multi-source heterogeneous data, and the deviation between resource supply and actual demand caused by relying on fixed rules or single indicators. It improves the real-time performance and accuracy of data processing, reduces resource waste, and enhances system stability. This ensures that the technical solution can be implemented by those skilled in the art and adapts to the dynamic changes of multi-source heterogeneous data, making it suitable for various multi-source heterogeneous data processing scenarios. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 A flowchart illustrating the dynamic relationship modeling process of this invention; Figure 3 This is a flowchart illustrating the resource optimization and adjustment process of the present invention. Figure 4 This is a system structure diagram of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1

[0019] Please see Figure 1-3 This embodiment provides a method for dynamically adjusting multi-source heterogeneous data resources, and details it in conjunction with the actual application scenario of a big data processing platform. The specific implementation includes the following steps: The platform needs to integrate heterogeneous data from multiple sources, including e-commerce transaction data, logistics trajectory data, user behavior data, and third-party public opinion data. The data sources cover various channels such as database storage, API interface push, file upload, and real-time streaming. The data formats include structured relational data, semi-structured JSON data, and unstructured text and image data. The update frequency of different data ranges from minutes to days, and there are complex dependencies in the data processing process. For example, user payment completion data needs to trigger the generation and processing of logistics scheduling data.

[0020] First, preliminary preparations are conducted, including initializing a heterogeneous data feature library and collecting and processing data based on a pre-defined data acquisition strategy. The data acquisition strategy is tailored to the characteristics of different data sources. For structured data stored in a database, incremental collection is performed using JDBC connections at a preset 5-minute interval. For semi-structured data pushed via API interfaces, HTTP requests are used to monitor interface responses in real time and retrieve data. For unstructured data uploaded from files, changes to the file storage directory are monitored and collection is triggered accordingly. For real-time streaming data, a streaming acquisition framework is used to continuously receive data. During the collection process, metadata, structured information, semi-structured information, and unstructured information of the multi-source heterogeneous data are acquired. Metadata includes key information such as data source identifier, data creation time, business module to which the data belongs, and data size. Structured information covers data table field names, data types, field lengths, primary key constraints, etc. Semi-structured information includes the key-value pair structure and nesting levels of JSON data. Unstructured information records file format, encoding method, and content theme. This information was then standardized, and a unified heterogeneous data feature description system was established using the JSON-LD format. This mapped the feature information of different types of data to description fields in a unified format, ensuring consistency in subsequent feature analysis. Simultaneously, historical resource consumption data for multi-source heterogeneous data under different processing tasks was continuously recorded. This data included computational resource utilization, storage space usage, network bandwidth usage, and task completion time. Computational resource utilization encompassed metrics such as CPU utilization, memory utilization, and GPU utilization. Storage space usage included the original data storage capacity, compressed storage capacity, and cache usage. Network bandwidth usage recorded uplink and downlink bandwidth usage during data transmission. Task completion time was the total time elapsed from data access to processing completion. This historical resource consumption data was then associated with and stored along with the corresponding heterogeneous data features, constructing a heterogeneous data processing behavior archive. This archive, indexed by data feature identifiers, linked and stored detailed resource consumption information for each processing task, providing data support for subsequent model training.

[0021] After preliminary preparations are completed, the first step is to acquire real-time status information of multi-source heterogeneous data. This real-time status information includes the structural differences, data formats, data update frequencies, and data processing dependencies of the multi-source heterogeneous data. During data processing, when the structure of the multi-source heterogeneous data changes, the structural change is detected by hash value comparison. This triggers structural parsing of newly added or modified data fields, using a parser to parse the field's attribute information, including field name, data type, and value range, and updating its structural difference description to ensure that the real-time status information accurately reflects the latest changes in the data structure. When the update frequency of the multi-source heterogeneous data exceeds a preset threshold, its data flow change rate and data volume growth trend are monitored in real time. The preset threshold is set using a statistical method, based on the average update frequency of the data over the past 30 days, with 1.5 times the average update frequency as the preset threshold. If the current update frequency is higher than this threshold for three consecutive collection cycles, it is determined to exceed the threshold. At this time, the data flow transmission rate is collected in real time using a traffic statistics tool, and the data volume growth trend is analyzed in conjunction with a time window, and its update frequency information is updated. When multi-source heterogeneous data is accessed or processed by a specific processing task, the input-output relationship of the data in the data stream log is analyzed to identify the dependency relationship between the data and other data. For example, order payment data is used as input to trigger the processing of logistics data, and its processing dependency relationship is updated at the same time to ensure that the real-time status information can fully present the correlation logic in the data processing process.

[0022] When the system is initially running and the amount of heterogeneous data processing behavior archive data is insufficient, a rule-based initial resource allocation strategy can be adopted, or a pre-trained model with similar data scenarios can be used for transition. After accumulating sufficient historical data, the machine learning model described in this solution can be activated and continuously optimized.

[0023] After acquiring real-time status information, the heterogeneous characteristics of multi-source heterogeneous data are analyzed based on this information. The computational, storage, and transmission resource requirements for data processing tasks are then identified based on these characteristics. During the identification of computational resource requirements, the computational resource overhead required for data cleaning, transformation, and integration is analyzed based on the structural differences and data formats of the multi-source heterogeneous data. Data structural differences are measured using a structural complexity coefficient, which is calculated by weighting the sum of indicators such as the number of fields and nesting levels. The weighting coefficients are determined using the Analytic Hierarchy Process (AHP), with a weight of 0.6 for the number of fields and 0.4 for the nesting level. A higher structural complexity coefficient results in greater computational overhead for field matching and format correction during data cleaning. Data formats are categorized according to compatibility levels: high, medium, and low. Different levels correspond to different computational volumes for format conversion. The computational resource overhead is assessed by weighting the structural complexity coefficient and the computational volume corresponding to the format compatibility level. The weighting coefficients are set at 0.7 for the structural complexity coefficient and 0.3 for the format compatibility level, thus identifying the computational resource requirements. In terms of storage resource demand identification, based on the update frequency and data volume of multi-source heterogeneous data, the growth in data storage capacity and the storage space requirements for data caching are predicted. The ARIMA model from time series analysis is used for storage capacity growth prediction; the model formula is as follows: ; in, Let be the storage capacity requirement at time t. For constant terms, These are the autoregressive coefficients. Let be the actual storage capacity at time ti. The moving average coefficient is... Let be the error term at time tj. This represents the random error at the current moment. Data caching requirements are determined based on data access frequency, calculated as the number of accesses per unit time. Higher access frequency requires more cache space. Storage resource requirements are comprehensively identified by combining storage capacity growth predictions with caching needs. During the identification of transmission resource requirements, the network bandwidth and data transmission latency requirements for data transmission between different processing nodes are determined based on the processing dependencies of multi-source heterogeneous data. Network bandwidth requirements are calculated as the ratio of data transmission volume to transmission time, using the formula: ; in, For the required network bandwidth, To transmit data volume, For the allowed transmission time, A bandwidth reservation factor, ranging from 0.1 to 0.3, is used, adjusted based on network stability. Data transmission latency requirements are determined according to the real-time level of the processing task, which is divided into three levels: urgent, normal, and low. Each level corresponds to a different maximum allowable latency: the maximum allowable latency for the urgent level is no more than 100ms, for the normal level no more than 500ms, and for the low level no more than 1s. Combining network bandwidth requirements with transmission latency requirements, the identification of transmission resource needs is completed. This process accurately matches the actual resource needs of data processing, avoiding the problem of resource supply and demand mismatch.

[0024] After identifying heterogeneous features and resource requirements, a dynamic correlation model is constructed between the heterogeneous features of multi-source heterogeneous data and the computational, storage, and transmission resource requirements. First, historical heterogeneous data features and corresponding historical resource consumption data are obtained from the heterogeneous data processing behavior archive. Historical data from the past six months is selected as the training dataset to ensure the timeliness and representativeness of the data. The historical heterogeneous data features are represented as feature vectors, which include dimensions such as data structure complexity, data volume, update rate, access frequency, processing chain length, and data sensitivity. Data structure complexity is quantified using the aforementioned structural complexity coefficient; data volume is quantified in bytes; update rate is quantified in the number of updates per unit time; access frequency is quantified in the number of accesses per unit time; processing chain length is quantified in the number of processing steps the data participates in; and data sensitivity is assigned a quantization value between 0 and 1 based on the data security level. Historical resource consumption data is represented as a resource consumption vector, which includes dimensions such as computing core usage, memory usage, storage throughput, network inbound traffic, network outbound traffic, and average task response time. Computing core usage is quantified in terms of CPU cores, memory usage in terms of bytes, storage throughput in terms of bytes per second, network inbound and outbound traffic in terms of bytes per second, and average task response time in terms of milliseconds. Based on the feature vector and resource consumption vector, a BP neural network model is used to establish a nonlinear mapping relationship between heterogeneous features of multi-source heterogeneous data and resource requirements. The number of nodes in the input layer of the BP neural network is 6, the number of nodes is the same as the dimension of the feature vector, and there are 3 hidden layers with 12, 8, and 6 nodes respectively. The number of nodes in the output layer is 6, the number of nodes is the same as the dimension of the resource consumption vector. The ReLU function is used as the activation function, and the mean squared error function is used as the loss function. The formula is: ; in, For the sample size, This represents the actual resource consumption value. The model predicts resource consumption values. During model training, gradient descent is used to optimize parameters. The initial learning rate is set to 0.01, gradually decreasing with each training iteration. The number of iterations is set to 1000, and training stops when the loss function value falls below a preset threshold of 0.001, thus forming a resource demand prediction model. In practical applications, the resource demand prediction model is used to predict newly incoming heterogeneous data in real time. Simultaneously, the prediction model is updated and optimized using the latest collected heterogeneous data features and resource consumption data at a preset 24-hour cycle, ensuring that the model can continuously adapt to the dynamic changes in data features and resource demands. Compared to traditional fixed-rule resource allocation methods, this model can more accurately capture the correlation between the two.

[0025] Preferably, the feature vector and resource consumption vector are normalized before model training.

[0026] After the dynamic correlation model is constructed, a resource optimization allocation scheme is generated based on the model and the real-time status information of multi-source heterogeneous data. First, based on the predicted resource demand output by the dynamic correlation model, and combined with the total available computing, storage, and transmission resources and their real-time load on the current data processing platform, data such as the platform's total CPU cores, remaining cores, total memory capacity, remaining memory capacity, total storage capacity, remaining storage capacity, total network bandwidth, and used bandwidth are collected in real-time using resource monitoring tools to comprehensively understand the resource supply status. Then, the NSGA-Ⅲ multi-objective optimization algorithm is used to comprehensively consider the real-time requirements of data processing, resource utilization, processing costs, and system stability constraints to generate detailed resource allocation decisions. The objective function of the multi-objective optimization is defined as follows: ; in, For resource utilization, To address the processing cost, it is quantified as the weighted sum of the unit-time costs of the computing, storage, and transmission resources used to execute the data processing task. The weight of each resource's unit-time cost is consistent with the proportion of that resource consumed in the task processing. The system stability coefficient is quantified using a negative correlation function between task failure rate and resource contention count. The task failure rate and resource contention count are normalized for extreme values ​​before being substituted into the calculation. The calculation formula is as follows: ,in This represents the normalized task failure rate per unit time. This represents the normalized number of resource contention attempts per unit of time. , This is a weighting coefficient, set according to the resource priority requirements of the business scenario. The value range is [0,1], and a higher value indicates a more stable system operation. This is the ratio of data processing latency to preset latency. , , , The weighting coefficients, determined using the analytic hierarchy process (AHP), are 0.3, 0.2, 0.2, and 0.3, respectively. These values ​​are used in the actual calculation of the objective function. Before, it is necessary to , , , Various indicators undergo extreme value normalization to uniformly map them to the interval [0,1]. The formula for extreme value normalization is as follows: ,in The original value of the indicator. This is the minimum value of the indicator in the historical sample data. This represents the maximum value of the indicator in the historical sample data. The values ​​are the normalized values ​​for the indicators. Constraints include: the allocated computing resources do not exceed the total available computing resources, the allocated storage resources do not exceed the total available storage resources, the allocated transmission resources do not exceed the total available transmission resources, and the data processing latency does not exceed the maximum allowable latency for the corresponding task. This optimization algorithm obtains the Pareto optimal solution for resource allocation. The resource allocation decision clearly specifies the amount of computing resources, storage resources, and transmission resources that each data processing task should receive within a specified time. For example, an e-commerce transaction data processing task might be allocated 8 CPU cores, 16GB of memory, 100GB of storage capacity, and 100Mbps of network bandwidth for the next hour, ensuring the scientific and reasonable allocation of resources.

[0027] The above , , , The values ​​are determined based on the priority of the current business scenario using decision-making methods such as the Analytic Hierarchy Process (AHP); these weights can be adjusted according to actual needs in different application scenarios.

[0028] After the resource optimization allocation plan is generated, computing, storage, and transmission resources are dynamically adjusted according to the plan. First, the resource adjustment instructions contained in the resource optimization allocation plan are sent to the resource scheduling service. These instructions include the resource type, target adjustment value, and effective time. The resource scheduling service, implemented using the Kubernetes scheduler, can receive and parse these instructions. The resource scheduling service then executes the specific adjustment operations. For computing resources, elastic scaling is performed according to the adjustment instructions. When the target adjustment value is higher than the current allocation, new computing nodes are automatically added or existing nodes are given additional computing resources, such as increasing the number of CPU cores from 4 to 8, or increasing memory capacity from 8GB to 16GB. During the scaling process, data processing tasks are ensured to remain uninterrupted. When the target adjustment value is lower than the current allocation, scaling down is performed to release idle computing resources. For storage resources, capacity expansion or data migration operations are performed. Capacity expansion is achieved by adding nodes to the distributed storage cluster. Data migration is based on data access frequency and storage cost, migrating infrequently accessed data to low-cost storage nodes and retaining frequently accessed data on high-performance storage nodes to ensure efficient utilization of storage resources. For transmission resources, bandwidth QoS adjustments or route optimizations are performed. Based on the adjustment instructions, corresponding bandwidth priorities are allocated to different processing tasks. Core business tasks are allocated high-priority bandwidth, while non-core tasks are allocated normal-priority bandwidth. Simultaneously, the optimal data transmission path is selected through route optimization algorithms to reduce transmission latency. After adjustment, the processing performance and actual resource utilization of multi-source heterogeneous data are continuously monitored. Performance monitoring tools such as Prometheus are used to collect real-time performance metrics including completion time, data throughput, error rate, and resource utilization (CPU utilization, memory utilization, disk IOPS, network latency). These performance metrics are compared and analyzed with the expected effects of the resource optimization allocation scheme, and the deviation rate between the actual and expected values ​​is calculated. The deviation rate formula is: ; in, These are actual performance index values. This represents the expected performance metric value. The current adjustment's fit is assessed; if the deviation rate is below 10%, the adjustment is considered fit. If a deviation is found between actual performance and expected results, and the deviation rate exceeds 10%, then retraining the dynamic correlation model or locally correcting the resource optimization allocation scheme is triggered. During retraining, the latest performance monitoring data and resource consumption data are used to update the training dataset. For local corrections, the resource allocation is adjusted by 10%-20% based on the deviation, ensuring that resource adjustments continuously adapt to changes in data processing needs.

[0029] Through the complete implementation process described above, this method can accurately capture the dynamic correlation between the heterogeneous characteristics of multi-source heterogeneous data and resource requirements, and dynamically adjust resource allocation according to the real-time status of the data and processing needs, effectively solving the problem of deviation between resource allocation and actual needs in existing technologies.

[0030] Example 2

[0031] Please see Figure 4 The present invention also provides a system for dynamically adjusting multi-source heterogeneous data resources, used to implement the method for dynamically adjusting multi-source heterogeneous data resources in Embodiment 1. The system includes: The system comprises a data status acquisition module, a heterogeneous feature analysis module, a dynamic correlation modeling module, a scheme generation module, and a resource adjustment execution module. These modules work together to dynamically adjust resources for multi-source heterogeneous data.

[0032] The data status acquisition module uses a combination of the Flume acquisition framework and JDBC connection tools to achieve comprehensive acquisition of real-time status information of multi-source heterogeneous data. It can monitor changes in data structure, update frequency, and processing dependencies, and transmit the acquired real-time status information to subsequent modules in a standardized format.

[0033] The heterogeneous feature analysis module is deployed on the Spark computing framework. It uses distributed computing capabilities to quickly process real-time status information, analyze the heterogeneous features of the data, accurately identify the three types of resource requirements (computing, storage, and transmission) through a preset resource requirement assessment algorithm, and synchronize the analysis results to the dynamic association modeling module.

[0034] The dynamic correlation modeling module is built on the TensorFlow deep learning framework and integrates BP neural network and regression analysis algorithm. It can read historical data from heterogeneous data processing behavior archives for model training, generate dynamic correlation model, and update and optimize the model with new data according to preset period. At the same time, it outputs the resource requirements predicted by the model to the solution generation module.

[0035] The scheme generation module integrates the NSGA-Ⅲ multi-objective optimization algorithm library and resource monitoring tools. It can receive the predicted resource demand output by the dynamic correlation modeling module, combine it with the current resource status of the platform, generate a resource optimization allocation scheme through optimization algorithms, and pass the scheme to the resource adjustment execution module.

[0036] The resource adjustment execution module connects to the cloud platform's resource scheduling API, which can parse the adjustment instructions in the resource optimization allocation scheme, execute elastic scaling of computing resources, expansion and migration of storage resources, and bandwidth adjustment and routing optimization of transmission resources. At the same time, it integrates performance monitoring tools to collect performance indicators and resource utilization data after adjustment in real time, and generate feedback information to be sent back to the dynamic correlation modeling module to achieve closed-loop operation of the system.

[0037] During system operation, the data status acquisition module first initiates data collection and status monitoring, transmitting real-time status information to the heterogeneous feature analysis module. After completing feature analysis and resource demand identification, the heterogeneous feature analysis module sends the results to the dynamic correlation modeling module. The dynamic correlation modeling module trains a model using historical data and performs real-time predictions, outputting predicted resource demands to the solution generation module. The solution generation module combines resource status to generate an optimized allocation scheme, which is then passed to the resource adjustment execution module. The resource adjustment execution module executes adjustment operations and feeds back monitoring data. The dynamic correlation modeling module optimizes the model based on the feedback information, and the solution generation module adjusts the allocation scheme based on the feedback, forming a continuous optimization operating mechanism to ensure that the system can stably and efficiently achieve dynamic resource adjustment based on multi-source heterogeneous data.

[0038] In summary, this invention first initializes a heterogeneous data feature library and a heterogeneous data processing behavior archive, establishes a unified heterogeneous data feature description system and associates it with historical resource consumption data, then accurately acquires real-time status information of multi-source heterogeneous data, analyzes the heterogeneous features of the data based on this information to identify computing, storage, and transmission resource requirements, then uses machine learning algorithms to construct a dynamic correlation model between heterogeneous features and resource requirements and optimizes it periodically, combining the current total resources and real-time load of the data processing platform, generating a resource optimization allocation scheme through a multi-objective optimization algorithm, and finally performs dynamic resource adjustments and continuously monitors feedback to iteratively optimize the model and scheme. This effectively solves the problems of existing technologies that are difficult to accurately capture the dynamic correlation between heterogeneous features and resource requirements of multi-source heterogeneous data, and that rely on fixed rules or single indicators, leading to deviations between resource supply and actual demand. It improves the real-time performance and accuracy of data processing, reduces resource waste, and enhances system stability, ensuring that the technical solution can be implemented by those skilled in the art and can adapt to the dynamic changes of multi-source heterogeneous data, making it suitable for various multi-source heterogeneous data processing scenarios.

[0039] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0040] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for dynamically adjusting multi-source heterogeneous data resources, characterized in that, include: S1. Obtain real-time status information of multi-source heterogeneous data. The real-time status information includes the structural differences, data format, data update frequency, and data processing dependencies of multi-source heterogeneous data. S2. Based on real-time status information, analyze the heterogeneous characteristics of multi-source heterogeneous data, and identify the computing resource requirements, storage resource requirements, and transmission resource requirements of data processing tasks according to the heterogeneous characteristics. S3. Construct a dynamic correlation model between the heterogeneous characteristics of multi-source heterogeneous data and the requirements for computing resources, storage resources, and transmission resources. S4. Based on the dynamic correlation model and the real-time status information of multi-source heterogeneous data, generate a resource optimization allocation scheme. The scheme is used to adjust computing resources, storage resources and transmission resources. S5. Dynamically adjust computing resources, storage resources, and transmission resources according to the resource optimization allocation plan.

2. The method for dynamic adjustment of multi-source heterogeneous data resources according to claim 1, characterized in that, Before acquiring the real-time status information of multi-source heterogeneous data, the method further includes: Initialize the heterogeneous data feature library, and based on the preset data acquisition strategy, acquire metadata, structured information, semi-structured information and unstructured information of multi-source heterogeneous data from multiple data sources according to the preset period, and perform standardized processing on the information to establish a unified heterogeneous data feature description system. Continuously record historical resource consumption data of multi-source heterogeneous data under different processing tasks. The historical resource consumption data includes computing resource utilization, storage space usage, network transmission bandwidth usage, and task processing completion time. The historical resource consumption data is associated with the corresponding data heterogeneity characteristics to construct a heterogeneous data processing behavior profile.

3. The method for dynamically adjusting multi-source heterogeneous data resources according to claim 1, characterized in that, Step S1 includes: S11. When the structure of multi-source heterogeneous data changes, the structure of the newly added or modified data fields is parsed and their structural difference descriptions are updated. S12. When the update frequency of multi-source heterogeneous data exceeds the preset threshold, the data flow change rate and data volume growth trend are monitored in real time, and the update frequency information is updated. S13. When multi-source heterogeneous data is accessed or processed by a specific processing task, analyze its input-output relationship with other data, identify and update its processing dependencies.

4. The method for dynamic adjustment of multi-source heterogeneous data resources according to claim 1, characterized in that, Step S2 includes: S21. Based on the structural differences and data formats of multi-source heterogeneous data, analyze the computational resource overhead required for data cleaning, transformation and integration; S22. Based on the update frequency and data volume of multi-source heterogeneous data, predict the growth of data storage capacity and the storage space requirements for data caching; S23. Based on the processing dependencies of multi-source heterogeneous data, determine the network bandwidth requirements and data transmission latency requirements for data transmission between different processing nodes.

5. The method for dynamically adjusting multi-source heterogeneous data resources according to claim 1, characterized in that, Step S3 includes: S31. Obtain historical heterogeneous data characteristics and corresponding historical resource consumption data from the heterogeneous data processing behavior archive; S32. Using machine learning algorithms, historical heterogeneous data features are taken as input and corresponding historical resource consumption data are taken as output to train a resource demand prediction model. The prediction model is used to map the resource demand of multi-source heterogeneous data under different heterogeneous features. S33. Based on the resource demand prediction model, make real-time predictions on newly incoming heterogeneous data, and update and optimize the prediction model according to a preset cycle using the latest collected heterogeneous data characteristics and resource consumption data.

6. The method for dynamically adjusting multi-source heterogeneous data resources according to claim 5, characterized in that, Step S32 includes: S321. Represent the characteristics of historical heterogeneous data as feature vectors. The feature vectors include data structure complexity, data volume, update rate, access frequency, processing chain length, and data sensitivity. S322. Represent historical resource consumption data as a resource consumption vector. The resource consumption vector includes computing core usage, memory usage, storage throughput, network inbound traffic, network outbound traffic, and average task response time. S323. Based on feature vectors and resource consumption vectors, establish a nonlinear mapping relationship between heterogeneous features of multi-source heterogeneous data and resource demand through regression analysis or neural network models.

7. The method for dynamically adjusting multi-source heterogeneous data resources according to claim 1, characterized in that, Step S4 includes: S41. Based on the predicted resource requirements output by the dynamic correlation model, combined with the total amount and real-time load of the available computing resources, storage resources and transmission resources of the current data processing platform. S42. Employ a multi-objective optimization algorithm to comprehensively consider the real-time requirements of data processing, resource utilization, processing costs, and system stability constraints, and generate detailed resource allocation decisions. S43. The resource allocation decision clearly stipulates the amount of computing resources, storage resources, and transmission resources that each data processing task should receive within a specified time.

8. The method for dynamically adjusting multi-source heterogeneous data resources according to claim 1, characterized in that, Step S5 includes: S51. Send the resource adjustment instruction contained in the resource optimization allocation plan to the resource scheduling service. The resource adjustment instruction includes the resource type, the adjustment target value, and the effective time. S52, The resource scheduling service performs elastic scaling of computing resources, capacity expansion or data migration of storage resources, and bandwidth QoS adjustment or routing optimization of transmission resources. S53. Continuously monitor the processing performance and actual resource utilization of the adjusted multi-source heterogeneous data, and use it as feedback information for iterative optimization of the dynamic correlation model and adjustment of the resource optimization allocation scheme.

9. A method for dynamically adjusting multi-source heterogeneous data resources according to claim 8, characterized in that, The continuous monitoring of the processing performance and actual resource utilization of the adjusted multi-source heterogeneous data includes: Real-time acquisition of performance metrics including completion time, data processing throughput, error rate, and resource utilization (CPU utilization, memory utilization, disk IOPS, network latency) of the adjusted data processing tasks. Compare and analyze the performance indicators with the expected effects of the resource optimization allocation plan to assess the suitability of the current adjustment; When a deviation is found between actual performance and expected results, the dynamic correlation model is retrained or the resource optimization allocation scheme is partially modified.

10. A system for dynamically adjusting multi-source heterogeneous data resources, applied to the method for dynamically adjusting multi-source heterogeneous data resources as described in any one of claims 1-9, characterized in that, include: The data status acquisition module is used to acquire real-time status information of multi-source heterogeneous data. The real-time status information includes the structural differences, data format, data update frequency, and data processing dependencies of the multi-source heterogeneous data. The heterogeneous feature analysis module is used to analyze the heterogeneous features of multi-source heterogeneous data based on real-time status information, and to identify the computing resource requirements, storage resource requirements and transmission resource requirements of data processing tasks based on the heterogeneous features. The dynamic correlation modeling module is used to construct a dynamic correlation model between the heterogeneous characteristics of multi-source heterogeneous data and the requirements for computing resources, storage resources, and transmission resources. The scheme generation module is used to generate resource optimization allocation schemes based on the dynamic correlation model and the real-time status information of multi-source heterogeneous data. The schemes are used to adjust computing resources, storage resources and transmission resources. The resource adjustment execution module is used to dynamically adjust computing resources, storage resources, and transmission resources according to the resource optimization allocation plan.