Intelligent data access and integration method based on rapid dynamic load adjustment
Through intelligent data access and integration methods, data sources are automatically detected and identified, data source characteristics are adapted, data source characteristics are self-learning, data quality problems are corrected, processing processes are dynamically adjusted, multiple databases and interface management are supported, and system performance is monitored, and the limitations of the existing technology in multi-source data adaptation, data quality assurance, load balancing and intelligent storage are solved, and efficient and flexible data processing and storage are achieved.
Patent Information
- Application Number
- CN202411833413.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data access and integration methods have great limitations in adapting multi-source heterogeneous data, ensuring data quality, dynamic expansion and load balancing, intelligent storage, interface security and system monitoring optimization, and are difficult to meet modern data processing needs.
An intelligent data access and integration method is proposed, including automatic detection and identification of multiple data sources, adapting according to the characteristics of the data source, predicting and automatically correcting data quality problems through a self-learning mechanism, automatically adjusting data processing processes according to real-time data load conditions, supporting multiple types of databases, developing and maintaining data access interfaces, and monitoring system performance and data processing processes.
It realizes efficient data access, processing, storage and monitoring functions, significantly improves data compatibility and access efficiency, ensures data integrity and accuracy, realizes flexible load balancing and intelligent storage management, and enhances the overall performance and stability of the system.
Smart Images

Figure CN119988467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data access and processing, and in particular to an intelligent data access and integration method based on fast dynamic load adjustment. Background Art
[0002] With the development of information technology and the popularity of data-driven decision-making, modern enterprises and institutions are increasingly relying on multi-source and diverse data for analysis and decision-making. Especially in the fields of finance, healthcare, Internet of Things and smart cities, the scale and complexity of data are increasing. These data are usually generated from different sources, including structured data and unstructured data, or processed in real time and batch mode. In order to effectively utilize these diverse data, data access and integration methods have gradually become a key component of data processing technology.
[0003] Traditional data access and integration methods are usually based on data warehouses and ETL (extraction, transformation, and loading) processes. These methods have good support for structured data and are suitable for data integration in predefined formats. However, in the face of unstructured data, real-time data, and heterogeneous data sources, traditional ETL methods are powerless. Data Lake, as a newer integration method, supports the storage of multiple data types, but due to the lack of effective support for data quality and governance, it is often difficult to meet the needs of efficient data processing in practical applications. Therefore, existing data integration methods have great limitations in adapting to heterogeneous data sources and ensuring data quality, and it is difficult to cope with the requirements of modern data applications for real-time and multi-source data processing.
[0004] In addition, when dealing with multiple data sources, existing data access and integration methods generally rely on manual configuration, which is cumbersome and error-prone when dealing with diverse and dynamically changing data sources. For example, the differences in data source formats, structures, and access rights require manual writing of adaptation rules or adjustment of data models, which makes the data processing process time-consuming and inflexible. As data sources continue to increase, existing methods are difficult to efficiently adapt to new data access requirements.
[0005] Data quality management is also a major challenge in traditional methods. Existing data quality management methods are mostly post-detection, that is, checking and cleaning after data is loaded, but this method leads to the accumulation of quality problems, and the timeliness and accuracy of data cannot be guaranteed. In addition, data quality problems usually include missing values, inconsistencies, and outliers. Although in traditional data quality management, problems can be troubleshooted through rule detection or experience-based settings, these manually configured methods are often not flexible enough and difficult to adapt to the needs of automated and intelligent data quality management.
[0006] Traditional integration methods also have obvious shortcomings in dealing with dynamic data loads. As the amount of data grows and traffic fluctuates, the system needs to have load adaptation and dynamic expansion capabilities to avoid resource waste or system bottlenecks. However, traditional integration methods usually lack dynamic adjustment of resource allocation, and load adjustment is often completed by manually expanding nodes or presetting the number of nodes, which makes it difficult to flexibly respond to changes in real-time load. This deficiency makes existing data access and integration methods prone to performance bottlenecks when facing peak traffic, and leads to low resource utilization when under low load.
[0007] In terms of data storage, there is usually a distinction between "hot and cold data" due to the different access frequencies and importance of data. Frequently accessed "hot data" requires fast reading and writing to meet real-time needs, while less frequently accessed "cold data" focuses more on storage costs. Existing methods often adopt static storage strategies, lack dynamic management and migration mechanisms for hot and cold data, and cannot achieve intelligent storage optimization based on data access patterns. In addition, data migration operations between different storage systems usually need to be handled manually, lack flexibility, and it is difficult to automatically adjust the storage method according to changes in data usage frequency.
[0008] In terms of data access interfaces, traditional methods are mostly based on a single API structure, such as RESTful API. Although this interface model can meet basic data access requirements, it lacks flexible support for multiple interfaces and management mechanisms for API versions. In addition, existing methods are usually limited in terms of user authentication and permission control, making it difficult to achieve fine-grained control of different users and applications, which poses a risk to data access security. At the same time, insufficient monitoring and compliance review of interface usage make it difficult to monitor and respond to API usage in a timely manner in different scenarios, reducing the management efficiency of data interfaces.
[0009] In terms of system monitoring and optimization, traditional methods usually only provide static status monitoring, such as simple collection of key performance indicators such as CPU and memory. However, in complex data integration scenarios, fluctuations in system load and the emergence of sudden performance bottlenecks are often difficult to predict, so it is difficult to reflect the operating status of the system in real time by relying solely on static monitoring. In addition, traditional methods usually lack automated optimization mechanisms. Even if the bottleneck of the system is detected, manual intervention is still required, and there is a lack of adaptive adjustment capabilities to load changes.
[0010] In summary, existing data access and integration methods have great limitations in adapting multi-source heterogeneous data, ensuring data quality, dynamic expansion and load balancing, intelligent storage, interface security, and system monitoring optimization. In order to meet the needs of modern data processing, an intelligent data access and integration method is needed that can achieve automated data adaptation, real-time quality management, flexible load balancing, and intelligent storage management to improve the overall efficiency and flexibility of the system. Summary of the invention
[0011] In order to solve the above problems in the prior art, the present invention proposes an intelligent data access and integration method, which is characterized in that the data access and integration method comprises the following steps: Automatically detect and identify multiple data sources and their characteristics; Adapting the data source according to the identified data source characteristics; Predict and automatically correct data quality issues through self-learning mechanisms; Clean and transform raw data; Automatically adjust data processing flow based on real-time data load, including simplifying data processing steps when data load is low, launching additional ETL services when data load is high, and starting or shutting down processing nodes based on data processing needs; Store data and support multiple types of databases; Develop and maintain data access interfaces; Monitor system performance and data processing flow.
[0012] The steps of automatically detecting and identifying multiple data sources include: Collecting characteristic information of various data sources, including data format, structure and access rights; Collect and analyze metadata information of data sources through automated scripts or standardized interfaces; Create a data source directory; Verify accessibility of data source connections; Sampling data from a data source.
[0013] The step of adapting the data source according to the data source characteristics comprises: receiving data source characteristic information; Convert data formats from different data sources; Generate field mapping rules based on data source characteristics to achieve automatic matching of fields; Standardize the structure of different data sources to ensure data compatibility.
[0014] The steps for predicting and automatically correcting data quality issues include: Optimize data quality detection rules based on historical data through self-learning mechanism; Monitor data for integrity, consistency, and anomaly detection; Use statistical analysis and machine learning algorithms to predict and automatically correct data quality issues; Trigger data correction actions when data quality issues are detected.
[0015] The step of adjusting the data processing flow according to the real-time data load condition comprises: The total system load is calculated using an automatic scheduling algorithm based on load thresholds. The formula of the automatic scheduling algorithm is:
[0016] when When , the number of expansion nodes is calculated according to the following formula : ; in: : CPU load ratio, which indicates the ratio of the current CPU usage to the total capacity; : Memory load ratio, which indicates the ratio of current memory usage to total capacity; : The load ratio of network bandwidth, which indicates the ratio of the current usage of network bandwidth to the total bandwidth; : Disk I / O load ratio, which indicates the ratio of the current usage of disk I / O to the total I / O capacity; is the weight coefficient; is the number of nodes that need to be expanded; is the current total system load; is the upper load threshold; is the expansion ratio function.
[0017] The expansion ratio function The calculation formula is: in, Indicates the rate of change of CPU usage per unit time; Indicates the rate of change of the load of memory resources; Indicates the rate of change of network bandwidth load; Indicates the rate at which the disk I / O load changes.
[0018] The steps to store data include: Dynamically select storage solutions based on data access frequency and importance to achieve separate storage of hot and cold data; Automatically migrate data between different storage systems when data access frequency changes.
[0019] The steps of developing and maintaining the data access interface include: Support data access via RESTful API and GraphQL interface; Manage multiple API versions; Control data access rights through authentication and authorization; Monitor the access frequency, response time, and usage patterns of API requests and generate log reports.
[0020] The steps of monitoring system performance and data processing flow include: Display system operation status through visual monitoring interface; Monitor key performance indicators such as CPU usage, memory usage, data processing latency, network traffic, and database query response time; Real-time assessment of system health through system log and data flow analysis; Dynamically adjust system resource configuration based on real-time monitoring data; Predict system resource bottlenecks through trend analysis and expand resources in advance.
[0021] The intelligent data access and integration method of the present invention realizes efficient data access, processing, storage and monitoring functions. By automatically detecting data sources and intelligently adapting, the method adapts to a variety of data formats and structures, significantly improving data compatibility and access efficiency. The self-learning data quality management module ensures data integrity and accuracy and reduces data error rates. Real-time load monitoring and dynamic expansion mechanisms enable the system to adjust processing resources according to load, maintain efficient operation, and effectively respond to high concurrency and large data volume scenarios. Dynamic storage management realizes the separation of hot and cold data, optimizes storage costs, and enhances the overall performance and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present application, but do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A flow chart of the intelligent data access and integration method of the present invention is shown.
[0023] Figure 2A flow chart showing the data load adjustment mechanism. DETAILED DESCRIPTION
[0024] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments, wherein the illustrative embodiments and descriptions are only used to explain the present invention but are not intended to limit the present invention.
[0025] like Figure 1 As shown, the present invention provides an intelligent data access and integration method, covering the full process of automatic data source identification, data adaptation, quality management, real-time load adjustment, data storage and interface management. The system can be widely used in complex environments that need to process large amounts of heterogeneous and dynamic data, such as the Internet of Things, financial data analysis, and smart cities.
[0026] Step 1: Automatically detect and identify data sources, including: After the system is started, the data source automatic identification module will automatically detect and identify multiple data sources and their characteristics. First, through automated scripts or standardized interfaces (APIs), the system collects characteristic information including data format, structure, access rights, etc. The collected characteristic information is classified and organized through data source directory units, and data source connections are verified to confirm the accessibility of each data source. Finally, the system performs preliminary sampling of data sources to obtain representative data.
[0027] Step 2: Data source adaptation, including: like Figure 1 As shown in the figure, after identifying the characteristics of the data source, the system completes the adaptation process of the data source through the data adaptation module. The adaptation process mainly includes the following operations: 1. Format conversion: Standardize data in different formats in the data source, such as converting XML, JSON and other formats into a unified structured data format.
[0028] 2. Field mapping: Generate field mapping rules based on the collected data source characteristic information to automatically achieve field matching and docking.
[0029] 3. Structural standardization: The structures of different data sources are standardized to ensure data compatibility, thus laying the foundation for subsequent data processing steps.
[0030] Step 3: Data quality management, including: In order to ensure data quality, the system starts the intelligent data quality management module after data access, and predicts and corrects data quality problems through a self-learning mechanism. The self-learning mechanism automatically optimizes data quality detection rules based on historical data and monitors the integrity, consistency and anomalies of the data in real time.
[0031] Data quality management is mainly divided into the following steps: 1. Completeness check: Ensure that there are no missing or empty values in the data set.
[0032] 2. Consistency check: Check whether the values of data fields meet the preset standards to avoid data conflicts or duplications.
[0033] 3. Anomaly detection: Statistical analysis and machine learning algorithms are used to detect anomalies in data sets. When data quality issues are detected, the system triggers the data preprocessing module to perform corresponding correction operations.
[0034] Step 4: Data preprocessing, including: After discovering quality problems, the intelligent data quality management module will call the data preprocessing module to clean and convert the data. The cleaning process includes missing value filling, data format standardization, and outlier processing. The preprocessing module processes missing and abnormal data by filling, deleting, or replacing to ensure that the data quality meets the standards.
[0035] Step 5: Data load adjustment, including: like Figure 2 As shown, the present invention adopts a dynamic load adjustment mechanism to adapt to the fluctuation of system data load. The mechanism is based on an automatic scheduling algorithm based on the load threshold. ) to decide whether to expand or reduce processing nodes in order to maintain processing performance during load peaks and save resources during load troughs. Specifically, the total load calculation formula is as follows: in, , , and Respectively represent the current usage load rate of CPU, memory, network bandwidth and disk I / O in the system. , , and They are the weight coefficients of various resources. The weight coefficients can be adjusted according to different application scenarios or resource sensitivity to ensure that the usage of each resource is reasonably reflected.
[0036] when Exceeding the preset upper load threshold (i.e. ), the system will trigger the node expansion mechanism and calculate the number of nodes to be expanded using the following formula :
[0037] in, is the expansion ratio function, which is used to determine the number of expansion nodes based on the current resource usage rate. Specifically, the calculation formula of the expansion ratio function is as follows:
[0038] The expansion ratio function dynamically captures the speed of the load increase trend by calculating the load change rate of each resource. When the load rate of a resource is high, the expansion ratio function will give the resource a higher weight so that expansion can be triggered faster. Specifically, Indicates the rate of change of CPU usage. Indicates the rate of change of memory usage. Indicates the load change rate of the network bandwidth. Indicates the change rate of disk I / O usage.
[0039] This dynamic load adjustment mechanism ensures that the system can respond quickly during peak load periods and expand the number of nodes on demand to avoid performance bottlenecks; when the load drops back to normal levels, the system automatically reduces unnecessary nodes, thereby optimizing resource utilization and reducing operating costs. This dynamic load balancing mechanism based on load change rate enables the system to flexibly adapt to different workload conditions and ensure that resources are optimally allocated in a changing load environment.
[0040] Step 6: Data storage management, including: The storage module implements dynamic data storage management through the hot and cold data separation strategy. The system will select the appropriate storage method based on the access frequency and importance of the data. Frequently accessed data (hot data) is preferentially stored in a high-performance relational database, while data with low access frequency (cold data) is stored in a low-cost NoSQL database or cloud storage. In addition, when the data access frequency changes, the system will automatically adjust the storage solution, migrate hot data to high-performance storage or migrate cold data to low-cost storage to reduce storage costs.
[0041] Step 7: Data access interface management, including: The system develops and maintains data access interfaces through the API management module. The module supports data access through RESTful API and GraphQL interfaces, and can provide flexible access options for different users and applications. At the same time, the API management module has multi-version management capabilities and supports simultaneous maintenance of multiple API versions so that updates can be pushed without affecting existing users. In addition, the system controls data access rights through identity authentication and authorization mechanisms to ensure data security.
[0042] Step 8: System monitoring and optimization, including: During the overall operation of the system, the monitoring and optimization module displays the real-time status of the system through a visual monitoring interface, including key performance indicators such as CPU usage, memory usage, data processing latency, network traffic, and database query response time. System log and data traffic analysis can help administrators evaluate the health of the system in real time.
[0043] When the monitoring module detects performance bottlenecks or excessive load, the system will automatically adjust resource configuration to ensure smooth operation of the system. In addition, through trend analysis, the monitoring module can predict system resource bottlenecks in advance and prompt administrators to expand necessary resources before the peak period arrives to ensure system efficiency and stability.
[0044] Through the above steps, the present invention provides an intelligent data access and integration method, which can efficiently realize automatic adaptation, quality management, real-time load adjustment, dynamic storage and access control of multi-source data.
[0045] The above description is only a preferred embodiment of the present invention, so all equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the scope of the patent application of the present invention.
Claims
1. An intelligent data access and integration method, characterized in that: The data access and integration method comprises the following steps: Step 1: Automatically detect and identify multiple data sources and their characteristics; Step 2: Adapt the data source according to the identified data source characteristics; Step 3: Predict and automatically correct data quality issues through self-learning mechanisms; Step 4: Clean and transform the raw data; Step 5: Automatically adjust the data processing flow based on real-time data load conditions, including simplifying data processing steps when data load is low, starting additional ETL services when data load is high, and starting or shutting down processing nodes based on data processing needs; Step 6: Store data and support multiple types of databases; Step 7: Develop and maintain data access interfaces; Step 8: Monitor system performance and data processing flow.
2. An intelligent data access and integration method as claimed in claim 1, characterized in that: The steps of automatically detecting and identifying multiple data sources include: Collecting characteristic information of various data sources, including data format, structure and access rights; Collect and analyze metadata information of data sources through automated scripts or standardized interfaces; Create a data source directory; Verify accessibility of data source connections; Sampling data from a data source.
3. The intelligent data access and integration method according to claim 1, characterized in that: The step of adapting the data source according to the data source characteristics comprises: receiving data source characteristic information; Convert data formats from different data sources; Generate field mapping rules based on data source characteristics to achieve automatic matching of fields; Standardize the structure of different data sources to ensure data compatibility.
4. The intelligent data access and integration method according to claim 1, characterized in that: The steps for predicting and automatically correcting data quality issues include: Optimize data quality detection rules based on historical data through self-learning mechanism; Monitor data for integrity, consistency, and anomaly detection; Use statistical analysis and machine learning algorithms to predict and automatically correct data quality issues; Trigger data correction actions when data quality issues are detected.
5. The intelligent data access and integration method according to claim 1, characterized in that: The step of adjusting the data processing flow according to the real-time data load condition comprises: The total system load is calculated using an automatic scheduling algorithm based on load thresholds. The formula of the automatic scheduling algorithm is: when When , the number of expansion nodes is calculated according to the following formula : ; in: is the current total system load; : CPU load ratio, indicating the current CPU usage With total capacity The ratio of : The memory load ratio, indicating the current memory usage With total capacity The ratio of : The load ratio of the network bandwidth, indicating the current usage of the network bandwidth Total bandwidth The ratio of : Disk I / O load ratio, indicating the current usage of disk I / O Total I / O Capacity The ratio of is the weight coefficient; is the number of nodes that need to be expanded; is the upper load threshold; is the expansion ratio function.
6. An intelligent data access and integration method as claimed in claim 5, characterized in that: The expansion ratio function The calculation formula is: in, Indicates the rate of change of CPU usage per unit time; Indicates the rate of change of the load of memory resources; Indicates the rate of change of network bandwidth load; Indicates the rate at which the disk I / O load changes.
7. The intelligent data access and integration method according to claim 1, characterized in that: The steps to store data include: Dynamically select storage solutions based on data access frequency and importance to achieve separate storage of hot and cold data; Automatically migrate data between different storage systems when data access frequency changes.
8. An intelligent data access and integration method as claimed in claim 1, characterized in that: The steps of developing and maintaining the data access interface include: Support data access via RESTful API and GraphQL interface; Manage multiple API versions; Control data access rights through authentication and authorization; Monitor the access frequency, response time, and usage patterns of API requests and generate log reports.
9. The intelligent data access and integration method according to claim 1, characterized in that: The steps of monitoring system performance and data processing flow include: Display system operation status through visual monitoring interface; Monitor key performance indicators such as CPU usage, memory usage, data processing latency, network traffic, and database query response time; Real-time assessment of system health through system log and data flow analysis; Dynamically adjust system resource configuration based on real-time monitoring data; Predict system resource bottlenecks through trend analysis and expand resources in advance.
Citation Information
Patent Citations
Data storage node adjustment method and system
CN104850634A
Self-adaptive configuration adjustment method and system applied to distributed system
CN118295732A
Data fragmentation and table division autonomous extension system and method
CN118606295A
Intelligent data cleaning system based on real-time database
CN118885473A
Data analysis and governance integrated platform based on multi-dimensional data
CN119025582A