Intelligent data access and integration system based on dynamic expansion architecture

The intelligent data access and integration system based on a dynamically expandable architecture solves the problems of low efficiency, poor scalability, and difficulty in guaranteeing data quality in the processing of multi-source heterogeneous data in existing technologies. It achieves efficient and secure data access and integration, and adapts to the needs of complex data scenarios.

CN119311754BActive Publication Date: 2025-10-21STATE GRID SIJI DIGITAL TECH (BEIJING) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411833297.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-10-21
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing data processing systems suffer from inefficiency, poor scalability, difficulty in guaranteeing data quality, slow processing speed, and insufficient security when dealing with large-scale, multi-source, heterogeneous data. In particular, they lack flexibility and intelligent adaptability when processing unstructured data.

Method used

An intelligent data access and integration system based on a dynamically scalable architecture is adopted, including a data source automatic identification module, a data adaptation module, an intelligent data quality management module, a data preprocessing module, a data integration module, a data storage module, an API management module, and a monitoring and optimization module. Through a self-learning mechanism, dynamic load balancing, and modular design, it achieves automatic identification, adaptation, and efficient integration of multi-source data.

Benefits of technology

It enables automatic identification and adaptation of various heterogeneous data sources, improves data access speed, ensures data quality, optimizes resource allocation, enhances system flexibility and scalability, meets data access needs of different application scenarios, and provides efficient and secure data processing and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311754B_ABST
    Figure CN119311754B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent data access and integration system based on a dynamic expansion architecture, comprising: a data source automatic identification module, which automatically detects and identifies various data sources and their characteristics; a data adaptation module, which adapts according to the characteristics of the data sources; an intelligent data quality management module, which predicts and automatically corrects data quality problems based on a self-learning mechanism; a data preprocessing module, which is used for cleaning and converting original data; a data integration module, which adjusts the data processing flow according to real-time load through a dynamic expansion architecture and a load balancing mechanism; a data storage module, which supports various database types and realizes cold and hot data separation storage; an API management module, which develops and maintains data access interfaces; and a monitoring and optimization module, which monitors system performance and data flow in real time and guarantees stable and efficient operation of the system. The application improves data access efficiency, quality and utilization of system resources and ensures system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an intelligent data access and integration system based on a dynamic expansion architecture. Background Art

[0002] With the rapid development of big data and the Internet of Things (IoT) technologies, more and more businesses and organizations are facing the challenge of processing and managing massive amounts of information from multiple data sources. These data sources can be structured, semi-structured, or unstructured, encompassing a variety of formats, protocols, and storage methods. Data access, integration, and management have become critical technical issues, particularly in areas such as financial services, healthcare, smart cities, and the Industrial Internet of Things. Acquiring high-quality, usable data from diverse sources and ensuring its efficient and secure processing and analysis have become pressing challenges in modern data-driven systems.

[0003] Traditional data access and integration methods rely primarily on manual or semi-automated approaches to identify and process data sources. These methods often require extensive predefined rules and manual intervention to adapt to the formats and protocols of diverse data sources. While this approach may be feasible for small amounts of data, its limitations become apparent as the number and scale of data sources increase. Traditional methods, in particular, face challenges in efficiency, flexibility, and scalability when dealing with large-scale and diverse data sources.

[0004] Common solutions in existing technologies include data warehouses, ETL (extract, transform, and load) tools, and data lakes.

[0005] A data warehouse is a centralized system for storing and managing large amounts of structured data, widely used in business intelligence (BI) and analytics scenarios. Data warehouses are designed to provide high-quality data for decision support systems by integrating and cleaning data from multiple sources to create a unified view, thereby supporting data analysis. However, a major limitation of traditional data warehouses is that they can only process structured data, with limited support for unstructured or semi-structured data. With the diversification of data types (such as text, images, and sensor data), relying solely on data warehouses can no longer meet the needs of modern data management.

[0006] ETL tools, as a common data processing tool, are mainly used to extract data from different data sources, transform it, and then load it into the data warehouse. ETL tools play an important role in the data integration process, especially in enterprise information systems. However, traditional ETL tools have shown some shortcomings when faced with large-scale, multi-source heterogeneous data. First, the ETL process usually requires complex manual configuration, and the data processing speed is slow when processing large-scale data. Second, ETL tools have poor adaptability to data sources, limited scalability, and cannot effectively process unstructured data. In addition, ETL tools often rely on hard-coded transformation rules and lack intelligent adaptation capabilities.

[0007] A data lake is a centralized storage platform capable of storing both structured and unstructured data, addressing the inability of data warehouses to handle unstructured data. Data lakes are typically used to store raw data without any preprocessing or transformation, enabling them to accommodate data of various types and formats. While data lakes offer advantages in handling large-scale, heterogeneous data, they suffer from significant deficiencies in data governance, data quality management, and data query efficiency. Because data in data lakes is unprocessed, users often need to perform complex preprocessing operations before using the data. The sheer volume of data can lead to data redundancy and poor data quality. Furthermore, data lakes are complex in architecture and expensive to maintain and manage.

[0008] As enterprise data demands grow and data sources diversify, traditional data warehouses, ETL tools, and data lakes face numerous technical limitations when coping with modern data environments. Specifically, existing technologies have problems in the following areas:

[0009] Existing systems often rely on manual configuration to handle access to diverse data sources. This approach is not only time-consuming and labor-intensive, but also difficult to adapt to the rapid changes in multi-source data. Each new data source often requires redesigning and reconfiguring the corresponding adaptation rules, resulting in inefficient data access. Furthermore, traditional data systems have limited capabilities when dealing with unstructured data (such as text, images, and video), lacking flexibility and scalability.

[0010] Data quality is a key challenge in large-scale data processing. Existing systems mostly rely on manual post-processing data cleansing, failing to monitor and correct data quality issues in real time. For example, in ETL tools, data quality issues are typically identified and corrected after data loading is complete. This approach not only prolongs data processing time but can also reduce the accuracy of subsequent analysis results. Furthermore, traditional systems lack intelligent self-learning mechanisms, making them unable to predict and optimize based on historical data quality issues.

[0011] Traditional data processing systems often suffer from slow processing speeds and long response times when faced with large amounts of data. This is especially true when processing heterogeneous data from multiple sources. Existing ETL tools, with their hard-coded rules and fixed processes, struggle to maintain both efficiency and flexibility. Large fluctuations in data traffic can easily lead to system overload and reduced processing efficiency. Furthermore, existing systems lack dynamic scalability, making it impossible to adjust the allocation of processing resources based on changes in data traffic.

[0012] Traditional data systems are mostly based on fixed architectures, resulting in poor scalability. As the number of data sources and data volumes increase, the load capacity and processing performance of existing systems often struggle to meet demand. Furthermore, the lack of modularity in traditional systems necessitates extensive system refactoring when adding new functionality, increasing development and maintenance costs.

[0013] Existing data systems typically provide only limited data access interfaces, failing to meet the flexible data access needs of diverse application scenarios. Especially when dealing with multiple data sources, diverse user groups, and varying permissions, existing systems' API management capabilities are limited, making it difficult to achieve granular control over data access. Furthermore, data access security is insufficient, making it unable to cope with complex permission management and data privacy protection requirements.

[0014] Therefore, developing an intelligent, efficient and automated multi-source data access and integration system has become a key direction to solve existing technical problems. Summary of the Invention

[0015] In order to solve the above problems in the prior art, the present invention proposes an intelligent data access and integration system based on a dynamic expansion architecture, comprising:

[0016] Data source automatic identification module, used to automatically detect and identify multiple data sources and their characteristics;

[0017] a data adaptation module, connected to the data source automatic identification module, for adapting the data source according to the identified characteristics;

[0018] Intelligent data quality management module, which uses a self-learning mechanism to predict and automatically correct data quality issues;

[0019] Data preprocessing module, used to clean and transform raw data;

[0020] The data integration module uses a dynamic expansion architecture to automatically adjust the data processing process according to the real-time data load. When the data load is low, the data processing steps are simplified, and when the data load is high, additional ETL services are started. The dynamic load balancing mechanism is used to start or shut down additional processing nodes according to data processing needs.

[0021] Data storage module, used to store data and support multiple types of databases;

[0022] API management module, used to develop and maintain data access interfaces;

[0023] Monitoring and optimization module, used to monitor system performance and data processing flow.

[0024] The data source automatic identification module includes: a data collection unit for detecting and identifying multiple data sources and their characteristics, including data format, structure and access rights; an automated script or standardized interface for collecting and analyzing metadata information of the data source; a data source directory establishment unit for establishing a data source directory based on the collected metadata information; a data connection verification unit for verifying the accessibility of the data source connection; and a data sampling unit for sampling data from the data source.

[0025] The data adaptation module includes: an interface connected to the data source automatic identification module, used to receive data source characteristic information; a data format conversion unit, used to convert data formats of different data sources;

[0026] The field mapping unit is used to generate field mapping rules based on the characteristics of the data source to achieve automatic matching of field correspondences;

[0027] The structure standardization unit is used to standardize the structures of different data sources to ensure data compatibility and accessibility.

[0028] The intelligent data quality management module includes: a self-learning mechanism for optimizing data quality detection rules through historical data;

[0029] a quality monitoring unit to monitor data integrity, consistency, and anomaly detection;

[0030] A statistical analysis and machine learning unit for predicting and automatically correcting data quality issues, wherein the machine learning unit includes algorithms for missing value detection, duplicate data detection, and anomaly detection;

[0031] A collaborative interface with the data preprocessing module, used to trigger data correction operations when data quality issues are detected.

[0032] The dynamic expansion architecture adopted by the data integration module automatically adjusts the data processing flow based on the automatic scheduling algorithm of the load threshold; the formula of the automatic scheduling algorithm is as follows:

[0033]

[0034] in:

[0035] : CPU load ratio, which indicates the ratio of the current CPU usage to the total capacity;

[0036] : Memory load ratio, which indicates the ratio of current memory usage to total capacity;

[0037] : The load ratio of network bandwidth, which indicates the ratio of the current usage of network bandwidth to the total bandwidth;

[0038] : Disk I / O load ratio, which indicates the ratio of the current disk I / O usage to the total I / O capacity;

[0039] is the weight coefficient.

[0040] when When , the system will expand the nodes, and the number of expanded nodes is calculated as follows:

[0041]

[0042] in: is the number of nodes that need to be expanded;

[0043] is the current total system load;

[0044] is the upper load threshold;

[0045] is the expansion ratio function.

[0046] The specific formula of the expansion ratio function is:

[0047]

[0048] in:

[0049] There are separate upper thresholds for CPU, memory, network bandwidth, and disk I / O.

[0050] The data storage module includes: a storage unit that supports multiple database types, including relational databases, NoSQL databases, and cloud-based storage systems;

[0051] Dynamic selection unit, used to dynamically select storage solutions based on data access frequency and importance, achieving separate storage of hot and cold data;

[0052] The data migration unit is used to automatically migrate data between different storage systems when the data access frequency changes.

[0053] The API management module includes:

[0054] Supports data access unit with multiple interfaces;

[0055] Multi-version management unit, used to support simultaneous maintenance of multiple API versions;

[0056] Identity authentication and authorization unit, which performs user authentication and controls data access rights;

[0057] Use monitoring and compliance review units to monitor the access frequency, response time, and usage patterns of API requests and generate log reports.

[0058] The monitoring and optimization module includes: a visual monitoring interface for displaying the system operation status; a key performance indicator monitoring unit for monitoring CPU utilization, memory utilization, data processing latency, network traffic, and database query response time; a system health assessment unit for real-time evaluation of system health through system log and data traffic analysis; an automatic optimization unit for dynamically adjusting system resource configuration based on real-time monitoring data; and a trend analysis unit for predicting system resource bottlenecks and expanding resources in advance.

[0059] The intelligent data access and integration system of the present invention realizes automatic identification, adaptation and efficient integration of multiple heterogeneous data sources through modular design, and has the following beneficial effects:

[0060] The automatic data source identification and data adaptation modules automatically identify and adapt to diverse data sources, reducing manual intervention and accelerating data access. The intelligent data quality management module utilizes a self-learning mechanism to automatically predict and correct data issues, ensuring data quality and accuracy. The data integration module utilizes a dynamically scalable architecture and load balancing mechanism to adjust processing nodes based on real-time load and optimize resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application, but do not constitute an improper limitation of the present invention. In the drawings:

[0062] Figure 1 Shows the overall structural block diagram of the intelligent data access and integration system of the present invention;

[0063] Figure 2 A functional structure diagram of the data integration module is shown. DETAILED DESCRIPTION

[0064] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The exemplary embodiments and descriptions are only used to explain the present invention but are not intended to limit the present invention.

[0065] like Figure 1 As shown, this application provides an intelligent data access and integration system. Through the collaborative operation of multiple modules, this system achieves automated access, intelligent processing, and efficient management of multi-source data. The system's main components include: an automatic data source identification module 101, a data adaptation module 102, an intelligent data quality management module 103, a data preprocessing module 104, a data integration module 105, a data storage module 106, an API management module 107, and a monitoring and optimization module 108.

[0066] The automatic data source identification module 101 is used to automatically detect and identify various data sources and their characteristics, including data format, structure, and access permissions. It can efficiently process heterogeneous data sources from different types, such as relational databases, NoSQL databases, file systems, and API interfaces. Through automated scripts or standardized interfaces, this module collects and analyzes metadata for each data source, automatically creates a data source catalog, and performs data connection verification and preliminary data sampling.

[0067] This module's design allows the system to autonomously identify and adapt to the formats and structures of different data sources without pre-defining data models, significantly reducing the complexity and time cost of manual intervention. In large-scale data access scenarios, the automatic data source identification module can greatly improve data access efficiency while ensuring the accuracy and consistency of data source access.

[0068] For example, when the system connects to multiple heterogeneous data sources, the automatic data source identification module can quickly determine the format (e.g., XML, JSON, CSV), structure (e.g., relational table structure, document structure), and access permissions of each data source, and pass this information to subsequent modules for further processing. This automated feature greatly improves the system's adaptability to diverse data sources and is particularly suitable for complex scenarios that require processing large amounts of real-time or dynamically updated data, such as IoT device data, financial data streams, and cloud service interfaces. This module not only supports common data sources but can also be expanded based on preset rules to support new data source types in the future.

[0069] Furthermore, the automatic data source identification module automatically verifies data connections during data source access, ensuring data source accessibility and data integrity. During preliminary data sampling, the module also verifies data source rationality based on system-defined rules, preventing the access of invalid or erroneous data sources, thereby improving overall system robustness and data processing efficiency.

[0070] The data adaptation module 102 is connected to the automatic data source identification module and is responsible for automatically adapting and standardizing data based on the identified data source characteristics. This module uses built-in algorithms to standardize the formats and structures of different data sources, ensuring data compatibility and accessibility. Adaptation operations include, but are not limited to, data format conversion, field mapping, and structure standardization, eliminating format and structure differences between data sources and achieving seamless data integration.

[0071] For example, if the structure of a data source doesn't match the system's standard data model, the data adaptation module can automatically generate corresponding mapping rules to convert the incompatible data structure into a unified format that the system can recognize and process. Typical adaptation operations include aligning the table structure from a relational database with the system's standard object model, converting JSON-formatted data into a unified structured data format, or automatically matching the mappings of different field names.

[0072] This flexible adaptation mechanism enables the system to process data from a variety of sources and formats, significantly enhancing the system's compatibility and adaptability to heterogeneous data. The adaptation module eliminates the need for manual adjustments through automated processes, significantly improving data processing efficiency and reducing the incidence of errors caused by format mismatches. The adaptation module also customizes the adaptation logic based on the characteristics of the data source and business needs, ensuring the system can adapt to the complex data requirements of different application scenarios.

[0073] Furthermore, the data adaptation module seamlessly integrates with subsequent processing modules, ensuring that adapted data can smoothly enter subsequent data processing and storage processes. This module is designed to be highly flexible and scalable, enabling the addition of new adaptation rules or algorithms as system requirements evolve, supporting future changes in data processing requirements.

[0074] Through the collaborative work of the data source automatic identification module and the data adaptation module, the system can automatically complete the entire process from data source detection and identification to format adaptation and standardized processing, providing an efficient and stable basic guarantee for subsequent data analysis and management.

[0075] The intelligent data quality management module 103 is responsible for comprehensive quality management of data entering the system, ensuring that data undergoes thorough quality testing and correction at all stages before entering the system. This module incorporates a self-learning mechanism that uses statistical analysis and machine learning algorithms to predict and automatically correct data quality issues, thereby improving the overall data quality of the system.

[0076] The intelligent data quality management module primarily monitors data integrity, consistency, and anomaly detection. Integrity testing ensures there are no missing or null values ​​in the dataset, consistency testing ensures there are no conflicts or duplications between data, and anomaly detection identifies unusual data points that fall outside a reasonable range.

[0077] To achieve these detection capabilities, the module utilizes a variety of statistical analysis and machine learning techniques. Statistical analysis methods include descriptive statistics (such as mean, variance, and standard deviation), while machine learning algorithms include supervised and unsupervised learning. Supervised learning algorithms are trained on existing labeled data to identify potential quality issues. Unsupervised learning can be used to detect new data anomalies, such as using the K-means clustering algorithm to detect possible anomalous groups in a dataset.

[0078] The data quality management module uses these methods to continuously update its understanding of different data sets and generate corresponding quality detection rules. When new data arrives, the module quickly verifies it according to established rules and predicts potential quality issues. The system also continuously refines and updates its detection algorithms based on common quality issues in historical data, gradually improving its ability to predict data quality issues.

[0079] The intelligent data quality management module incorporates a self-learning mechanism that continuously learns from historical data to enhance the accuracy of quality management. Specifically, the system regularly collects processed data sets and analyzes common quality issues (such as missing values, format errors, and numerical anomalies). Based on this analysis, the system continuously optimizes its detection and correction strategies. For example, using decision trees or random forest algorithms, it can identify the most common features that cause data issues and re-adjust the parameters of the detection model based on these feature weights.

[0080] This self-learning mechanism can identify new types of issues based on historical data behavior patterns. For example, if a specific data source shows persistent anomalies in a specific field, the system can dynamically adjust detection rules to detect similar errors earlier. Furthermore, the system can automatically generate remediation plans. For example, if certain fields frequently contain missing values, the module can automatically recommend the most appropriate fill values ​​(such as mean fill or nearest neighbor interpolation) based on historical data, reducing manual intervention.

[0081] When the intelligent data quality management module detects data quality issues, the system collaborates with the data preprocessing module 104 to automatically trigger appropriate processing operations. For example, if the module detects missing values ​​in a dataset, the system invokes the corresponding operations in the data preprocessing module to fill the missing values. The system supports a variety of filling strategies, including common mean filling, mode filling, and interpolation methods (such as linear interpolation or polynomial interpolation). The specific strategy is automatically selected based on the data type and context.

[0082] Furthermore, for duplicate data detection and removal, the Intelligent Data Quality Management module compares different records within the same dataset to identify data redundancy. For example, by comparing string similarity algorithms, the module can effectively identify potential duplicate records and automatically remove them by collaborating with the Data Preprocessing module.

[0083] Abnormal data processing is automatically implemented through anomaly detection algorithms. The intelligent data quality management module uses anomaly detection algorithms such as those based on support vector machines (SVMs) or the isolation forest algorithm to identify outliers in the data. These data points are then flagged as potential issues and passed to the data preprocessing module for correction. This correction can include simple outlier discarding, as well as outlier replacement or interpolation based on business needs.

[0084] The intelligent data quality management module not only detects and corrects issues in existing data but also, through self-learning mechanisms, predicts potential future problems. For example, by analyzing common error types and patterns in historical data, the system can generate a data quality early warning model, providing quality warnings for data entering the system. For example, if a particular data source has frequently encountered null values ​​in the past, the system will automatically increase monitoring of that data source and even proactively correct potential data issues.

[0085] This prediction function effectively avoids the accumulation of erroneous data, prevents data quality issues from spreading in the system, and ensures the long-term stable and efficient operation of the system.

[0086] The Intelligent Data Quality Management module also features a report generation function that produces detailed data quality analysis reports. These reports include detailed statistics on data integrity, consistency, and anomaly detection, and display historical trends and resolution records for various data issues. This reporting feature not only provides system administrators with a comprehensive data quality monitoring tool but also provides a reliable basis for data analysis and business decision-making.

[0087] Through the self-learning mechanism and real-time data correction function of the intelligent data quality management module, this system can continuously optimize the data quality management process, achieve efficient quality detection, prediction and processing, and ensure the accuracy and consistency of all data in the system.

[0088] The data preprocessing module 104 is closely integrated with the intelligent data quality management module and is used to clean and convert raw data. Data preprocessing includes missing value handling, error correction, format standardization, and data fusion. The system performs preprocessing operations based on different rule sets (e.g., data cleaning rules, format conversion rules) to ensure that the data meets the requirements of subsequent storage and analysis.

[0089] For example, when processing data from multiple sources, the data preprocessing module can automatically perform encoding conversion operations based on the different data characteristics to ensure format compatibility across all data sources. In this way, the system can effectively eliminate redundant and abnormal information in the data, laying the foundation for subsequent data analysis.

[0090] The data integration module 105 is based on a containerized microservice architecture and uses a dynamic expansion architecture and a dynamic load balancing mechanism (such as Figure 2 This module uses ETL tools to extract and transform pre-processed data from source systems and load it into target data stores, enabling unified data management and analysis.

[0091] Automatically adjust data processing based on the system's real-time data load. When data load is low, processing steps are simplified to save resources; when data load is high, the system automatically activates additional ETL services to ensure data processing efficiency.

[0092] Dynamically start or shut down processing nodes based on system load, ensuring on-demand expansion of data processing capabilities. When dealing with large-scale data access and high-concurrency access, the load balancing mechanism can effectively allocate system resources to ensure smooth and efficient data processing.

[0093] The data storage module 106 efficiently and securely stores integrated and processed data, supporting various database systems, including relational databases, NoSQL databases, and cloud-based storage solutions. The core of this module lies in its ability to dynamically select the optimal storage solution based on data access frequency and importance, enabling separate storage of hot and cold data, thereby meeting the needs of efficient queries while minimizing storage costs.

[0094] For frequently accessed "hot data," such as data that requires real-time processing or frequent queries, the system prioritizes high-performance storage options. Common storage options include relational databases (such as MySQL and PostgreSQL), which offer strong transaction processing capabilities and efficient query performance, enabling rapid response to query requests. Furthermore, in scenarios requiring high concurrent access, the data storage module can also utilize distributed databases (such as Cassandra and CockroachDB) to ensure high data availability and throughput.

[0095] For less frequently accessed "cold data," such as historical data or infrequently updated data, the system utilizes low-cost storage options, such as NoSQL databases (e.g., MongoDB, Couchbase) or cloud storage services (e.g., Amazon S3, Azure Blob Storage). These storage options not only offer lower operating costs but also provide large-scale data storage capabilities through a distributed storage architecture, making them suitable for storing large datasets that require less frequent access. Furthermore, cloud storage solutions offer flexible scalability, dynamically adjusting capacity based on data storage needs.

[0096] The data storage module ensures the system's efficiency and cost-effectiveness when processing large amounts of data through categorized and hierarchical data storage. By integrating data access patterns, update frequency, and importance, the module automatically migrates data between different storage systems, ensuring a balance between timely data access and storage costs. For example, if certain data is frequently accessed over a period of time, it will be marked as "hot data" by the system and stored in a high-performance relational database. When access frequency decreases, the system will automatically migrate it to a lower-cost NoSQL database or cloud storage, reducing unnecessary high-performance storage usage.

[0097] The API management module 107 is used to develop and maintain the system's data access interfaces, supporting modern data interface standards such as RESTful APIs and GraphQL interfaces. This module enables the system to provide flexible and unified data access for different types of users, applications, or third-party services, ensuring data accessibility, scalability, and security.

[0098] The API Management module features powerful multi-version management capabilities, allowing developers to maintain multiple versions of the API simultaneously to meet the needs of different applications and user groups. This version management allows the system to push new features or interface updates without impacting existing services, ensuring users can choose the appropriate API version based on their needs, enabling a smooth transition and avoiding business interruptions.

[0099] The API Management module also features comprehensive authentication and authorization mechanisms to ensure data access security. For example, it supports standard authentication methods such as OAuth 2.0 and JWT (JSON Web Token), allowing the system to verify user identities and control data access based on user permissions. Fine-grained permission control ensures that users can only access data consistent with their permissions, preventing the leakage of sensitive data.

[0100] Furthermore, the API management module features real-time usage monitoring and compliance review. The system monitors the access frequency, response time, and usage patterns of each API request, generating detailed log reports based on user behavior. This information not only helps administrators identify potential abuse but also facilitates system optimization, adjusting data access policies, and improving API response efficiency. By monitoring API usage, the system can also dynamically adjust access permissions or restrictions based on individual user behavior, ensuring data access security and compliance.

[0101] The monitoring and optimization module 108 is responsible for real-time monitoring of the overall system operation and displays the system's operating status through a visual monitoring interface. This module monitors key performance indicators (KPIs) such as CPU and memory usage, data processing latency, network traffic, and database query response time. It also assesses the system's health in real time through system log and data traffic analysis.

[0102] The data processing latency analysis provided by the Monitoring and Optimization module helps system administrators identify performance bottlenecks in the system. For example, if the latency of a processing node in the data processing process continues to increase, the system can automatically allocate more resources based on the module's feedback, or launch new processing nodes in a dynamically scalable architecture, ensuring that the overall system processing capacity is not affected by local bottlenecks.

[0103] The monitoring and optimization module also features automatic optimization, dynamically adjusting system resource allocation based on real-time monitoring data to ensure optimal resource utilization. In the event of a surge in data traffic, the monitoring module can promptly detect and automatically initiate an expansion mechanism to add additional processing nodes to prevent system overload. When traffic returns to normal, the system automatically releases excess resources to reduce unnecessary resource consumption.

[0104] Furthermore, the monitoring and optimization module uses trend analysis to proactively predict potential system resource bottlenecks. For example, if data traffic or processing requests show an increasing trend over a certain period, the system can proactively scale resources to avoid performance degradation or response delays during peak periods. This predictive capability enhances system resilience, ensuring efficient and stable operation even in the face of high concurrency and large-scale data processing.

[0105] The specific implementation of the data integration module is as follows:

[0106] like Figure 2 As shown, data integration module 105, based on a containerized microservices architecture, efficiently processes pre-processed data from multiple data sources and integrates it into the target data storage system through extract, transform, and load (ETL) operations, enabling unified data management and analysis. This module achieves efficient and flexible data integration through a dynamically scalable architecture and dynamic load balancing mechanism, making it particularly suitable for the integrated processing of large-scale, multi-source, heterogeneous data.

[0107] The data integration module utilizes a containerized microservices architecture, enabling each ETL task or data processing service to be encapsulated as an independent microservice unit. This containerization enables independent deployment, operation, and scalability of modules, ensuring high isolation and flexibility. Each ETL task runs as an independent microservice in its own dedicated container environment, ensuring that different data processing tasks do not interfere with each other. Container orchestration tools (such as Kubernetes and Docker Swarm) can be used to dynamically scale the number of ETL services to accommodate varying data processing needs. Containerized services can be quickly started and shut down, making them suitable for data processing scenarios requiring rapid response times.

[0108] The dynamic expansion architecture of the data integration module can automatically adjust the data processing process according to the current data load of the system to ensure data processing efficiency and resource utilization. The specific implementation plan is as follows:

[0109] Data load monitoring: The system monitors data source input traffic and system load in real time. By analyzing data flow statistics and identifying historical data patterns, the system can predict future data load trends and plan resource scheduling in advance.

[0110] Low-load scenarios: When the system detects low data load, the dynamically scalable architecture automatically streamlines data processing. For example, the ETL service automatically adjusts the frequency of extraction and transformation, deferring some processing tasks to times of higher load to conserve resources. This may reduce non-essential data conversion operations, retaining only the core data integration processes and ensuring efficient resource utilization.

[0111] High-load scenarios: When system load increases (such as with large-scale data ingestion or high-concurrency access), the dynamic scaling architecture automatically launches additional ETL services to ensure system processing capacity matches data traffic. The system automatically increases the number of ETL container instances based on the current load to ensure smooth data processing during peak periods.

[0112] The key technology behind this dynamic scaling architecture is an automatic scheduling algorithm based on load thresholds. This algorithm dynamically expands or contracts system resources by setting upper and lower thresholds for system load (such as CPU usage, memory utilization, and data traffic).

[0113] The specific formula of the automatic scheduling algorithm based on multi-dimensional load is as follows:

[0114]

[0115] in:

[0116] The current total system load is an overall load index calculated by summing up the load ratios of various resources within the system. It is used to indicate the comprehensive utilization of the current system resources.

[0117] Under reasonable conditions, Usually in within a certain range, making comparison and monitoring easier.

[0118] Load ratios: The formula includes four load ratios, representing the utilization of CPU, memory, network bandwidth, and disk I / O:

[0119] : CPU load ratio, which indicates the ratio of the current CPU usage to the total capacity;

[0120] : Memory load ratio, which indicates the ratio of current memory usage to total capacity;

[0121] : The load ratio of network bandwidth, which indicates the ratio of the current usage of network bandwidth to the total bandwidth;

[0122] : Disk I / O load ratio, which indicates the ratio of current disk I / O usage to total I / O capacity.

[0123] These ratios quantify the utilization of each resource, ranging from 0 to 1. The closer the ratio is to 1, the more heavily loaded the resource is.

[0124] Weight coefficient ,The weight coefficients are used to reflect the relative importance of each resource in the ,overall load.,They flexibly adjust the system scheduling strategy by assigning different ,priorities to different resources.

[0125] For example, for memory-intensive tasks, increasing The value can increase the weight of the memory load ratio in the overall load calculation; for I / O intensive tasks, increase This emphasizes the impact of disk I / O load on system scheduling. This weight can be dynamically adjusted based on different task types.

[0126] Among them, when When , the system will expand the node, and the expansion number is calculated as follows:

[0127]

[0128] in: Number of nodes to be expanded: Indicates the number of nodes that need to be expanded. This number is calculated by the ratio of the system load exceeding the upper threshold and is calculated by the function Precise adjustments ensure that the system can quickly replenish resources and relieve pressure in overload situations.

[0129] Since the number of expanded nodes should be an integer, the ceiling function is used To ensure that the calculation results are consistent with the actual number of scalable nodes.

[0130] Overload Ratio: Indicates the relative percentage by which the current load exceeds the threshold. This quantifies the deviation of the system's current load from the acceptable upper limit, reflecting the current level of system resource insufficiency.

[0131] The calculation method is: use the current system total load Subtract the upper load threshold , then divided by , thus obtaining the overload ratio.

[0132] (Expanded ratio function): Function Used to adjust the expansion ratio so that the expansion process can allocate resources more accurately and efficiently.

[0133] It is a weight coefficient for different resources (such as CPU, memory, network bandwidth, disk I / O). Each coefficient has different priorities under different types of tasks. For example, the weight can be increased for I / O intensive tasks. , thereby giving priority to expanding I / O-related nodes.

[0134] Generates a scaling factor based on the system's load imbalance. Monitors whether the load on each resource exceeds its individual upper limit, allowing for more targeted scaling of the most needed resources.

[0135] The specific formula for the expansion adjustment coefficient is:

[0136]

[0137] in:

[0138] These are separate upper thresholds for CPU, memory, network bandwidth, and disk I / O, and their values ​​range from 0 to 1. They are set based on the system's performance requirements and load capacity.

[0139] The max function selects the maximum value between 0 and the difference; when the ratio of each resource exceeds the corresponding threshold, the expansion ratio function increases the contribution of the weight, otherwise it will be regarded as zero.

[0140] function Calculate the expansion ratio based on resource usage patterns to make expansion more accurate and effective.

[0141] The data integration module's dynamic load balancing mechanism ensures that data processing capacity can be scaled and allocated on demand based on load fluctuations. The system continuously monitors the load of each ETL node. When the load on a node exceeds a threshold, tasks are redistributed to less loaded nodes to ensure load balancing. When the load increases significantly, the system triggers an automatic scaling mechanism, launching new ETL nodes and assigning new tasks. When the load decreases, redundant nodes are automatically shut down to reduce resource waste. The ETL tool in the data integration module is highly adaptable, including adaptive data extraction, incremental data processing, and transformation load optimization.

[0142] Adjust the extraction strategy based on the real-time status of the data source and the output of the preprocessing module. Frequently updated data sources will have their extraction frequency increased; infrequently updated data sources will have their extraction frequency reduced through batch processing.

[0143] ETL tools only extract and convert new or updated data, avoiding duplicate processing and significantly improving processing efficiency.

[0144] During the data loading phase, the system optimizes the import speed through batch loading technology.

[0145] To further improve system efficiency, this application develops a load forecasting model based on time series analysis. The specific forecasting model formula is:

[0146]

[0147] in:

[0148] For predicted future loads;

[0149] is the average load, and are the autoregressive coefficient and the moving average coefficient, respectively.

[0150] (Future Time Forecasted Load): This is the load forecast for the next time t, used to estimate future system resource requirements. If the forecasted load exceeds a certain threshold, the system can expand resources in advance; conversely, if the load decreases, some resources can be released to optimize system costs.

[0151] (Load Average): is a constant that represents the average level of historical load data. This constant ensures that the forecast value is based on the overall average load of the system in the past, reducing errors caused by occasional fluctuations.

[0152] The specific value of can be obtained by calculating the mean of historical data or by fitting the model.

[0153] (Autoregressive coefficients): It is the coefficient of the autoregressive term, which indicates the weight of the impact of historical load values ​​on the current forecast and reflects the continuity of past load data in time.

[0154] The parameter p is the order of the autoregressive component, which determines how many past load values ​​are included in the prediction calculation. By optimally selecting p, the model can achieve a good balance between prediction accuracy and computational complexity.

[0155] (Moving average coefficient): It is the coefficient of the moving average (MA) term, which represents the correction amount of the previous error term to the forecast value, and is usually used to smooth short-term fluctuations caused by incidental events.

[0156] The parameter q is the order of the moving average, indicating how many past error terms are used in the current forecast calculation. By choosing a reasonable q value, the cumulative deviation of the error term can be effectively reduced.

[0157] (Error term): The error term represents the prediction error at time i in the past and is a white noise sequence, i.e., a random variable with zero mean. The error term is introduced to correct for random disturbances and make the model more robust. This model enables the system to predict load peaks in advance, expand or release resources in advance, and optimize processing capacity.

[0158] Represents the actual load value at the past i-th moment. It is the historical data used by the autoregressive part, that is, when predicting the load at the future time t, the model will refer to the actual load data of the past i time steps ; These historical load values ​​are used to capture the trends and patterns of system load changes over time, enabling the model to predict future load conditions by analyzing past load behavior.

[0159] The data integration module uses containerized architecture, dynamic expansion and load balancing mechanisms, combined with adaptive ETL tools and predictive models, to ensure efficient data integration and unified management under different loads.

[0160] The above description is only a preferred embodiment of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the scope of the patent application of the present invention.

Claims

1. An intelligent data access and integration system, characterized by: The intelligent data access and integration system includes: Data source automatic identification module, used to automatically detect and identify multiple data sources and their characteristics; a data adaptation module, connected to the data source automatic identification module, for adapting the data source according to the identified characteristics; Intelligent data quality management module, which uses a self-learning mechanism to predict and automatically correct data quality issues; Data preprocessing module, used to clean and transform raw data; The data integration module uses a containerized, dynamically scalable architecture to automatically adjust data processing based on real-time data load. It simplifies data processing steps when data load is low and starts additional ETL services when data load is high. It also uses a dynamic load balancing mechanism to start or shut down additional processing nodes based on data processing needs. Data storage module, used to store data and support multiple types of databases; API management module, used to develop and maintain data access interfaces; Monitoring and optimization module, used to monitor system performance and data processing flow; The dynamic expansion architecture automatically adjusts the data processing process according to the real-time data load, including: An overall load index is calculated based on the real-time data load status of each load in the system using an automatic scheduling algorithm, and an overload ratio is obtained by subtracting an upper load threshold from the overall load index and then dividing the result by the upper load threshold. The number of expansion nodes is calculated based on the overload ratio, the individual online threshold of each load, and the weight coefficient corresponding to each load; The loads include CPU, memory, network bandwidth and disk I / O; The intelligent data quality management module is further used to: Analyze the acquired historical data to determine error types and models, and generate a data quality warning model, which is used to issue quality warnings for data about to enter the system; The data storage module includes: a storage unit that supports multiple database types, including relational databases, NoSQL databases, and cloud-based storage systems; A dynamic selection unit is used to dynamically select a storage solution based on the access frequency and importance of the data to achieve separate storage of hot and cold data. The storage solution includes a NoSQL database and a cloud-based storage system for cold data, and a relational database for hot data. The data migration unit is used to automatically migrate data between different storage systems when the data access frequency changes.

2. The intelligent data access and integration system according to claim 1, characterized in that: The data source automatic identification module includes: a data collection unit for detecting and identifying multiple data sources and their characteristics, including data format, structure and access rights; an automated script or standardized interface for collecting and analyzing metadata information of the data source; a data source directory establishment unit for establishing a data source directory based on the collected metadata information; a data connection verification unit for verifying the accessibility of the data source connection; and a data sampling unit for sampling data from the data source.

3. The intelligent data access and integration system according to claim 1, characterized in that: The data adaptation module includes: an interface connected to the data source automatic identification module, used to receive data source characteristic information; a data format conversion unit, used to convert data formats of different data sources; The field mapping unit is used to generate field mapping rules based on the characteristics of the data source to achieve automatic matching of field correspondences; The structure standardization unit is used to standardize the structures of different data sources to ensure data compatibility and accessibility.

4. The intelligent data access and integration system according to claim 1, characterized in that: The intelligent data quality management module includes: a self-learning mechanism for optimizing data quality detection rules through historical data; a quality monitoring unit to monitor data integrity, consistency, and anomaly detection; A statistical analysis and machine learning unit for predicting and automatically correcting data quality issues, wherein the machine learning unit includes algorithms for missing value detection, duplicate data detection, and anomaly detection; A collaborative interface with the data preprocessing module, used to trigger data correction operations when data quality issues are detected.

5. The intelligent data access and integration system according to claim 1, characterized in that: The dynamic expansion architecture adopted by the data integration module automatically adjusts the data processing flow based on the automatic scheduling algorithm of the load threshold; the formula of the automatic scheduling algorithm is as follows: in: L total is the current total system load; CPU load ratio, indicating the current CPU usage C used With the total capacity C total The ratio of The memory load ratio indicates the current memory usage M used With total capacity M total The ratio of The load ratio of the network bandwidth, which indicates the current usage of the network bandwidth N used With the total bandwidth N total The ratio of The load ratio of disk I / O indicates the current usage of disk I / O. used Total I / O Capacity total The ratio of α, β, γ, δ are weight coefficients.

6. The intelligent data access and integration system according to claim 5, characterized in that: When L total >L upper When , the system will expand the nodes, and the number of expanded nodes is calculated as follows: Where: n new is the number of nodes that need to be expanded; L total is the current total system load; L upper is the upper load threshold; f(α,β,γ,δ) is the expansion ratio function.

7. The intelligent data access and integration system according to claim 6, characterized in that: The specific formula of the expansion ratio function is: in: T C 、T M 、T N 、T I Separate upper thresholds for CPU, memory, network bandwidth, and disk I / O.

8. The intelligent data access and integration system according to claim 1, characterized in that: The API management module includes: Supports data access unit with multiple interfaces; Multi-version management unit, used to support simultaneous maintenance of multiple API versions; Identity authentication and authorization unit, which performs user authentication and controls data access rights; Use monitoring and compliance review units to monitor the access frequency, response time, and usage patterns of API requests and generate log reports.

9. The intelligent data access and integration system according to claim 1, characterized in that: The monitoring and optimization module includes: a visual monitoring interface for displaying the system operation status; a key performance indicator monitoring unit for monitoring CPU utilization, memory utilization, data processing latency, network traffic, and database query response time; a system health assessment unit for real-time evaluation of system health through system log and data traffic analysis; an automatic optimization unit for dynamically adjusting system resource configuration based on real-time monitoring data; and a trend analysis unit for predicting system resource bottlenecks and expanding resources in advance.

Citation Information

Patent Citations

  • Load balancing method and device for server cluster and electronic equipment

    CN114461389A

  • Intelligent data cleaning system based on real-time database

    CN118885473A

  • Data analysis and governance integrated platform based on multi-dimensional data

    CN119025582A