Enterprise big data analysis processing method and system based on cloud computing

By using cloud computing-based technologies such as multi-source heterogeneous data access and intelligent data preprocessing, the problems of resource waste and slow speed in traditional big data processing methods have been solved, dynamic resource scheduling and data quality assurance have been achieved, and the efficiency and security of enterprise big data processing have been improved.

CN120849489APending Publication Date: 2025-10-28NANJING YUNNUO NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510915306.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional stand-alone or small cluster processing methods are insufficient to cope with the surge in data volume, diversification, and high-speed processing needs of enterprise big data, resulting in resource waste, high costs, slow analysis speed, and difficulty in forming a unified data view and ensuring data security and compliance.

Method used

It adopts cloud computing-based multi-source heterogeneous data access, intelligent data preprocessing, adaptive hybrid storage, elastic resource scheduling, distributed parallel computing and multi-dimensional analysis, combined with modules such as unified access gateway, metadata management and security management, to achieve dynamic resource scheduling and data quality assurance.

Benefits of technology

It enables automatic scaling when the load changes, reduces hardware investment and storage costs, improves data processing speed and efficiency, breaks down data silos, provides a unified data view, and ensures data security and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849489A_ABST
    Figure CN120849489A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise big data analysis and processing method and system based on cloud computing, and belongs to the technical field of big data process.The enterprise big data analysis and processing method and system based on cloud computing comprises the following specific steps that firstly, multi-source heterogeneous data access and metadata perception are conducted through a configured connector; data are accessed in real time in batches from a plurality of data sources inside and outside an enterprise, the accessed data sources are automatically scanned, data structure, format, mode and data quality preliminary information metadata are extracted and registered to a unified metadata management library, and sensitive data fields are identified. Through dynamic resource scheduling based on cloud computing, the system can automatically stretch and retract according to the load, explosive growth of the data size and the computing requirement is easily coped with, huge hardware investment in the early stage is not needed, and the computing and storage cost is greatly reduced through an intelligent mixed instance strategy and a self-adaptive storage strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of big data processing technology, specifically relating to a cloud computing-based enterprise big data analysis and processing method and system. Background Technology

[0002] Modern enterprises face challenges such as a surge in data volume, diversification of data sources and formats, high demands for data processing speed, low data value density, and the need to ensure data authenticity. Traditional stand-alone or small cluster processing methods are insufficient to address these challenges.

[0003] Traditional technical solutions rely on fixed hardware resources, making it difficult to handle peak business demand and growth. Expansion is costly and time-consuming, resource allocation is static, resulting in wasted resources during off-peak hours and insufficient resources during peak hours. When dealing with petabyte / electrode (PB / EB) level data, traditional architectures are slow and cannot meet real-time or near-real-time analysis needs. Hardware maintenance, software deployment, and cluster management also impose heavy operational burdens, requiring specialized teams. Furthermore, data from different business systems is scattered, making it difficult to integrate and analyze it to form a unified view. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a cloud computing-based enterprise big data analysis and processing method and system.

[0005] The technical solution adopted to solve the above technical problem includes the following specific steps: Step 1: Multi-source heterogeneous data access and metadata awareness. Through configurable connectors, data is accessed in batches and in real time from multiple data sources inside and outside the enterprise. The accessed data sources are automatically scanned, and preliminary information on data structure, format, mode, and data quality metadata is extracted and registered in a unified metadata management library to identify sensitive data fields. Step 2: Intelligent data preprocessing and quality enhancement. Based on predefined rules and machine learning models, data cleaning is automatically performed, data quality monitoring and feedback are implemented, low-quality data is marked and alarms and repair processes are triggered, and sensitive data is dynamically desensitized and encrypted. Step 3: Adaptive hybrid storage strategy, which intelligently decides the data storage hierarchy based on data access frequency, popularity, analysis needs, and cost requirements; Step 4: Provide visualization and declarative tools to support users in defining business logic and data analysis goals, and automatically optimize data models; Step 5: Elastic resource scheduling and computing optimization, real-time monitoring and analysis of task queue load, data volume, and SLA requirements; Step Six: Distributed Parallel Computing Execution. Utilize the distributed computing framework provided by the cloud computing platform to execute computing tasks in parallel. Step 7: Multidimensional analysis and insight generation. Execute predefined multidimensional analysis queries to generate visual reports and business insight results. Step 8: Result delivery and feedback loop: Push the analysis results to downstream business systems and mobile application development to continuously optimize data models, algorithm parameters and resource scheduling strategies.

[0006] Through the above technical solutions, the system can automatically scale according to the load by using cloud computing-based dynamic resource scheduling, easily coping with the explosive growth of data volume and computing needs, without the need for huge upfront hardware investment. The intelligent hybrid instance strategy and adaptive storage strategy significantly reduce computing and storage costs.

[0007] Furthermore, step four, which involves automatically optimizing the data model, includes automatically generating an optimized data model based on the analysis query pattern and business logic, and automatically executing the ETL / ELT process.

[0008] Furthermore, the elastic resource scheduling in step five includes dynamically scaling the computing cluster size based on predictive models and real-time monitoring, allocating computing tasks to the most suitable resource pool, considering data locality, intelligently combining on-demand instances, reserved instances, and spot instances to execute tasks, maximizing cost-effectiveness, and a critical task guarantee mechanism to ensure that core tasks are not interrupted by spot instances.

[0009] Through the above technical solutions, the distributed parallel computing framework, combined with optimized resource scheduling and data locality considerations, as well as the optimization of computing engine parameters, significantly improves the speed of massive data processing and the execution efficiency of complex analysis tasks.

[0010] Furthermore, it includes a unified access gateway module, a metadata management module, an intelligent preprocessing engine, an adaptive storage manager, a data modeling and optimization center, an elastic resource scheduler, a distributed computing execution cluster, an analysis computing engine, a visualization platform, an integration and feedback API, a security management center, and a system management platform.

[0011] The above technical solutions enable convenient API-based push notifications to business systems, allowing feedback to be collected and forming a closed loop. This truly integrates data analysis into business processes, driving business decisions and optimization.

[0012] Furthermore, the unified access gateway module includes a connector management submodule, which supports configuring and expanding new data source connectors, is responsible for connecting to various data sources, and provides security authentication, protocol conversion, and data buffering functions.

[0013] The above technical solutions effectively break down data silos through unified access gateways and metadata management, providing a unified data view and clear data lineage. Intelligent preprocessing ensures data quality. A security management center guarantees data security and compliance.

[0014] Furthermore, the metadata management module is responsible for storing and managing the metadata of all access data sources. The intelligent preprocessing engine includes a data cleaning submodule, a data quality monitoring submodule, and a data desensitization and encryption submodule. The adaptive storage manager includes a storage strategy engine that automatically or by rules determines the data storage location and level based on data characteristics. The data modeling and optimization center is responsible for providing visual modeling tools and declarative interfaces.

[0015] The above technical solution allows for preprocessing of the data, facilitating subsequent operations.

[0016] Furthermore, the elastic resource scheduler includes a monitoring agent, a prediction engine, a scheduling strategy engine, and a cost optimizer, which dynamically decides on cluster scaling, task assignment, and instance type selection strategies based on load, prediction, SLA, and cost targets.

[0017] The above technical solutions can greatly improve resource utilization efficiency and avoid waste.

[0018] Furthermore, the distributed computing execution cluster is a computing resource pool managed by the cloud platform, running mainstream distributed computing frameworks. The analysis computing engine is responsible for providing machine learning libraries, the visualization and insight platform is responsible for displaying analysis results, the integration and feedback API is responsible for receiving business feedback, the security management center is responsible for ensuring the security and compliance of the entire system, and the system management and monitoring console is responsible for providing a graphical interface.

[0019] The above technical solutions can intuitively display data to users. Based on the cloud platform, it eliminates the burden of hardware procurement and data center operation and maintenance, and the system management console provides a convenient operation and maintenance monitoring interface.

[0020] The beneficial effects of this invention are as follows: This invention enables the system to automatically scale according to load through cloud-based dynamic resource scheduling, easily coping with explosive growth in data volume and computing demands without requiring huge upfront hardware investments. Intelligent hybrid instance strategies and adaptive storage strategies significantly reduce computing and storage costs. The distributed parallel computing framework, combined with optimized resource scheduling and data locality considerations, as well as computing engine parameter tuning, significantly improves the speed of massive data processing and the execution efficiency of complex analysis tasks. A unified access gateway and metadata management effectively break down data silos, providing a unified data view and clear data lineage. Intelligent preprocessing ensures data quality. A security management center guarantees data security and compliance. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0023] like Figure 1 As shown in this embodiment, a cloud computing-based enterprise big data analysis and processing method and system includes the following specific steps: Step 1: Multi-source heterogeneous data access and metadata awareness. Through configurable connectors, data is accessed in batches and in real time from multiple data sources inside and outside the enterprise. The accessed data sources are automatically scanned, and preliminary information on data structure, format, mode, and data quality metadata is extracted and registered in a unified metadata management library to identify sensitive data fields. Step 2: Intelligent data preprocessing and quality enhancement. Based on predefined rules and machine learning models, data cleaning is automatically performed, data quality monitoring and feedback are implemented, low-quality data is marked and alarms and repair processes are triggered, and sensitive data is dynamically desensitized and encrypted. Step 3: Adaptive hybrid storage strategy, which intelligently decides the data storage hierarchy based on data access frequency, popularity, analysis needs, and cost requirements; Step 4: Provide visualization and declarative tools to support users in defining business logic and data analysis goals, and automatically optimize data models; Step 5: Elastic resource scheduling and computing optimization, real-time monitoring and analysis of task queue load, data volume, and SLA requirements; Step Six: Distributed Parallel Computing Execution. Utilize the distributed computing framework provided by the cloud computing platform to execute computing tasks in parallel. Step 7: Multidimensional analysis and insight generation. Execute predefined multidimensional analysis queries to generate visual reports and business insight results. Step 8: Result delivery and feedback loop. Push the analysis results to downstream business systems and mobile application development to continuously optimize data models, algorithm parameters and resource scheduling strategies. Through cloud-based dynamic resource scheduling, the system can automatically scale according to the load, easily cope with the explosive growth of data volume and computing needs, without the need for huge upfront hardware investment. Intelligent hybrid instance strategy and adaptive storage strategy greatly reduce computing and storage costs.

[0024] Step four, which involves automatically optimizing the data model, includes automatically generating an optimized data model based on the analysis query pattern and business logic, and automatically executing the ETL / ELT process.

[0025] The elastic resource scheduling in step five includes dynamically scaling the computing cluster size based on predictive models and real-time monitoring, allocating computing tasks to the most suitable resource pool, considering data locality, and intelligently combining on-demand instances, reserved instances, and spot instances to execute tasks to maximize cost-effectiveness. The critical task guarantee mechanism ensures that core tasks are not interrupted by spot instances. The distributed parallel computing framework, combined with optimized resource scheduling and data locality considerations, as well as computing engine parameter tuning, significantly improves the speed of massive data processing and the execution efficiency of complex analysis tasks.

[0026] It includes a unified access gateway module, a metadata management module, an intelligent preprocessing engine, an adaptive storage manager, a data modeling and optimization center, an elastic resource scheduler, a distributed computing execution cluster, an analysis computing engine, a visualization platform, an integration and feedback API, a security management center, and a system management platform. It can conveniently push data to business systems through APIs and collect feedback to form a closed loop, so that data analysis can be truly integrated into business processes and drive business decisions and optimization.

[0027] The unified access gateway module includes a connector management submodule, which supports configuring and expanding new data source connectors, is responsible for connecting to various data sources, and provides security authentication, protocol conversion, and data buffering functions. The unified access gateway and metadata management effectively break down data silos, providing a unified data view and clear data lineage. Intelligent preprocessing ensures data quality. A security management center guarantees data security and compliance.

[0028] The metadata management module is responsible for storing and managing the metadata of all access data sources. The intelligent preprocessing engine includes a data cleaning submodule, a data quality monitoring submodule, and a data desensitization and encryption submodule. The adaptive storage manager includes a storage strategy engine that automatically or by rules decides the data storage location and level based on data characteristics. The data modeling and optimization center is responsible for providing visual modeling tools and declarative interfaces to preprocess the data for easier subsequent operations.

[0029] The elastic resource scheduler includes a monitoring agent, a prediction engine, a scheduling strategy engine, and a cost optimizer. Based on load, prediction, SLA, and cost targets, it dynamically decides on cluster scaling, task assignment, and instance type selection strategies, which can greatly improve resource utilization efficiency and avoid waste.

[0030] The distributed computing execution cluster is a computing resource pool managed by the cloud platform, running mainstream distributed computing frameworks. The analysis computing engine is responsible for providing machine learning libraries, the visualization and insight platform is responsible for displaying analysis results, the integration and feedback API is responsible for receiving business feedback, the security management center is responsible for ensuring the security and compliance of the entire system, and the system management and monitoring console is responsible for providing a graphical interface that can intuitively display data to users. Based on the cloud platform, it eliminates the burden of hardware procurement and data center operation and maintenance, and the system management console provides a convenient operation and maintenance monitoring interface.

[0031] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.

Claims

1. A cloud computing-based enterprise big data analysis and processing method, characterized in that... This includes the following specific steps: Step 1: Multi-source heterogeneous data access and metadata awareness. Through configurable connectors, data is accessed in batches and in real time from multiple data sources inside and outside the enterprise. The accessed data sources are automatically scanned, and preliminary information on data structure, format, mode, and data quality metadata is extracted and registered in a unified metadata management library to identify sensitive data fields. Step 2: Intelligent data preprocessing and quality enhancement. Based on predefined rules and machine learning models, data cleaning is automatically performed, data quality monitoring and feedback are implemented, low-quality data is marked and alarms and repair processes are triggered, and sensitive data is dynamically desensitized and encrypted. Step 3: Adaptive hybrid storage strategy, which intelligently decides the data storage hierarchy based on data access frequency, popularity, analysis needs, and cost requirements; Step 4: Provide visualization and declarative tools to support users in defining business logic and data analysis goals, and automatically optimize data models; Step 5: Elastic resource scheduling and computing optimization, real-time monitoring and analysis of task queue load, data volume, and SLA requirements; Step Six: Distributed Parallel Computing Execution. Utilize the distributed computing framework provided by the cloud computing platform to execute computing tasks in parallel. Step 7: Multidimensional analysis and insight generation. Execute predefined multidimensional analysis queries to generate visual reports and business insight results. Step 8: Result delivery and feedback loop: Push the analysis results to downstream business systems and mobile application development to continuously optimize data models, algorithm parameters and resource scheduling strategies.

2. The enterprise big data analysis and processing method based on cloud computing according to claim 1, characterized in that, Step four, which involves automatically optimizing the data model, includes automatically generating an optimized data model based on the analysis query pattern and business logic, and automatically executing the ETL / ELT process.

3. The enterprise big data analysis and processing method based on cloud computing according to claim 2, characterized in that, The elastic resource scheduling in step five includes dynamically scaling the computing cluster size based on predictive models and real-time monitoring, allocating computing tasks to the most suitable resource pool, considering data locality, and intelligently combining on-demand instances, reserved instances, and spot instances to execute tasks to maximize cost-effectiveness. The critical task guarantee mechanism ensures that core tasks are not interrupted or affected by spot instances.

4. The system for enterprise big data analysis and processing based on cloud computing according to claim 3, characterized in that, It includes a unified access gateway module, a metadata management module, an intelligent preprocessing engine, an adaptive storage manager, a data modeling and optimization center, an elastic resource scheduler, a distributed computing execution cluster, an analysis and computing engine, a visualization platform, an integration and feedback API, a security management center, and a system management platform.

5. The system for enterprise big data analysis and processing based on cloud computing according to claim 4, characterized in that, The unified access gateway module includes a connector management submodule, which supports configuring and expanding new data source connectors, is responsible for connecting to various data sources, and provides security authentication, protocol conversion, and data buffering functions.

6. The system for enterprise big data analysis and processing based on cloud computing according to claim 5, characterized in that, The metadata management module is responsible for storing and managing the metadata of all access data sources. The intelligent preprocessing engine includes a data cleaning submodule, a data quality monitoring submodule, and a data desensitization and encryption submodule. The adaptive storage manager includes a storage strategy engine that automatically or by rules decides the data storage location and level based on data characteristics. The data modeling and optimization center is responsible for providing visual modeling tools and declarative interfaces.

7. The system for enterprise big data analysis and processing based on cloud computing according to claim 6, characterized in that, The elastic resource scheduler includes a monitoring agent, a prediction engine, a scheduling strategy engine, and a cost optimizer. Based on load, prediction, SLA, and cost targets, it dynamically decides on cluster scaling, task assignment, and instance type selection strategies.

8. The system for enterprise big data analysis and processing based on cloud computing according to claim 7, characterized in that, The distributed computing execution cluster is a computing resource pool managed by the cloud platform, running mainstream distributed computing frameworks. The analysis computing engine is responsible for providing machine learning libraries, the visualization and insight platform is responsible for displaying analysis results, the integration and feedback API is responsible for receiving business feedback, the security management center is responsible for ensuring the security and compliance of the entire system, and the system management and monitoring console is responsible for providing a graphical interface.