Data processing method for automatic deployment of enterprise-level PaaS platform
By combining multi-source data collection and intelligent resource scheduling with machine learning and big data analytics, the problems of rigid data processing workflows and inflexible resource scheduling in the automated deployment of enterprise-level PaaS platforms have been solved, achieving efficient, stable deployment and continuous optimization.
Patent Information
- Application Number
- CN202511083681.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
In the existing automated deployment of enterprise-level PaaS platforms, data processing workflows are rigid, resource scheduling is inflexible, and there is a lack of effective monitoring and optimization, resulting in low deployment efficiency, poor stability, and resource waste.
By employing multi-source data acquisition, intelligent task allocation and resource scheduling, real-time monitoring and optimization mechanisms, combined with machine learning and big data analysis, we can achieve accurate data acquisition, dynamic resource allocation and real-time monitoring, ensuring a smooth and optimized deployment process.
It improved deployment efficiency, enhanced stability, optimized resource utilization, and enabled continuous optimization and efficient operation of the platform.
Smart Images

Figure CN120973384A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of enterprise-level platform deployment, in particular to a data processing method for automatic deployment of an enterprise-level PaaS platform. BACKGROUND
[0002] In the existing automatic deployment process of an enterprise-level PaaS platform, there are many problems: Firstly, the data processing flow is relatively rigid, and the collection and verification of the basic environment data before deployment is single, which is difficult to quickly adapt to diversified enterprise IT environments, resulting in low deployment efficiency. Secondly, in the deployment task allocation, there is a lack of dynamic and intelligent scheduling of resources, and a fixed resource allocation strategy is often used, which cannot be flexibly adjusted according to the actual business load and resource usage, and is likely to cause resource waste or resource shortage. Thirdly, in the deployment process, the data interaction and collaborative work between components lack effective monitoring and management mechanisms, and once a problem occurs, it is difficult to quickly locate and solve, affecting the stability and reliability of the deployment. Fourthly, the existing deployment data processing method lacks effective monitoring and optimization mechanisms for the resources and data after deployment, and cannot timely discover resource bottlenecks and data anomalies, making it difficult to achieve continuous optimization and efficient operation of the platform. SUMMARY
[0003] The purpose of the present application is to provide a data processing method for automatic deployment of an enterprise-level PaaS platform, to solve the above-mentioned problems in the prior art, improve the efficiency, stability and resource utilization of the automatic deployment of the PaaS platform, and realize the continuous optimization of the platform.
[0004] To achieve the above-mentioned purpose, the present application provides the following technical solution: comprising the following steps: Basic environment data collection and preprocessing: before the automatic deployment of the PaaS platform, multi-source data collection technology is used to collect basic environment data from various data sources such as enterprise network equipment, servers, storage systems, etc., including network topology information, server hardware configuration, storage capacity, etc. After the collection is completed, data cleaning algorithm is used to clean the collected data, remove duplicate, incorrect and incomplete data, and use data standardization technology to convert the data into a unified format for subsequent processing. At the same time, a data verification rule library is established to verify the cleaned and standardized data, to ensure the accuracy and integrity of the data. Intelligent task allocation and resource scheduling: Based on the collected basic environment data and the information of the PaaS platform components to be deployed, a resource demand prediction model is constructed. This model is trained based on historical deployment data and business load data using machine learning algorithms, which can accurately predict the amount of resources required for each deployment task. Then, combined with real-time resource usage, a dynamic resource scheduling algorithm is used to allocate deployment tasks to the most suitable computing resources. During the allocation process, the load balancing and priority of resources are fully considered, and the resource requirements of critical business components are prioritized. At the same time, a resource monitoring mechanism is established to monitor the usage status of resources in real time, and when the resource usage reaches a certain threshold, automatic resource expansion or contraction operations are triggered. Deployment process data interaction and collaborative management: During the automatic deployment process of the PaaS platform, a unique data interaction identifier is assigned to each component, and data interaction protocols and interface specifications between components are established. Through message queues and event-driven mechanisms, real-time data interaction and collaborative work between components are achieved. At the same time, a deployment monitoring platform is built to monitor the data transmission, task execution status, etc. in real time during the deployment process. Once abnormal conditions such as data transmission delay or task execution failure are found, the alarm mechanism is triggered immediately, and the problem is located and analyzed through intelligent diagnosis algorithms to provide corresponding solutions. Post-deployment resource and data monitoring optimization: After deployment, a resource monitoring system is established to collect real-time resource index data such as CPU usage, memory usage, and disk I / O of servers. Using big data analysis techniques and machine learning algorithms, the collected resource data is analyzed to predict resource usage trends and identify resource bottlenecks in advance. At the same time, real-time monitoring and analysis of the data generated by the platform are performed, and a data quality evaluation model is established to evaluate the accuracy, completeness, and consistency of the data. Based on the evaluation results, data optimization and adjustment operations such as data backup, archiving, and cleaning are automatically performed to ensure data quality and availability. In addition, based on the monitoring results of resources and data, dynamic adjustments and optimizations of the PaaS platform configuration can be made to improve platform performance and operational efficiency.
[0005] In summary, due to the use of the above-mentioned technologies, the beneficial effects of the present application are: 1. Improve deployment efficiency: Through multi-source data collection and intelligent data preprocessing, the enterprise's basic environment data can be quickly and accurately obtained, providing a reliable basis for deployment and avoiding deployment delays caused by data problems. At the same time, intelligent task allocation and resource scheduling mechanisms can dynamically allocate resources according to actual needs, improving resource utilization efficiency and thus speeding up deployment. 2. Enhance deployment stability: The deployment process data interaction and collaborative management mechanism ensures smooth data interaction and collaboration between components. Through real-time monitoring and intelligent diagnosis, problems that occur during deployment can be found and solved in a timely manner, improving the stability and reliability of deployment. 3. Optimize resource utilization: The post-deployment resource and data monitoring optimization mechanism can monitor resource usage in real time, predict resource bottlenecks in advance, and optimize data based on data quality assessment results, achieving dynamic adjustment and optimization of resources, improving resource utilization, and reducing enterprise operating costs.
[0006] 4. Realize the continuous optimization of the platform: The present application can continuously collect and analyze data through data processing and monitoring during the whole deployment process, providing data support for the optimization of the PaaS platform, realizing the continuous optimization and efficient operation of the platform, and meeting the changing business needs of enterprises. BRIEF DESCRIPTION OF DRAWINGS
[0007] The accompanying drawings, which form a part of this application, are used to provide a further understanding of the application, make the other features, purposes and advantages of the application more apparent. The illustrative embodiments of the drawings and their descriptions serve to explain the present application, and do not constitute an improper limitation on the present application. In the drawings: Figure 1 The flowchart of the present application. DETAILED DESCRIPTION
[0008] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme of the embodiments of the present application will be described clearly and completely below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but only to represent selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0009] In the description of the present application, it should be understood that the terms indicating position or location relationship are based on the position or location relationship shown in the drawings, only for the convenience of describing the present application and simplifying the description, and cannot be understood as limiting or implying that the indicated device or element must have a specific position, be constructed and operated in a specific position, therefore cannot be understood as limiting the present application.
[0010] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting", "fixing" and other terms should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrated; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances of the specification.
[0011] The present application provides a data processing method for automatic deployment of an enterprise-level PaaS platform, Embodiment one: basic environment data collection and preprocessing A large enterprise plans to deploy a new enterprise-level PaaS platform. Before deployment, the method of the present application is used for basic environment data collection and preprocessing. Using multi-source data collection technology, data is collected in parallel from 500 servers, 20 network switches and 10 storage arrays of the enterprise through a distributed collection architecture. The collected basic environment data includes the CPU model (such as Intel Xeon Platinum 8380) of the server, the core number (48 cores), the memory capacity (256 GB), the storage capacity (1 TB of local storage for each server, and a total storage capacity of 500 TB for the storage array), and the connection relationship, bandwidth (10 Gbps) and other information of each switch in the network topology structure.
[0012] After the collection is completed, the data is cleaned using an abnormal value detection algorithm based on statistical analysis and an error data correction algorithm based on rule matching. For example, through statistical analysis, it is found that the CPU usage rate of a certain server continuously exceeds 100%, which is determined as abnormal data and corrected; according to the preset rules, the records with incorrect storage capacity data format are uniformly corrected. Then, using data standardization technology, all collected data is converted into JSON format that meets the requirements of the platform. Finally, according to the established data verification rule library, the cleaned and standardized data is verified. The verification rules include checking whether the server IP address is within the specified network segment and whether the storage capacity value is a positive integer, etc., to ensure the accuracy and integrity of the data, and to provide a reliable data basis for subsequent deployment operations.
[0013] Embodiment two: intelligent task allocation and resource scheduling After completing the basic environment data collection and preprocessing, intelligent task allocation and resource scheduling are started. The PaaS platform to be deployed includes 10 components such as database service, application server and message queue. According to the collected basic environmental data and component information, a resource demand prediction model is constructed. Using the historical data of similar PaaS platform deployment of the enterprise in the past six months and the load data of the recent business system, a random forest algorithm is used to train the model. The trained model can accurately predict the resource amount required for each component deployment task, for example, predicting that the database service deployment requires 8-core CPU, 16GB memory and 100GB storage. Combined with real-time resource usage, a dynamic resource scheduling algorithm based on genetic algorithm is used to allocate deployment tasks to appropriate computing resources. The current enterprise data center has 100 idle servers, through algorithm calculation, the database service deployment task is allocated to 3 servers with low CPU load and sufficient memory and storage, and the resources are reasonably allocated to ensure load balancing of resources. At the same time, the resource needs of application servers and other key business components are prioritized. A resource monitoring mechanism is established, and CPU usage exceeding 80% and memory usage exceeding 90% are set as resource usage thresholds. When the CPU usage of a server reaches 85%, an automatic resource expansion operation is triggered, and 2-core CPU and 4GB memory are allocated from the standby resource pool to supplement the server, ensuring the smooth progress of the deployment task and efficient use of resources. Example Three: Deployment Process Data Interaction and Collaborative Management During the automatic deployment process of the PaaS platform, deployment process data interaction and collaborative management are implemented. Each component such as database service and application server is assigned a unique data interaction identifier, such as "DB-001" and "APP-002", and a data interaction interface specification based on HTTP / 2 protocol between components is established.
[0014] Through Kafka distributed message queue and event-driven mechanism, real-time data interaction and collaborative work between components are realized. For example, after the application server completes initialization, it sends an "initialization complete" event through the message queue, and the database service receives the event and starts establishing a connection with the application server and performing data synchronization. A deployment monitoring platform based on Prometheus and Grafana is built to monitor data transmission rate, task execution status and other information in real time during the deployment process. When the data transmission delay between the database service and the message queue exceeds 500ms, an alarm mechanism is triggered immediately to notify the operation and maintenance personnel through email and SMS. At the same time, using intelligent diagnostic algorithms based on fault tree analysis and case reasoning technology, the cause of the transmission delay is quickly located to be insufficient network bandwidth, and solutions such as increasing network bandwidth or optimizing data transmission protocol are provided to ensure the stable progress of the deployment process. Example Four: Post-deployment Resource and Data Monitoring Optimization After the PaaS platform is deployed, the post-deployment resource and data monitoring optimization process is started. A resource monitoring system is established, and the Zabbix tool is used to collect resource indicator data such as CPU usage, memory usage, and disk I / O of the server in real time, with a collection frequency of once per minute. Using big data analysis technology and the LSTM (Long Short-Term Memory) algorithm, the collected resource data is analyzed to predict the resource usage trend in the next 24 hours. For example, it is predicted that the memory usage of a server will reach 95% at 3 pm, an early warning is issued, and an automatic resource adjustment operation is triggered to release some cache data, avoiding service interruption due to insufficient memory. A data quality evaluation model is established, and the analytic hierarchy process is used to determine the weights of evaluation indicators such as data accuracy (weight 0.4), completeness (weight 0.3), and consistency (weight 0.3). The fuzzy comprehensive evaluation method is used to evaluate the data generated by the platform. When it is found that some order data in the database is missing, the data repair process is automatically started according to the evaluation results to recover the missing order data from the backup data and archive the data to ensure the quality and availability of the data. At the same time, according to the monitoring results of resources and data, the configuration of the PaaS platform is dynamically adjusted, such as adjusting the cache strategy of the database and optimizing the thread pool configuration of the application server, to improve the performance and efficiency of the platform.
Claims
1. A data processing method for enterprise-level PaaS platform automatic deployment, characterized in that: The method comprises the following steps: S1: basic environment data collection and preprocessing: using multi-source data collection technology, collecting network topology information, server hardware configuration, storage capacity and other basic environment data from enterprise network equipment, servers and storage systems; using data cleaning algorithm to clean the data, remove duplicate, incorrect and incomplete data, and convert the data to a unified format using data standardization technology; establishing a data verification rule library to verify the cleaned and standardized data; S2: intelligent task allocation and resource scheduling: based on the collected basic environment data and the information of the PaaS platform components to be deployed, a resource demand prediction model trained based on historical deployment data and business load data using machine learning algorithm is constructed to predict the resource amount required for each deployment task; combining real-time resource usage, using dynamic resource scheduling algorithm, the deployment task is allocated to appropriate computing resources, taking into account resource load balancing and priority, and a resource monitoring mechanism is established to automatically trigger resource expansion or contraction operation when the resource usage reaches a threshold; S3: deployment process data interaction and collaborative management: in the automatic deployment process of the PaaS platform, a unique data interaction identifier is assigned to each component, and a data interaction protocol and interface specification between components are established; Through message queue and event-driven mechanism, real-time data interaction and collaborative work between components are realized; a deployment monitoring platform is built to monitor data transmission and task execution status in real time during deployment, trigger alarm mechanism when abnormality is found, and provide solutions through intelligent diagnosis algorithm to locate and analyze problems; S4: post-deployment resource and data monitoring optimization: after deployment, a resource monitoring system is established to collect real-time server CPU usage, memory usage and disk resource index data; using big data analysis technology and machine learning algorithm, resource data is analyzed to predict resource usage trend; a data quality evaluation model is established to evaluate the accuracy, integrity and consistency of platform generated data, and according to the evaluation results, data backup, archiving, cleaning and other optimization adjustment operations are automatically performed, and according to the resource and data monitoring results, the PaaS platform configuration is dynamically adjusted and optimized.
2. The data processing method for enterprise-level PaaS platform automation deployment according to claim 1, characterized in that, The multi-source data collection technology adopts a distributed collection architecture to realize parallel collection of different types of data sources.
3. The data processing method for enterprise-class PaaS platform automation deployment according to claim 1, characterized in that, The data cleaning algorithm includes an outlier detection algorithm based on statistical analysis and an error data correction algorithm based on rule matching.
4. The data processing method for enterprise-class PaaS platform automation deployment according to claim 1, characterized in that, The dynamic resource scheduling algorithm uses genetic algorithm or particle swarm optimization algorithm to optimize resource utilization and task completion time.
5. The data processing method for enterprise-class PaaS platform automation deployment according to claim 1, characterized in that, The message queue uses a high-availability distributed message queue architecture to ensure the reliability and real-time performance of data interaction.
6. The data processing method for enterprise-class PaaS platform automation deployment according to claim 1, characterized in that, The intelligent diagnosis algorithm is based on fault tree analysis and case-based reasoning technology to quickly locate and analyze problems in the deployment process.
7. The data processing method for enterprise PaaS platform automation deployment according to claim 1, characterized in that, The data quality evaluation model uses analytic hierarchy process to determine the weight of data quality evaluation indicators, and uses fuzzy comprehensive evaluation method to evaluate data quality.