Information technology analysis system and method based on cloud computing

By adopting virtualization, containerization and intelligent scheduling technologies in the cloud-based information technology analysis system, combined with distributed file systems and data lakes, the problems of low resource utilization efficiency and limited data analysis capabilities are solved, efficient resource utilization and in-depth data analysis are realized, and scientific basis is provided for decision-making.

CN120091019AInactive Publication Date: 2025-06-03BAODING MAOXING INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510323070.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cloud-based information technology analysis system has problems such as low resource utilization efficiency and limited data analysis capabilities, resulting in performance bottlenecks and resource waste.

Method used

Virtualization and containerization technology are adopted, combined with intelligent scheduling algorithms, to achieve dynamic allocation and efficient utilization of resources. Build a distributed file system and data lake to achieve high-reliability data storage and comprehensive data management. Establish a professional algorithm team, develop and optimize data analysis algorithms, and use distributed computing to accelerate training and perform multi-source data fusion and in-depth mining.

Benefits of technology

It realizes flexible resource utilization, high-reliability data storage and in-depth data analysis, ensures key application performance, avoids resource waste, and provides scientific decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091019A_ABST
    Figure CN120091019A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and discloses an information technology analysis system and method based on cloud computing, and the method comprises the steps that a network building module caches static data to nodes nearby a user through a CDN; the resource scheduling module monitors and dynamically allocates resources in real time; the data storage module realizes data redundancy storage; the data acquisition module performs multi-source data acquisition; the data integration module integrates data of different data sources; the data management module manages the whole data process; the data analysis module carries out multi-source data fusion analysis and mines a potential relationship; the result display module provides an intuitive and easy-to-use interface; the security management module guarantees data storage and transmission security; and the monitoring management module deploys a monitoring tool to monitor system operation in real time. According to the method, dynamic allocation and efficient utilization of resources are achieved, deep data analysis is achieved, training is accelerated through distributed computing, automatic model updating is achieved, multi-source data fusion and deep mining are conducted, and a scientific basis is provided for decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and more specifically, to an information technology analysis system and method based on cloud computing. Background Art

[0002] Cloud computing provides nearly unlimited computing resources. When an information technology analysis system processes massive amounts of data, such as the large number of transaction data processed by financial institutions every day, and user behavior data of e-commerce platforms, etc., traditional local computing resources may be overwhelmed due to the large amount of data. Based on cloud computing, the system can quickly call a large number of computing resources and quickly complete complex data mining and analysis algorithms, greatly improving the analysis efficiency, enabling enterprises to obtain valuable information from the data faster and make decisions in a timely manner. Enterprises do not need to invest huge amounts of money to build and maintain a large-scale local data center. The purchase and maintenance of servers, storage devices, etc. not only have high upfront procurement costs, but also continuous subsequent operation and maintenance, upgrade and other expenses. By adopting an information technology analysis system based on cloud computing, enterprises only need to rent cloud services on demand and pay according to the actual amount of resources used, which greatly reduces the construction cost and operation and maintenance cost of hardware facilities. Especially for small and medium-sized enterprises, this low-cost model enables them to carry out advanced information technology analysis work with a relatively low threshold.

[0003] The prior art document with the publication number CN118170845A provides an information technology analysis system based on cloud computing. The analysis system includes a data collection module, a data storage module, a data processing and analysis module, a visualization module, a security and privacy module, an interaction and user interface module, an elastic expansion and management module. The data collection module collects data from various data sources, including databases, APIs, sensors, etc., and uploads it to the cloud platform for storage and processing. The data storage module stores the data collected from the data collection module. The data processing and analysis module processes and analyzes the data stored in the cloud database. The visualization module displays the analyzed data to users in a visual way. This invention can provide users with efficient, safe, and flexible data analysis and application services, meet the needs of different user groups, and provide strong support for business decision-making and innovation of enterprises.

[0004] Although the above prior art solution can achieve relevant beneficial effects through the structure of the prior art, there are still the following defects: 1. The resource utilization efficiency is low, and it is difficult to flexibly divide and efficiently utilize physical server resources. Resource allocation may be static and unreasonable, resulting in performance bottlenecks and resource waste. 2. The data analysis ability is limited, and the comprehensive analysis and in-depth data mining ability are insufficient, making it difficult to mine the potential value in the data.

[0005] In view of this, we propose an information technology analysis system and method based on cloud computing. Summary of the Invention

[0006] 1. Technical Problem to be Solved

[0007] The purpose of this application is to provide an information technology analysis system and method based on cloud computing, which solves the technical problems proposed in the above background technology, realizes flexible resource utilization, utilizes virtualization and containerization technologies, combines intelligent scheduling algorithms, realizes dynamic allocation and efficient utilization of resources, ensures the performance of critical applications, and avoids resource waste; realizes highly reliable data storage; realizes comprehensive data management, constructs a relational database warehouse, implements full-life cycle management, formulates backup and recovery strategies, optimizes storage resources, and ensures data security and efficient utilization; realizes in-depth data analysis, uses distributed computing to accelerate training, realizes automatic model update, conducts multi-source data fusion and in-depth mining, and provides a scientific basis for decision-making.

[0008] 2. Technical Solution

[0009] The technical solution of this application provides an information technology analysis system based on cloud computing, including:

[0010] Network construction module: Build a high-speed network, use 10 Gigabit Ethernet technology to construct the core network, deploy software-defined network (SDN) technology, realize intelligent scheduling and optimization of network traffic, dynamically allocate network bandwidth, and reduce network latency. Use a content delivery network (CDN) to cache common static data (such as pictures, videos, analysis report templates, etc.) to the node closest to the user.

[0011] Resource scheduling module: Use virtualization technologies such as KVM to divide physical servers into multiple virtual machines to realize flexible allocation of computing resources. Introduce Docker containerization technology to package application programs and their dependencies into independent containers, which is convenient for rapid deployment and migration, and improves resource utilization and application isolation.

[0012] Conduct intelligent resource scheduling, construct a resource scheduling algorithm based on machine learning, and real-time monitor the resource usage (CPU, memory, disk I / O, etc.) of each virtual machine and container in the system and the business requirements of users. According to this information, dynamically allocate computing resources.

[0013] Data storage module: Use distributed file systems such as Ceph to disperse data storage on multiple storage nodes to realize redundant storage and high availability of data. Construct a data lake for storing various original format data, including structured, semi-structured and unstructured data. Use technologies such as Apache Hudi to realize efficient management of the data lake, support real-time writing, updating and querying of data, and provide rich data sources for data analysis.

[0014] Data Acquisition Module: Develop a general data acquisition interface to support data collection from various data sources. Use ETL tools to implement data extraction, transformation, and loading, converting data in different formats into a unified format for subsequent processing.

[0015] During the data acquisition process, use data quality monitoring tools to continuously monitor the accuracy, integrity, and consistency of data. By setting data quality rules and thresholds, verify the collected data. For data that does not meet quality requirements, mark and process it in a timely manner, such as data cleaning, repair, or discard.

[0016] Data Integration Module: Establish unified data standards, including data formats, encoding methods, field definitions, data dictionaries, etc. Build a data standard management platform to centrally manage and maintain data standards, ensuring that data from all data sources follows the unified standards.

[0017] Use data integration tools to integrate data from different data sources. During the integration process, clean the data by removing duplicate data, correcting incorrect data, and filling in missing data. Apply data mapping and transformation rules to convert data from different data sources into a unified format and structure, and store it in a data warehouse or data lake.

[0018] Data Management Module: Build a data warehouse based on relational databases (such as MySQL, Oracle) to store structured data that has been cleaned, transformed, and integrated. Design the architecture of the data warehouse using the star model or snowflake model to improve the efficiency of data query and analysis.

[0019] Establish a full data life cycle management platform to achieve the whole process management of data from generation, collection, storage, use to archiving and destruction. Develop data backup and recovery strategies, and regularly back up data to ensure data security and recoverability. At the same time, classify and store data according to its usage frequency and importance, and archive infrequently used data to low-cost storage media to improve the utilization rate of storage resources.

[0020] Data analysis module: Establish a professional algorithm team to develop and optimize data analysis algorithms and models for different business scenarios and data characteristics. Combine artificial intelligence, machine learning, deep learning and other technologies to continuously improve the performance and accuracy of the algorithm. Use distributed computing frameworks (such as Apache Spark, TensorFlow, etc.) to achieve distributed model training of large-scale data. Distribute the training data to multiple computing nodes for parallel processing, which greatly shortens the model training time. At the same time, use model parallelism and data parallelism technologies to further improve training efficiency. Establish a model automatic update mechanism to monitor data changes and model performance indicators in real time. Conduct multi-source data fusion analysis, and use multi-source data fusion analysis technology to fuse and analyze data from different data sources. Use data association, feature extraction, pattern recognition and other technologies to explore the potential relationship and value between different data sources. Use advanced technologies such as deep learning and graph computing to conduct in-depth mining of large-scale data.

[0021] Results display module: Integrates professional visualization analysis tools (such as Tableau, Power BI, etc.) to provide users with an intuitive and easy-to-use data analysis interface. Users can quickly generate various data visualization reports and charts through operations such as dragging and clicking, making it easier for users to understand and analyze data.

[0022] Security management module: During data storage and transmission, SSL / TLS encryption protocols are used to encrypt data to ensure data security. Encryption algorithms are used to encrypt and store sensitive data to prevent data leakage. A complete access control system is established, and multi-factor authentication, role-based access control and other technologies are used to strictly limit user access rights to data and system resources.

[0023] Monitoring and management module: deploy comprehensive system monitoring tools to monitor the system's operating status in real time. By setting monitoring indicators and thresholds, timely discover system failures and performance bottlenecks, and send alarm information. Establish a fault management process to respond to and handle system failures in a timely manner. When a system failure occurs, use automated fault diagnosis tools to quickly locate the cause of the failure and take appropriate measures to repair it. At the same time, record the time, cause, and handling process of the failure for subsequent analysis and improvement.

[0024] The present invention provides an information technology analysis method based on cloud computing, comprising the following steps:

[0025] S1. The network construction module uses 10 Gigabit Ethernet to build a high-speed core network, combines SDN technology to intelligently schedule traffic, and uses CDN to cache static data to nodes near users, thereby increasing transmission speed and reducing pressure on data centers.

[0026] S2, the resource scheduling module uses virtualization technologies such as KVM to divide physical servers into virtual machines, and combines Docker containerization technology to improve resource utilization and application isolation. Through intelligent scheduling algorithms based on machine learning, real-time monitoring and dynamic allocation of resources ensure that high-priority and real-time applications obtain sufficient resources to avoid performance bottlenecks.

[0027] S3, data storage module uses distributed file systems such as Ceph to achieve data redundant storage and high availability, automatically repair and balance data. Build a data lake to store data in multiple formats, use technologies such as Apache Hudi to efficiently manage data, support real-time operations, and provide rich data sources for analysis.

[0028] S4. The data acquisition module supports multi-source data acquisition, develops a universal interface, and uses ETL tools to convert data formats to achieve unified processing. At the same time, data quality monitoring tools are used to monitor data quality in real time, mark and process unqualified data, and ensure data accuracy, completeness, and consistency.

[0029] S5. The data integration module formulates and manages unified data standards, uses data integration tools to integrate data from different data sources, performs data cleansing, converts data formats and structures, and stores them in data warehouses or data lakes.

[0030] S6. The data management module builds a relational database data warehouse and adopts efficient model design; establishes a data life cycle management platform to manage the entire data process; formulates backup and recovery strategies to ensure data security; implements hierarchical storage to optimize storage resource utilization.

[0031] S7. Data analysis module: Establish a professional algorithm team, develop and optimize data analysis algorithms and models, establish an algorithm library and centrally manage it. Use a distributed computing framework to implement large-scale data training and establish an automatic model update mechanism. Conduct multi-source data fusion analysis to explore potential relationships. Deep data mining uses advanced technology to conduct in-depth analysis of large-scale data to provide strong support for decision-making.

[0032] S8. The result display module integrates visual analysis tools, provides an intuitive and easy-to-use interface, and supports drag-and-drop generation of reports and charts. Develop customized data analysis applications to meet specific business needs. Achieve seamless integration with the company's existing business systems and data sharing and collaboration.

[0033] S9. The security management module uses SSL / TLS and encryption algorithms to ensure the security of data storage and transmission, establishes an access control system to limit user permissions, formulates a data privacy policy to protect user privacy, and ensures that data is used legally and in compliance with regulations.

[0034] S10. The monitoring and management module deploys monitoring tools to monitor the system operation in real time, establishes a fault management process for quick response and handling, formulates operation and maintenance plans and specifications, and uses automated tools to improve operation and maintenance efficiency.

[0035] 3. Beneficial Effects

[0036] One or more technical solutions provided in the technical solution of this application have at least the following technical effects or advantages:

[0037] 1. The present invention can achieve flexible resource utilization. By using virtualization and containerization technologies and combining intelligent scheduling algorithms, it realizes the dynamic allocation and efficient utilization of resources, ensures the performance of critical applications, and avoids resource waste.

[0038] 2. Achieve highly reliable data storage. Use a distributed file system to achieve data redundancy and high availability, automatic repair and balancing, build a data lake to support multi-source data storage and efficient management, and provide a solid foundation for analysis.

[0039] 3. Achieve comprehensive data management. Build a relational database warehouse, implement full life cycle management, formulate backup and recovery strategies, optimize storage resources, and ensure data security and efficient utilization.

[0040] 4. Achieve in-depth data analysis. Form a professional team to research and develop optimization algorithm models, use distributed computing to accelerate training, realize automatic model update, conduct multi-source data fusion and in-depth mining, and provide a scientific basis for decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of an information technology analysis method based on cloud computing disclosed in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following further describes this application in detail with reference to the accompanying drawings of the specification.

[0043] Refer to Figure 1 , the embodiment of this application provides an information technology analysis system based on cloud computing, including:

[0044] Network construction module: Build a high-speed network. Use 10 Gigabit Ethernet technology to build the core network to ensure high-speed data transmission inside the data center and between the data center and external users. At the same time, deploy software-defined network (SDN) technology to realize intelligent scheduling and optimization of network traffic, dynamically allocate network bandwidth, and reduce network latency. Use a content delivery network (CDN) to cache common static data (such as pictures, videos, analysis report templates, etc.) to the nodes closest to users. When users request this data, they directly obtain it from the CDN nodes, which greatly improves the data transmission speed and reduces the network pressure on the data center.

[0045] Resource Scheduling Module: Using virtualization technologies such as KVM (Kernel-based Virtual Machine), physical servers are divided into multiple virtual machines to achieve flexible allocation of computing resources. At the same time, Docker containerization technology is introduced to package application programs and their dependencies into independent containers, facilitating rapid deployment and migration, and improving resource utilization and application isolation.

[0046] Perform intelligent resource scheduling, build a resource scheduling algorithm based on machine learning, and real-time monitor the resource usage (CPU, memory, disk I / O, etc.) of each virtual machine and container in the system as well as the business requirements of users. Based on this information, dynamically allocate computing resources to ensure that high-priority tasks and applications with high real-time requirements can obtain sufficient resources and avoid performance bottlenecks caused by uneven resource allocation.

[0047] Data Storage Module: Adopt a distributed file system such as Ceph to disperse data storage across multiple storage nodes to achieve redundant data storage and high availability. Ceph has automatic repair and data balancing functions, which can automatically recover data to other normal nodes when some nodes fail, ensuring data integrity.

[0048] Build a data lake for storing various raw data formats, including structured, semi-structured, and unstructured data. Adopt technologies such as Apache Hudi to achieve efficient management of the data lake, support real-time writing, updating, and querying of data, and provide rich data sources for data analysis.

[0049] Data Collection Module: Conduct multi-source data collection, develop a general data collection interface, and support data collection from various data sources (such as databases, file systems, sensors, log files, etc.). Use ETL tools (such as Apache Sqoop, Kettle, etc.) to achieve data extraction, transformation, and loading, and convert data in different formats into a unified format for subsequent processing.

[0050] During the data collection process, use data quality monitoring tools (such as Informatica Data Quality, etc.) to real-time monitor the accuracy, integrity, and consistency of data. By setting data quality rules and thresholds, verify the collected data, and for data that does not meet the quality requirements, mark and process it in a timely manner, such as data cleaning, repair, or discard.

[0051] Data Integration Module: Formulate unified data standards, including data formats, encoding methods, field definitions, data dictionaries, etc. Establish a data standard management platform to centrally manage and maintain data standards to ensure that data from all data sources follows unified standards.

[0052] Use data integration tools (such as Talend, etc.) to integrate data from different data sources. During the integration process, clean the data by removing duplicate data, correcting incorrect data, and filling in missing data. Adopt data mapping and transformation rules to convert the data from different data sources into a unified format and structure, and store it in a data warehouse or a data lake.

[0053] Data management module: Build a data warehouse based on relational databases (such as MySQL, Oracle) to store the structured data that has been cleaned, transformed, and integrated. Design the architecture of the data warehouse using the star model or the snowflake model to improve the efficiency of data query and analysis.

[0054] Establish a full life cycle management platform for data to achieve the whole process management of data from generation, collection, storage, use to archiving and destruction. Develop data backup and recovery strategies, and regularly back up the data to ensure the security and recoverability of the data. At the same time, classify the storage of data according to the usage frequency and importance of the data, and archive the infrequently used data to low-cost storage media to improve the utilization rate of storage resources.

[0055] Data analysis module: Form a professional algorithm team, and develop and optimize data analysis algorithms and models according to different business scenarios and data characteristics. Combine technologies such as artificial intelligence, machine learning, and deep learning to continuously improve the performance and accuracy of the algorithms, such as developing deep learning-based image recognition algorithms and machine learning-based prediction models. Establish an algorithm library to centrally manage and maintain various algorithms. The algorithm library should have functions such as version management, algorithm description, and parameter setting to facilitate users to query and use. At the same time, evaluate and monitor the performance of the algorithms, and update and optimize the algorithms in a timely manner to ensure the effectiveness and adaptability of the algorithms.

[0056] Use distributed computing frameworks (such as Apache Spark, TensorFlow, etc.) to achieve distributed model training for large-scale data. Disperse the training data to multiple computing nodes for parallel processing, which greatly shortens the model training time. At the same time, adopt model parallelism and data parallelism technologies to further improve the training efficiency. Establish an automated model update mechanism to monitor the changes in data and the performance indicators of the model in real time. When significant changes occur in the data or the model performance deteriorates, automatically trigger the model update process. Use technologies such as incremental learning and online learning to incrementally update the model to avoid the high cost and high time consumption brought by retraining the entire model.

[0057] Conduct multi-source data fusion analysis. Through multi-source data fusion analysis technology, fuse and analyze data from different data sources (such as structured data, unstructured data, image data, video data, etc.). Adopt technologies such as data association, feature extraction, and pattern recognition to explore the potential relationships and values between different data sources, providing more comprehensive and in-depth analysis results for decision-making.

[0058] Conduct in-depth data mining. Utilize advanced technologies such as deep learning and graph computing to conduct in-depth mining of large-scale data. For example, adopt deep learning algorithms to conduct sentiment analysis and topic model mining on text data; utilize graph computing algorithms to analyze social network data and knowledge graph data to discover the relationships and patterns between nodes, providing strong support for business decision-making.

[0059] Result display module: Integrate professional visualization analysis tools (such as Tableau, PowerBI, etc.) to provide users with an intuitive and easy-to-use data analysis interface. Users can quickly generate various data visualization reports and charts, such as bar charts, line charts, pie charts, maps, etc., through operations such as dragging and clicking, facilitating users to understand and analyze data. Develop customized data analysis applications according to the needs of different industries and users. For example, develop risk assessment applications for the financial industry, disease prediction applications for the medical industry, and user behavior analysis applications for the e-commerce industry, etc. Through customized applications, meet the specific business needs of users and improve the pertinence and practicality of data analysis.

[0060] Develop unified interfaces and adapters to achieve seamless integration of this system with the enterprise's existing business systems (such as ERP, CRM, OA, etc.). Through data sharing and business process collaboration, timely feedback the data analysis results to the existing systems to provide support for enterprise decision-making. Ensure that the system can operate normally on different operating systems (such as Windows, Linux, MacOS), browsers (such as Chrome, Firefox, Safari, etc.) and mobile devices (such as mobile phones, tablets). Adopt responsive design and cross-platform development technologies to optimize the user experience of the system and improve the compatibility and accessibility of the system.

[0061] Security management module: During the data storage and transmission process, use the SSL / TLS encryption protocol to encrypt the data to ensure data security. At the same time, use encryption algorithms (such as AES, RSA, etc.) to encrypt and store sensitive data (such as user personal information, financial data, etc.) to prevent data leakage.

[0062] Establish a sound access control system, adopt technologies such as multi-factor authentication (such as username / password + SMS verification code + fingerprint recognition), role-based access control (RBAC), etc., and strictly limit users' access rights to data and system resources. Only authorized users can access specific data and functions to ensure the confidentiality and integrity of data. Develop a strict data privacy policy to clarify the rules for data collection, use, storage, and sharing. Adopt technologies such as data anonymization and differential privacy to process user data and protect user privacy. When sharing data and providing services externally, ensure that the use of data complies with the requirements of relevant laws, regulations, and privacy policies.

[0063] Monitoring and management module: Deploy comprehensive system monitoring tools (such as Zabbix, Prometheus, etc.) to monitor the running status of the system in real time, including server performance (CPU, memory, disk I / O, network bandwidth, etc.), application status, database performance, etc. By setting monitoring metrics and thresholds, promptly detect system failures and performance bottlenecks and send alarm messages.

[0064] Establish a fault management process to respond to and handle system failures in a timely manner. When a system failure occurs, use an automated fault diagnosis tool to quickly locate the cause of the failure and take corresponding measures for repair. At the same time, record the time, cause, and handling process of the failure for subsequent analysis and improvement.

[0065] Develop a system operation and maintenance plan and specifications, and regularly maintain and upgrade the system, including software updates, hardware maintenance, data backup, etc. Adopt automated operation and maintenance tools (such as Ansible, SaltStack, etc.) to automate the execution of operation and maintenance tasks and improve the efficiency and quality of operation and maintenance.

[0066] Furthermore, the resource scheduling module realizes the flexible allocation of computing resources, including the following steps:

[0067] 1. Preliminary preparation: Conduct a comprehensive inspection of the physical server hardware configuration, including parameters such as CPU model, number of cores, memory size, disk type and capacity, network bandwidth, etc. Based on the expected business load and resource requirements, plan the number of physical servers and the role each server will play in the virtualized environment. Install an operating system that supports KVM virtualization, such as mainstream Linux distributions (CentOS, Ubuntu Server, etc.), on the physical server, and ensure that the system kernel version meets the running requirements of KVM. Install the Docker runtime environment, which can be conveniently installed through the official software source or script, to ensure good compatibility with the operating system and the applications to be deployed later. Design the network architecture in the virtualized environment, divide network subnets for different purposes, such as a management network for administrators to manage virtual machines and containers, and a business network for applications to provide services externally and communicate with each other. Configure the physical server network interface to ensure its effective connectivity with the internal and external networks, and set appropriate firewall rules to ensure network security without affecting resource scheduling and application operation.

[0068] 2. Virtualization and containerization operations: Load the KVM kernel module in the server operating system, and ensure the normal operation of the module through command-line tools (such as modprobe kvm and modprobe kvm-intel or modprobe kvm-amd, depending on the CPU type). Install virtualization management software packages such as qemu-kvm. These tools provide command-line or graphical interfaces for creating and managing virtual machines. Use commands such as virt-install to create virtual machines. During the creation process, specify parameters such as the virtual machine name, allocated number of CPU cores, memory size, disk space, and network configuration. Write a Dockerfile to define the running environment of the application and its dependencies. In the Dockerfile, specify the base image (such as official Python, Java images, etc.), install the libraries, frameworks, and other dependency packages required by the application through commands, copy the application code to the specified directory in the image, and set the command to be executed when the container starts.

[0069] In the directory containing the Dockerfile, build the image. Run the container, and parameters such as the container name, mapped port, and mounted data volume can be specified.

[0070] 3. Build a resource scheduling algorithm model:

[0071] 3.1. Data collection and collation: Install resource monitoring tools such as collectd and sysstat in virtual machines and containers. These tools can collect data such as CPU usage, memory usage, and disk I / O read and write speeds. For example, in a Linux-based virtual machine or container, after installing sysstat through the package manager, the sar command can be used to obtain the system resource usage. Use a time series database such as InfluxDB to store resource monitoring data, and use a data collection agent such as Telegraf to send the data collected by the monitoring tools to InfluxDB. Configure Telegraf to connect to the monitoring tool interfaces of each virtual machine and container, collect data regularly, and store it in InfluxDB in the specified format.

[0072] Conduct user business requirement analysis, communicate with the business department to understand the business requirements of different applications and tasks, including task priorities, expected running durations, resource peak demands, etc. Quantify these requirements and store them in the database, establishing an association with the resource monitoring data.

[0073] 3.2. Machine learning model training: Select appropriate machine learning algorithms according to resource scheduling requirements and data characteristics, such as the Q-learning algorithm in reinforcement learning, deep learning models based on neural networks, etc. For example, for dynamic resource allocation scenarios, the Q-learning algorithm can learn the optimal resource allocation strategy through continuous trial and error. The resource scheduling model is:

[0074] Q(s,a)←Q(s,a)+A[w*r(1-β)+γmax a’[[Q(s’, a’)-Q(s, a)]]; where Q(s, a) is the Q-value function, which represents the discounted sum of the expected cumulative rewards obtained after taking action a in state s. It is a key metric in reinforcement learning for evaluating the goodness of taking a certain action in a specific state. In the resource scheduling scenario, state s can be understood as the current resource usage status of virtual machines and containers, and action a is various resource allocation operations, such as increasing the number of CPU cores for a certain virtual machine or allocating more memory for a container. A is the learning rate, with a value range between 0 and 1. It determines the degree of update of the original Q-value by new information. γ is the discount factor, also with a value range between 0 and 1. It reflects the degree of emphasis on future rewards. r is the immediate reward, which is the reward obtained immediately after executing action a. In the resource scheduling scenario, it is the feedback on the direct result of the current resource allocation decision. For example, if sufficient resources are successfully allocated to a high-priority task, enabling it to execute smoothly, then this action may receive a positive immediate reward; conversely, if the resource allocation is unreasonable, resulting in task failure or delay, a negative immediate reward may be obtained. w is the task urgency weight, set according to the urgency of the task. β is the resource volatility coefficient, with a value range between 0 and 1, which can be obtained by calculating the standard deviation of resource usage over a period of time, etc. s' is the new state, which is the new state of the system after executing action a. In resource scheduling, it is the new resource usage status of virtual machines and containers after the resource allocation operation is completed. max a' Q(s', a') represents the maximum Q-value that can be obtained among all possible actions a' in the new state s'. It enables the agent to comprehensively consider the current reward and future potential rewards when updating the Q-value of the current state-action, that is, to consider which action can obtain the maximum benefit in the new state, so as to optimize the current decision.

[0075] Extract the historical resource usage data and corresponding business requirement data from InfluxDB, clean and preprocess the data to remove outliers and noisy data. Divide the data into a training set, a validation set, and a test set according to a certain ratio for model training and evaluation.

[0076] Use deep learning frameworks such as TensorFlow and PyTorch in Python or reinforcement learning toolkits such as OpenAI Gym to build a model training environment. During the training process, continuously adjust the model parameters and optimize the model performance to achieve the purpose of accurately predicting resource requirements and reasonably allocating resources. Evaluate the generalization ability of the model through the validation set to prevent overfitting.

[0077] Model Evaluation and Optimization: Use the test set to evaluate the trained model and calculate the model's performance on metrics such as resource allocation accuracy and task completion efficiency. According to the evaluation results, further optimize the model, such as adjusting the network structure, increasing the amount of training data, improving algorithm parameters, etc., until the model performance meets the expected requirements.

[0078] 4. Actual Scheduling Implementation:

[0079] 4.1 Real-time Resource Monitoring and Data Update: Continuously collect real-time resource usage data of virtual machines and containers through Telegraf and update the data to InfluxDB in real time. At the same time, if the business department has new business requirements or task priority adjustments, update them in a timely manner to the records related to resource scheduling in the database.

[0080] 4.2 Execution of Resource Scheduling Decision: Input the real-time resource usage data and business requirement data into the trained machine learning resource scheduling model. The model generates resource allocation decisions according to the learned strategies. The scheduling system makes dynamic adjustments to the resources of virtual machines and containers according to the decisions, using virtualization management tools (such as using the virsh command to manage KVM virtual machine resources) and Docker commands (such as using docker update to adjust container resource limits). For example, when the model determines that a high-priority real-time task has insufficient resources, the scheduling system can use the virsh setvcpus command to increase the number of CPU cores for the corresponding virtual machine or use the docker update -memory command to allocate more memory to the container.

[0081] 4.3 Effect Feedback and Continuous Optimization: After the resource scheduling is executed, continuously monitor the running status of the application and the resource usage efficiency, and collect feedback data such as task completion status and response time. Compare these feedback data with the expected goals. If it is found that the resource scheduling effect is not good, re-evaluate the model performance, and if necessary, re-train the model or adjust the model parameters to continuously optimize the resource scheduling strategy to adapt to the changing business requirements and system operating environment.

[0082] Furthermore, the data integration module uses data integration tools to integrate data from different data sources, including the following steps:

[0083] 1. Select a data integration tool: Evaluate the mainstream data integration tools in the market, such as Talend, Informatica, Apache NiFi, etc., based on factors such as data source type, data volume, data processing complexity, and budget. Select Talend, which has rich data source connectors, supports various data format conversions, and provides a visual development interface, facilitating developers to design and debug data integration processes. Ensure that its version is compatible with other technical components used in the project, and purchase corresponding licenses according to the project scale and data processing volume.

[0084] 2. Connect to data sources: Use the selected data integration tool (such as Talend) to configure connections to each data source. For relational database data sources, provide connection information such as database address, port, username, and password; for file data sources, specify the file path and format; for web service data sources, set the API address and authentication information, etc. Through the tool's test connection function, ensure that the connections to each data source are normal. According to business requirements, formulate a data extraction strategy. Adopt full extraction (extracting all data from the data source at once) or incremental extraction (only extracting data that has changed since the last extraction). For example, for the customer basic information table with a low change frequency, full extraction can be adopted; for the transaction record table, due to the rapid growth of data volume, incremental extraction is adopted, and by recording the timestamp of the last extraction, only new transaction records are extracted.

[0085] 3. Data cleaning: Include removing duplicate data, correcting incorrect data, and filling in missing data;

[0086] Removing duplicate data: After data extraction, use the deduplication function of the data integration tool to identify and remove duplicate data. Deduplication can be based on a single field (such as a unique identifier field) or a combination of multiple fields.

[0087] Correcting incorrect data: Formulate data quality rules and check and correct data according to the rules.

[0088] Filling in missing data: Identify missing values in the data. For missing values in numeric fields, statistical methods such as mean, median, and mode can be used for filling; for missing values in character fields, if a default value is set, the default value can be used for filling, and if there is no default value, it can be inferred and filled according to business logic or marked as unknown.

[0089] 4. Data Mapping and Transformation: Define mapping rules. Based on the established data standards, define mapping rules for data fields from different data sources to map the fields in the data source to the corresponding fields in the target data format. For example, map the "Customer Name" field in Data Source A to the "customer_name" field in the target data format. For fields with inconsistent data types, also define data type conversion rules, such as converting the "Age" field with string type in the data source to integer type in the target format.

[0090] Execute data transformation. Use the transformation function of the data integration tool to transform the data according to the mapping rules. For example, in Talend, configure mapping components and transformation functions to transform the data in the data source into a format and structure that conforms to the unified standard. During the transformation process, conduct quality checks on the transformed data to ensure the accuracy of the transformation results. Perform data transformation according to the following formula:

[0091] Q 综合 = w 1 * P map + w 2 * P conv + w 3 * C; w 1 + w 2 + w 3 * C = 1; In the formula, Q 综合 is the comprehensive index of data transformation quality, used to comprehensively evaluate the overall quality of data during the conversion from the data source format to the target format. It is a quantitative value that comprehensively considers multiple key factors, and the theoretical value range is between 0 and 1. P map is the mapping accuracy rate, which measures the accuracy of the mapping from the data source fields to the target data format fields. The calculation method is the percentage of the number of correctly mapped fields in the total number of mapped fields. P conv is the conversion success rate. For operations such as data type conversion, it reflects the proportion of successful conversion processes. It measures the reliability of data during format or type conversion. A high P conv value means higher stability and accuracy of the data type conversion operation. C is the degree of consistency between the transformed data and the target format, and the value range is 0 - 1. It is obtained by comparing the compliance of the format of the transformed data (such as whether the date format meets the requirements of the target format, whether the order of data fields is correct, etc.), data range (such as whether the numerical data is within the reasonable range specified by the target format), etc. with the target format standard. w 1 、w 2 、w 3 are the weights of each index. They determine the weights of P map 、P conv and C in the calculation of Q 综合The proportion it occupies at a certain value. The value range is between (0 - 1). By reasonably adjusting the weight, the importance of different factors to data conversion quality can be flexibly emphasized according to different business requirements and data characteristics.

[0092] 5. Data storage: Evaluate whether to use a data warehouse or a data lake for data storage based on the characteristics of the data, usage scenarios, and business requirements. A data warehouse is suitable for storing structured data that has been cleaned, transformed, and integrated, and is used to support enterprise decision-making analysis; a data lake can store data in various formats (structured, semi-structured, unstructured), and is more suitable for data exploration and complex analysis scenarios. For example, if an enterprise mainly conducts standardized report analysis, a data warehouse can be selected; if it needs to conduct in-depth mining and exploratory analysis on multi-source heterogeneous data, a data lake may be more appropriate.

[0093] In the case of determining to use a data warehouse, select a suitable data warehouse architecture. If a data lake is selected, determine the storage technology to be used, such as a data lake based on the Hadoop Distributed File System (HDFS), or an object storage data lake based on cloud storage (such as AWS S3, Alibaba Cloud OSS, etc.). Develop a data loading strategy, including the data loading method (such as batch loading, real-time loading) and loading frequency. For batch loading, determine a suitable batch size according to the data volume and system performance; for real-time loading, adopt stream processing technologies (such as Apache Flink, Kafka Streams, etc.) to ensure that data can be stored in the target storage system in a timely manner. For example, for transaction data, due to high real-time requirements, adopt real-time loading, receive transaction data through the Kafka message queue, and then use Apache Flink for real-time processing and storage; for daily business statistical data, adopt batch loading and load data during the low business period at night. Use data integration tools to store the cleaned and transformed data into the data warehouse or data lake according to the selected storage scheme and loading strategy.

[0094] Furthermore, the data analysis module conducts multi-source data fusion analysis and in-depth data mining, including the following steps:

[0095] 1. Build an algorithm model: Deeply understand the requirements and pain points of different business scenarios. At the same time, conduct a comprehensive investigation of the existing data, including data types (structured, unstructured, images, videos, etc.), data volume, data quality, data distribution, etc. For example, in the e-commerce business scenario, understand the requirements for user purchase behavior analysis and investigate data such as user transaction records, browsing logs, and product information.

[0096] According to the business scenario and data characteristics, select the appropriate algorithm direction and design the model. For image recognition business, if the data volume is large and the accuracy requirement is extremely high, select the convolutional neural network (CNN) algorithm based on deep learning; for the scenario of predicting user churn, adopt machine learning algorithms such as logistic regression and random forest. During the algorithm design process, fully consider the scalability, computational complexity, and hardware resource requirements of the algorithm. Use programming languages such as Python and R, combined with deep learning frameworks such as TensorFlow and PyTorch or machine learning libraries such as Scikit-learn to implement the selected algorithm. During the implementation process, follow the code specifications, conduct modular design, and improve the readability and maintainability of the code. After completing the code writing, use test data to debug the algorithm, check whether the algorithm runs as expected, and troubleshoot and fix the errors and exceptions that occur.

[0097] 2. Algorithm Library Construction and Management: Design the architecture of the algorithm library, adopting a hierarchical structure, including an algorithm interface layer, an algorithm implementation layer, and a data storage layer. The algorithm interface layer provides a unified call interface for users, facilitating users to query and use algorithms; the algorithm implementation layer stores the specific codes of various algorithms; the data storage layer is used to store algorithm-related configuration information, training data, model parameters, etc. For example, in the algorithm interface layer, design a unified function call format, and users can call the algorithm for calculation by simply passing in the corresponding parameters.

[0098] Put the algorithms that have been developed and tested into the library according to the established architecture specifications. Assign a unique identifier to each algorithm and record the version information of the algorithm. When the algorithm is updated, use a version control tool (such as Git) to record the change history of the algorithm, which is convenient for tracing and management. At the same time, establish an algorithm version release process to ensure that the new version of the algorithm is officially released only after sufficient testing and verification.

[0099] Write a detailed description document for each algorithm in the algorithm library, including the principle of the algorithm, applicable scenarios, input and output parameter descriptions, performance characteristics, etc. Set a parameter configuration interface in the algorithm library to facilitate users to adjust algorithm parameters according to actual needs. For example, for a clustering algorithm, users can set parameters such as the number of clusters and distance measurement methods in the parameter configuration interface.

[0100] 3. Distributed Model Training:

[0101] 3.1. Framework Evaluation and Selection: Evaluate mainstream distributed computing frameworks such as Apache Spark and TensorFlow based on factors such as the data scale, computing resources, and algorithm types of the project. Apache Spark is suitable for batch processing and stream processing tasks of large-scale data, with efficient in-memory computing capabilities and rich data analysis libraries; TensorFlow, on the other hand, performs excellently in the distributed training of deep learning models, supporting model parallelism and data parallelism.

[0102] 3.2. Framework Installation and Configuration: Install and configure according to the selected distributed computing framework. During the installation process, ensure that the version of the framework is compatible with other technical components used in the project. For Apache Spark, parameters such as cluster node information, memory allocation, and disk storage need to be configured; for TensorFlow, parameters such as the server address, port, and GPU resource allocation for distributed training need to be set. Ensure the correct installation and configuration by testing the basic functions of the framework.

[0103] 3.3. Implementation of Data Parallelism and Model Parallelism: Includes the implementation of data parallelism and the implementation of model parallelism;

[0104] Implementation of Data Parallelism: Split the training data into multiple subsets and distribute them to different computing nodes for parallel processing. On each computing node, use the same model to train their respective data subsets, and then synchronously update the model parameters on the computing nodes. For example, in the distributed training based on TensorFlow, use the tf.distribute.Strategy API to implement data parallelism and synchronize model parameters among multiple GPUs through MirroredStrategy.

[0105] Implementation of Model Parallelism: For complex deep learning models, allocate different parts of the model (such as different neural network layers) to different computing nodes for parallel computing. Ensure data transfer and synchronization between different parts of the model by designing a reasonable communication mechanism. For example, in a multi-layer neural network model, allocate the first few layers to one computing node and the last few layers to another computing node, and implement data transmission between layers through network communication.

[0106] 3.4. Model Training and Optimization: Preprocess the original training data, including operations such as data cleaning, data standardization, and data augmentation (for image data). Split the preprocessed data according to the requirements of data parallelism and store it in a distributed file system (such as HDFS) so that computing nodes can quickly read the data. During the model training process, monitor indicators such as training progress, loss function values, and accuracy in real time. Use visualization tools (such as TensorBoard) to display the changes in indicators during the training process, facilitating the algorithm team to timely understand the training status of the model. When problems such as overfitting, underfitting, or training stagnation occur during training, promptly adjust algorithm parameters, model structures, or training strategies. Adopt optimization algorithms (such as Stochastic Gradient Descent, Adagrad, Adadelta, etc.) to update model parameters, improving the convergence speed and performance of the model. At the same time, perform operations such as model pruning and quantization on the model to reduce the complexity and storage space of the model and improve the running efficiency of the model in practical applications. For example, use L1 and L2 regularization methods to prevent model overfitting and remove unnecessary neurons and connections through model pruning.

[0107] 4. Model Automatic Update:

[0108] 4.1. Data and Model Monitoring:

[0109] Data Change Monitoring: Establish a data monitoring system to monitor the data changes in the data source in real time. By regularly collecting the feature statistics of the data (such as mean, variance, data distribution, etc.) and comparing them with the historical data features, determine whether the data has changed significantly. For example, for user behavior data, monitor the changes in statistics such as user access frequency and browsing duration. The data change monitoring model is:

[0110] D = Σ n i=1 {w i * |c i - h i | * [1 + Σ N i=1 (R ij * τ i )] + O i}; Σ n i=1 (w i ) = 1; where D is the data change difference index, used to comprehensively measure the changes between the set of current collected data feature statistics and the set of historical data feature statistics. n is the number of data feature statistics. In actual data monitoring scenarios, multiple different data features may be involved. w iis the weight of the i-th feature statistic, with a value range between 0 and 1. The setting of the weight reflects the relative importance of each data feature in the overall data change monitoring. c i is the value of the i-th data feature statistic collected currently. It represents the actual observed value of this data feature at the current moment. h i is the value of the i-th data feature statistic in historical data. This is the statistical value of this data feature within a certain past time period, serving as a comparison benchmark for measuring the change of the current data feature. R is the data feature correlation coefficient matrix, and its element R ij (i, j = 1, 2,..., n)) represents the correlation coefficient between the i-th feature and the j-th feature, with a value range between -1 and 1. τ i is the data change trend factor, used to measure the change trend of the i-th data feature over time, with a value range between 0 and 1. This value is obtained through time series analysis methods (such as linear regression to fit the change of the data feature over time). O i is the outlier correction term. When the data feature value c i is detected as an outlier (which can be detected through statistical methods such as box plots, etc.), O i is a value adjusted according to the degree of abnormality; if c i is not an outlier, then O i = 0. Outliers may be caused by data collection errors, special events, etc., and will have a greater interference on the calculation of the data change difference degree. The existence of O i can correct this kind of interference, making the D value more accurately reflect the real change situation of the data.

[0111] Model performance monitoring: After the model is deployed to the production environment, continuously monitor the performance indicators of the model, such as accuracy rate, recall rate, F1 value, etc. Through an online evaluation system, obtain the prediction results of the model for new data in real time, and compare them with the true values to calculate the performance indicators. For example, in a model for predicting customer purchase intention, monitor the matching degree between the model prediction results and the actual purchase behavior.

[0112] 4.2. Setting of trigger update conditions:

[0113] Data change trigger condition: Set the threshold for data change. When the change of the feature statistic of the data exceeds this threshold, trigger the model update process. For example, if the mean change of user behavior data exceeds 20%, it is considered that the data has changed significantly and the model needs to be updated.

[0114] Model performance trigger condition: Set the threshold for model performance. When the performance indicator of the model is lower than the set threshold, trigger the model update. For example, if the accuracy rate of the prediction model drops from 80% to 70% and remains at a low level for a continuous period of time, start the model update mechanism.

[0115] 4.3. Model Incremental Update Implementation: When the model update condition is triggered, incremental learning technology is used to train with new data on the basis of the original model and update the model parameters. For example, for a classification model based on decision trees, the incremental decision tree algorithm is adopted to gradually add new data to the decision tree and update the tree structure and node parameters. For scenarios with high real-time requirements, online learning technology is adopted, and the model immediately learns and updates when new data is received. For example, in a real-time recommendation system, the online gradient descent algorithm is used to continuously adjust the parameters of the recommendation model according to the real-time behavior data of users to improve the accuracy of recommendations.

[0116] 5. Multi-source Data Fusion Analysis:

[0117] 5.1. Data Preprocessing: Collect multi-source data from different data sources (such as databases, file systems, API interfaces, etc.), covering structured data (such as user information tables, transaction record tables), unstructured data (such as user comments, news articles), image data (such as product pictures, user avatars), video data (such as product promotion videos, surveillance videos), etc. Preprocess the collected multi-source data and adopt different processing methods for different types of data. For structured data, perform operations such as data cleaning, duplicate removal, and missing value filling; for unstructured data, perform operations such as text tokenization, part-of-speech tagging, word vector conversion (for text data), image normalization, and feature extraction (for image data); for video data, perform operations such as video decoding, key frame extraction, and video feature extraction.

[0118] 5.2. Data Association and Fusion: Analyze the potential association relationships between multi-source data and formulate data association rules. For example, in user information data and transaction data, establish an association through the user ID; in user comment data and product data, establish an association through the product ID. For image and video data, relevant features (such as product categories, brands, etc.) related to other data sources can be extracted through image recognition technology or video content analysis technology to establish an association.

[0119] Select appropriate data fusion methods according to the data type and association relationship, such as feature-level fusion methods (directly splicing the features of different data sources or combining them after feature transformation) and decision-level fusion methods (integrating the decision results of different data sources). For example, for user behavior data and user portrait data, a feature-level fusion method is adopted to splice the user's behavior features (such as purchase frequency, browsing duration) and portrait features (such as age, gender, region) to form a new feature vector.

[0120] 5.3. Data Analysis and Mining: Utilize the data after data association and fusion, and adopt data mining techniques (such as association rule mining, clustering analysis, anomaly detection, etc.) to mine the potential relationships and patterns between different data sources. For example, through association rule mining, discover the association relationship between users' purchase of a certain type of product and their browsing of specific advertisements; through clustering analysis, group users with similar behaviors and characteristics into one category to provide a basis for precision marketing. Provide support for business decisions based on the mined potential relationships and values. For example, based on the analysis results of users' purchase behaviors and product association relationships, optimize the product recommendation system to improve the accuracy and conversion rate of recommendations; according to the results of user clustering analysis, formulate differentiated marketing strategies to improve marketing effectiveness.

[0121] 6. Deep Data Mining: Include deep mining of text data and deep mining of graph data;

[0122] 6.1. Deep Mining of Text Data: For the deep mining of text data, select appropriate deep learning algorithms, such as recurrent neural networks (RNN) and their variants long short-term memory networks (LSTM), gated recurrent units (GRU), and the Transformer architecture, etc. For example, in sentiment analysis tasks, the LSTM network can be used to capture the context information in the text and judge the sentiment tendency of the text (positive, negative, or neutral); in topic model mining, the BERT model based on the Transformer can be adopted, and through pre-training and fine-tuning, extract the topic features of the text. Prepare text training data and perform data annotation (such as sentiment labels, topic labels). Train the selected deep learning model, adjust the model parameters, and optimize the model performance. After training, apply the model to the deep mining tasks of actual text data. For example, use the trained sentiment analysis model to perform sentiment analysis on user review data, count the proportions of positive and negative reviews, and understand users' satisfaction with products or services; utilize the topic model to mine the topic distribution of a large number of documents to provide support for document classification and information retrieval.

[0123] 6.2. Deep Mining of Graph Data: For graph data such as social network data and knowledge graph data, select appropriate graph computing algorithms, such as the PageRank algorithm (used to calculate the importance of nodes in a graph), community discovery algorithms (such as the Louvain algorithm and Label Propagation algorithm, used to discover the community structure in a graph), and graph neural network algorithms (such as Graph Convolutional Network, GCN; Graph Attention Network, GAT, used for tasks such as feature learning and node classification of graph data). For example, in social network analysis, use the PageRank algorithm to identify user nodes with greater influence; in knowledge graph construction, use the GCN algorithm to perform feature learning and classification on entities and relationships in the knowledge graph. Construct social network data, knowledge graph data, etc. into a graph structure, and preprocess the graph data, including operations such as attribute extraction of nodes and edges and graph normalization. Use the selected graph computing algorithms to analyze the graph data and discover the relationships and patterns between nodes. For example, through community discovery algorithms, identify interest groups or communities in social networks; use graph neural network algorithms to predict potential relationships between entities in the knowledge graph and improve the construction of the knowledge graph. Perform deep mining of image data according to the following formula:

[0124] PR(i) = (1 - d)A i C i + (d / N)Σ j∈Bi [PR(j) * T ij C j ; T ij = [T max - (t now - t create )] / T max ;

[0125] In the formula, (PR(i) is the PageRank value obtained by iterative calculation of node i, which is used to measure the comprehensive importance of node i in the graph structure. d is the damping factor, usually taking a value of about 0.85. It is an empirical parameter used to balance random browsing behavior and behavior through link jumps. A i is the node activity factor, which is used to measure the activity level of node i in the network. It is calculated by counting the number of operations of node i within a certain time window, such as the number of posts and comments in a social network and the update frequency in a knowledge graph. Its value range is between 0 and 1, and the higher the activity, the closer A i is to 1. C iis the community influence factor, representing the influence of node i in its affiliated community. It is comprehensively obtained by calculating indicators such as the degree centrality and betweenness centrality of node i within its community, and its value range is between 0 and 1. N is the total number of nodes in the graph. B i is the set of nodes pointed to by the out-links of node i. This set contains all the nodes that can be directly reached from node i through links. PR(j) is the PageRank value of node j, where node j belongs to the set Bi of nodes pointed to by the out-links of node i. T ij is the connection timeliness factor, reflecting the timeliness of the connection between node i and node j. For connections with timestamp records (such as the follow-up time between users in a social network, the establishment time of entity relationships in a knowledge graph, etc.), t now is the current time, t create is the connection establishment time, T max is the effective duration of the connection (which can be set according to the business scenario. For example, in a social network, it can be the average duration of the user active period), and its value range is between 0 and 1. L j is the number of out-links of node j. It represents the number of out-links from node j to other nodes.

[0126] The present invention provides an information technology analysis method based on cloud computing, including the following steps:

[0127] S1. The network construction module uses 10 Gigabit Ethernet to build a high-speed core network, intelligently schedules traffic in combination with SDN technology, and uses CDN to cache static data to nodes near users, improving the transmission speed and reducing the pressure on the data center.

[0128] S2. The resource scheduling module uses virtualization technologies such as KVM to divide physical servers into virtual machines, and combines Docker containerization technology to improve resource utilization and application isolation. Through an intelligent scheduling algorithm based on machine learning, it monitors and dynamically allocates resources in real time to ensure that high-priority and real-time applications obtain sufficient resources and avoid performance bottlenecks.

[0129] S3. The data storage module uses a distributed file system such as Ceph to achieve redundant data storage and high availability, and automatically repairs and balances data. It constructs a data lake to store multi-format data, and uses technologies such as Apache Hudi to efficiently manage data, support real-time operations, and provide rich data sources for analysis.

[0130] S4. The data collection module supports multi-source data collection, develops general interfaces, uses ETL tools to convert data formats, and realizes unified processing. At the same time, it uses data quality monitoring tools to monitor data quality in real time, mark and process unqualified data to ensure data accuracy, integrity, and consistency.

[0131] S5. The data integration module formulates and manages unified data standards, uses data integration tools to integrate data from different data sources, performs data cleansing, converts data formats and structures, and stores them in data warehouses or data lakes.

[0132] S6. The data management module builds a relational database data warehouse and adopts efficient model design; establishes a data life cycle management platform to manage the entire data process; formulates backup and recovery strategies to ensure data security; implements hierarchical storage to optimize storage resource utilization.

[0133] S7. Data analysis module: Establish a professional algorithm team, develop and optimize data analysis algorithms and models, establish an algorithm library and centrally manage it. Use a distributed computing framework to implement large-scale data training and establish an automatic model update mechanism. Conduct multi-source data fusion analysis to explore potential relationships. Deep data mining uses advanced technology to conduct in-depth analysis of large-scale data to provide strong support for decision-making.

[0134] S8. The result display module integrates visual analysis tools, provides an intuitive and easy-to-use interface, and supports drag-and-drop generation of reports and charts. Develop customized data analysis applications to meet specific business needs. Achieve seamless integration with the company's existing business systems and data sharing and collaboration.

[0135] S9. The security management module uses SSL / TLS and encryption algorithms to ensure the security of data storage and transmission, establishes an access control system to limit user permissions, formulates a data privacy policy to protect user privacy, and ensures that data is used legally and in compliance with regulations.

[0136] S10. The monitoring management module deploys monitoring tools to monitor system operation in real time, establishes fault management processes for rapid response, formulates operation and maintenance plans and specifications, and uses automated tools to improve operation and maintenance efficiency.

[0137] The present invention can realize flexible resource utilization, utilize virtualization and containerization technology, and combine intelligent scheduling algorithms to realize dynamic allocation and efficient utilization of resources, ensure key application performance, and avoid resource waste. Realize highly reliable data storage, use distributed file system to realize data redundancy and high availability, automatic repair and balancing, build data lake to support multi-data storage and efficient management, and provide a solid foundation for analysis. Realize comprehensive data management, build relational database warehouse, implement full life cycle management, formulate backup and recovery strategies, optimize storage resources, and ensure data security and efficient utilization. Realize in-depth data analysis, form a professional team to develop and optimize algorithm models, use distributed computing to accelerate training, realize automatic model update, perform multi-source data fusion and deep mining, and provide a scientific basis for decision-making.

[0138] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An information technology analysis method based on cloud computing, characterized in that: The following steps are involved: S1. The network construction module uses 10 Gigabit Ethernet to build a high-speed core network, combines SDN technology to intelligently schedule traffic, and uses CDN to cache static data to nodes near users to improve transmission speed; S2, the resource scheduling module uses KVM virtualization technology to divide physical servers into virtual machines, combined with Docker containerization technology to improve resource utilization and application isolation; real-time monitoring and dynamic allocation of resources; S3, the data storage module uses distributed file systems such as Ceph to achieve data redundant storage and high availability, and automatically repair and balance data; S4, data acquisition module collects data from multiple sources and monitors data quality in real time; S5. The data integration module integrates data from different data sources, cleans the data, converts the data format and structure, and stores it in a data warehouse or data lake; S6. The data management module builds a relational database data warehouse, establishes a data life cycle management platform, and manages the entire data process; Develop backup and recovery strategies to ensure data security; S7, the data analysis module establishes a model automatic update mechanism, conducts multi-source data fusion analysis, and mines potential relationships; S8. The result display module integrates visual analysis tools and provides an intuitive and easy-to-use interface; S9, security management module ensures data storage and transmission security; S10. The monitoring management module deploys monitoring tools to monitor system operation in real time.

2. The information technology analysis method based on cloud computing according to claim 1 is characterized in that: Step S2 includes the following steps: S21. Preliminary preparation: Check the physical server hardware, plan the number and roles of servers, install a Linux system that supports KVM, configure the Docker environment, design a virtualized network architecture, divide different subnets, configure network interfaces and firewall rules, and ensure connectivity and security; S22, virtualization and containerization operations: load the KVM module in the server operating system and install the qemu-kvm tool to create and manage virtual machines; build images and run containers; S23. Construct resource scheduling algorithm model: S23.

1. Data collection and collation: Install monitoring tools in virtual machines and containers to collect resource data, use time series databases to store data, and regularly collect and store data through data collection agents; perform user business demand analysis, quantify demand and associate it with resource data; S23.2, Machine learning model training: select Q-learning algorithm, process historical data, build training environment, adjust parameters to optimize model performance; in the model evaluation and optimization stage, use test set to evaluate model performance, and further optimize the model based on the results until the performance meets the standard; S24. Actual scheduling implementation: S24.

1. Real-time resource monitoring and data update: Continuously collect virtual machine and container resource data and update it to the database, adjust it according to the needs of the business department, and update records related to resource scheduling in a timely manner; S24.

2. Resource scheduling decision execution: Input real-time data and business requirements into the machine learning model to generate resource allocation decisions, and use management tools to dynamically adjust the resources of virtual machines and containers based on the decisions; S24.

3. Effect feedback and continuous optimization: Monitor application operating status and resource utilization efficiency, compare feedback data with expected goals, evaluate model performance and optimize resource scheduling strategies to adapt to changes in business needs and system operating environment.

3. The information technology analysis method based on cloud computing according to claim 2 is characterized in that: In step S23.2, the resource scheduling model is: Q(s,a)←Q(s,a)+A[w*r(1-β)+γmax a’ Q(s',a')-Q(s,a)]; where Q(s,a) represents the discounted sum of the expected cumulative rewards after taking action a in state s; A is the learning rate; γ is the discount factor; r is the immediate reward; w is the task urgency weight; β is the resource volatility coefficient; s' is the new state; max a' Q(s',a') represents the maximum Q value that can be obtained by taking all possible actions a' in the new state s'.

4. The information technology analysis method based on cloud computing according to claim 1, characterized in that: Step S5 includes the following steps: S51. Select data integration tool: Select Talend data integration tool based on data source type, data volume, data processing complexity and budget factors; S52. Connect to data sources: Use data integration tools to configure connections with various data sources, ensure normal connections, and develop full or incremental data extraction strategies based on business needs; S53, Data cleaning: including removing duplicate data, correcting erroneous data and filling missing data; S54, Data Mapping and Conversion: Develop mapping rules to map data fields from different data sources to the target data format, and use data integration tools to perform data conversion; S55. Data storage: Evaluate the use of data warehouses or data lakes to store data, select appropriate storage solutions based on data characteristics, usage scenarios, and business needs, develop data loading strategies, and use data integration tools to store cleaned and converted data in the target system according to the selected solution.

5. The information technology analysis method based on cloud computing according to claim 4 is characterized in that: In step S54, data conversion is performed according to the following formula: Q 综合 =w1*P map +w2*P conv +w3*C;w1+w2+w3*C=1; In the formula, Q 综合 It is a comprehensive indicator of data conversion quality; P map is the mapping accuracy; P conv is the conversion success rate; C is the consistency between the converted data and the target format; w1, w2, w3 are the mapping accuracy P map , converted into power P conv and the consistency degree C weight.

6. The information technology analysis method based on cloud computing according to claim 1, characterized in that: Step S7 includes the following steps: S71. Build algorithm model: select appropriate algorithm direction and design model according to scenario and data, implement algorithm using programming language and framework, ensure code is standardized and modular, and use test data for debugging and repair; S72. Algorithm library construction and management: Design a layered architecture, store the tested algorithms in the library according to the specifications and manage the versions, and use version control tools to record the change history; S73. Distributed model training: S73.

1. Framework evaluation and selection: Evaluate and select a suitable distributed computing framework based on the project's data scale, computing resources, and algorithm type; S73.2, Framework installation and configuration: According to the selected distributed computing framework, perform compatibility installation and configuration, including setting cluster parameters and resource allocation, and verify the correctness of the installation and configuration through basic function testing; S73.

3. Data parallel and model parallel implementation: including data parallel implementation and model parallel implementation; Data parallel implementation: The training data is split and distributed to different computing nodes for parallel processing. Each node uses the same model to train its own data subset and updates the model parameters synchronously. Model parallel implementation: Distribute different parts of complex deep learning models to different computing nodes for parallel computing, and ensure data delivery and synchronization through reasonable communication mechanisms; S73.4, Model training and optimization: including data preprocessing, data parallel storage, real-time monitoring of training indicators, adjustment of training strategies, use of optimization algorithms to update model parameters, model pruning, and quantization operations to improve model performance and operating efficiency; S74, Model Automation Update: S74.

1. Data and model monitoring: Establish a data monitoring system to monitor data changes in real time from the data source; after the model is deployed to the production environment, continuously monitor the performance indicators of the model; S74.

2. Setting trigger update conditions: Setting a threshold for data change. When the characteristic statistic of the data changes beyond the threshold, the model update process is triggered; Setting a threshold for model performance. When the performance index of the model is lower than the set threshold, the model update is triggered; S74.

3. Implementation of incremental model update: When the update condition is triggered, incremental learning or online learning technology is used to train the original model based on new data to update model parameters and structure to meet real-time requirements; S75. Conduct multi-source data fusion analysis; S76. Deep data mining: including text data deep mining and graph data deep mining; S76.

1. Deep mining of text data: Select appropriate deep learning algorithms, prepare and annotate training data, train and optimize model parameters, and apply the model to actual text mining tasks; S76.

2. Deep mining of graph data: Select appropriate graph computing algorithms to perform node importance calculation, community structure discovery, feature learning and node classification tasks on graph data.

7. The information technology analysis method based on cloud computing according to claim 6 is characterized in that: Step S75 includes the following steps: S75.

1. Data preprocessing: Collect structured, unstructured, image, and video data from multiple sources, and take appropriate preprocessing measures for each data type, including data cleaning, text processing, and image and video feature extraction to ensure data quality and extract useful information; S75.2, Data association and fusion: Analyze the potential associations between multi-source data, formulate association rules, establish data connections through ID matching or feature extraction, and select appropriate data fusion methods to integrate information from different data sources; S75.

3. Data analysis and mining: Using data association and fusion, data mining techniques are used to explore potential relationships and patterns between different data sources.

8. The information technology analysis method based on cloud computing according to claim 6 is characterized by: In step S76.2, the image data deep mining model is: PR(i)=(1-d)A i C i +(d / N)Σ j∈Bi [PR(j)*T ij C j ];T ij =[T max -(t now -t create )] / T max , Where PR(i) is the PageRank value of node i calculated by iteration; d is the damping coefficient; A i is the node activity factor; C i is the community influence factor, which indicates the influence of node i in the community to which it belongs; B i is the set of nodes pointed to by outbound links of node i; PR(j) is the PageRank value of node j, where node j belongs to the set of nodes Bi pointed to by outbound links of node i; T ij is the connection timeliness factor; t now is the current time, t create is the connection establishment time, T max is the effective duration of the connection; L j is the number of outbound links of node j.

9. The information technology analysis method based on cloud computing according to claim 6, characterized in that: In step S74.1, the data change monitoring model is: D=Σ n i=1 {w i *|c i -h i |*[1+Σ N i=1 (R ij *τ i )]+O i };Σ n i=1 (w i )=1; where D is the data change difference index; n is the number of data feature statistics; w i is the weight of the ith feature statistic; c i is the value of the characteristic statistic of the i-th data currently collected; h i is the value of the ith data feature statistic in the historical data; R is the data feature correlation coefficient matrix, whose element R ij (i,j=1,2,...,n) represents the correlation coefficient between the i-th feature and the j-th feature; τ i is the data change trend factor; i is the outlier correction term.

10. An information technology analysis system based on cloud computing, comprising: Network construction module, monitoring management module, resource scheduling module, data storage module, data acquisition module, data integration module, data management module, data analysis module, result display module and security management module; characterized by: Network construction module: Use 10 Gigabit Ethernet to build a high-speed core network, combine SDN technology to intelligently schedule traffic, and use CDN to cache static data to nodes near users to improve transmission speed; Resource scheduling module: Use KVM virtualization technology to divide physical servers into virtual machines, and combine Docker containerization technology to improve resource utilization and application isolation; monitor and dynamically allocate resources in real time; Data storage module: using distributed file systems such as Ceph to achieve data redundant storage and high availability, automatic repair and balancing of data; Data collection module: collects data from multiple sources and uses data quality monitoring tools to monitor data quality in real time; Data integration module: Integrate data from different data sources, perform data cleaning, convert data formats and structures, and store them in a data warehouse or data lake; Data management module: build a relational database data warehouse, establish a data life cycle management platform, manage the entire data process; formulate backup and recovery strategies to ensure data security; Data analysis module: establish an automatic model update mechanism, conduct multi-source data fusion analysis, and explore potential relationships; Result display module: integrates visual analysis tools and provides an intuitive and easy-to-use interface; Security management module: ensure data storage and transmission security; Monitoring management module: deploy monitoring tools to monitor system operation in real time.

Citation Information

Patent Citations

  • Information technology analysis system based on cloud computing

    CN118170845A

Cited By

  • Enterprise information visual management system and method based on data sharing

    CN120256620A

  • Technical supervision dynamic management and control system applied to thermal power-new energy scene

    CN120810935A

  • Power grid equipment fault analysis method and device, computer equipment and storage medium

    CN120892898A

  • ETL scheduling method based on conch data platform, computer and storage medium

    CN120950587A

  • New plastic material research and development data processing method based on cloud platform

    CN120952719A