Intelligent detection AI model training cloud platform system
By designing intelligent detection AI model training cloud platform system on cloud computing platform, using the collaborative work of multi-objective optimization, resource scheduling, reinforcement learning, Bayesian optimization, data processing and security guarantee modules, the problems of resource scheduling in the existing technology cannot be dynamically adapted, the efficiency of hyperparameter optimization, obvious bottlenecks in data processing and transmission, and insufficient data security guarantees are solved, and efficient resource utilization, rapid hyperparameter optimization and comprehensive data security protection are achieved.
Patent Information
- Application Number
- CN202510465756.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cloud computing platforms cannot dynamically adapt to multi-task concurrent scenarios in resource scheduling, low efficiency of hyperparameter optimization, obvious bottlenecks in data processing and transmission, and insufficient data security guarantees.
An intelligent detection AI model training cloud platform system is designed, including multi-objective optimization module, resource scheduling module, reinforcement learning module, Bayesian optimization module, data processing module and security guarantee module. Through the collaborative work of these modules, dynamic optimization of resource allocation strategies, efficient management of data transmission and full-process protection of data security are achieved.
It effectively solves the problem of low resource utilization in multi-task concurrency scenarios, significantly improving training efficiency; quickly determines the optimal hyperparameter configuration, avoiding inefficient processes of a large number of trials; through end-to-end encryption and zero-trust access control technology, the security protection of the entire data process is strengthened, ensuring privacy and security in data transmission and storage processes.
Smart Images

Figure CN119987979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing and artificial intelligence technology, and specifically to a cloud platform system for training an intelligent detection AI model. Background Art
[0002] With the continuous development of artificial intelligence technology and the increasing complexity of application scenarios, the training of deep learning models has become the core link in realizing various intelligent tasks. However, the training process of deep learning models usually consumes a lot of computing resources, and the training tasks are highly dynamic and complex. In this context, cloud computing platforms have gradually become the main supporting environment for AI model training, providing powerful computing power and flexible resource allocation modes to cope with the large-scale data processing and high-density computing requirements in training tasks.
[0003] Existing cloud computing platforms usually provide stable computing support for different types of training tasks based on fixed resource allocation strategies. This approach can ensure the predictability of resource allocation and the ease of management in scenarios with relatively simple task loads. At the same time, existing technologies achieve basic balanced resource allocation through preset task scheduling rules, which can meet conventional training needs in low-concurrency situations.
[0004] However, with the increase in the scale of deep learning tasks, existing technologies have gradually shown obvious deficiencies in dynamic resource scheduling, parameter optimization efficiency and data security. On the one hand, fixed resource allocation strategies cannot adapt to dynamic load changes in multi-task concurrent scenarios, resulting in over-allocation and waste of some resources and delays in training progress due to insufficient supply of other resources. On the other hand, parameter optimization methods such as grid search are inefficient and require a large amount of computing resources and time in high-dimensional parameter space. In addition, existing data processing systems lack dynamic traffic control mechanisms, which leads to bandwidth bottlenecks when data is transmitted across multiple nodes, seriously affecting the overall efficiency of the task. Finally, security deficiencies are even more obvious. Traditional static encryption or simple permission management modes lack real-time protection capabilities in complex environments. Security risks in data transmission and access are frequent, and there is a lack of a transparent operation traceability mechanism, making it difficult to achieve comprehensive system security. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides an intelligent detection AI model training cloud platform system, which solves the problems in the prior art that resource scheduling cannot dynamically adapt to multi-task concurrent scenarios, hyperparameter optimization efficiency is low, data processing and transmission bottlenecks are obvious, and data security is insufficient.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a cloud platform system for training an intelligent detection AI model, the system comprising: Multi-objective optimization module, which is used to optimize the resource allocation strategy of the cloud platform to achieve multi-objective optimization of resource utilization, training time and energy consumption; Resource scheduling module, which is used to dynamically schedule computing resources according to the Markov decision process model to ensure that resources are reasonably allocated according to task requirements; Reinforcement learning module, used to automatically learn the optimal resource scheduling strategy through deep Q learning or policy gradient algorithm; Bayesian optimization module, which is used to optimize the hyperparameters in resource scheduling to further improve the scheduling efficiency of the system; Data processing module, used to optimize the transmission and processing of large-scale data, reduce data bottlenecks, and improve data flow efficiency; The security module is used to ensure the privacy and security of data during transmission and prevent data leakage and illegal access.
[0007] Preferably, the multi-objective optimization module includes: An objective function generating unit, used to generate an objective function according to resource utilization, training time and energy consumption of the cloud platform; The optimization calculation unit is used to optimize the objective function through the weighted sum method to solve the optimal resource allocation plan.
[0008] Preferably, the objective function generating unit comprises: A resource utilization calculation unit, used to calculate resource utilization according to resource allocation conditions of each node of the cloud platform; A training time calculation unit is used to estimate the time required to complete the training task; The energy consumption calculation unit is used to calculate the energy consumption under a given resource configuration.
[0009] Preferably, the resource scheduling module includes: A state evaluation unit, used to evaluate the current resource usage state of the cloud platform and generate a current state representation; An action selection unit, used to select resource allocation actions based on current status and historical experience; The state transition control unit is used to calculate and control the transition probability from the current state to the next state.
[0010] Preferably, the Markov decision process model includes: The state space definition unit is used to define the resource allocation state of each computing node in the cloud platform; An action space definition unit, used to define resource adjustment actions and generate possible resource allocation solutions; The transition probability calculation unit is used to calculate the state transition probability based on the current resource state and the action taken.
[0011] Preferably, the reinforcement learning module includes: A deep Q network training unit, used to update the Q value function and train the scheduling strategy through a deep Q learning algorithm; A policy gradient calculation unit, used to optimize the scheduling strategy through the policy gradient method; The reward feedback unit is used to calculate the immediate reward based on the feedback of task execution and optimize the strategy.
[0012] Preferably, the Bayesian optimization module includes: Gaussian process modeling unit, used to model the objective function of the scheduling system and estimate the posterior distribution; The acquisition function calculation unit is used to select the optimal resource allocation strategy according to the posterior distribution of the current resource configuration.
[0013] Preferably, the data processing module includes: Data storage optimization unit, used to distribute and store training data on different computing nodes of the cloud platform, supporting distributed parallel processing; The data flow control unit is used to automatically adjust the data flow path based on the task priority and data volume to ensure efficient data transmission.
[0014] Preferably, the security assurance module includes: End-to-end encryption unit, used to encrypt data during transmission to ensure data privacy; Blockchain audit unit, which is used to record data access logs to ensure the traceability of the data transmission process; Zero Trust authentication unit for authenticating and controlling access to users and devices.
[0015] Preferably, the system further comprises: Resource monitoring unit, which is used to monitor the usage of cloud platform resources in real time and provide visual reports on resource usage; The adaptive optimization unit is used to automatically adjust the resource scheduling strategy according to the platform operation status.
[0016] The present invention provides a cloud platform system for intelligent detection AI model training. It has the following beneficial effects: 1. This invention combines reinforcement learning and multi-objective optimization technology to achieve efficient scheduling and energy consumption optimization by dynamically adjusting resource allocation. Compared with the traditional static allocation strategy, it effectively solves the problem of low resource utilization in multi-task concurrent scenarios and significantly improves training efficiency.
[0017] 2. The present invention uses a Bayesian optimization module to quickly determine the optimal hyperparameter configuration through Gaussian process modeling, avoiding the inefficient process of a large number of experiments. Compared with the traditional grid search method, it is faster and more accurate, solving the problem of time-consuming and inefficient hyperparameter tuning.
[0018] 3. This invention strengthens the security protection of the entire data process through end-to-end encryption and zero-trust access control technology, ensuring privacy during transmission and storage. Compared with the existing permission management method, it eliminates the hidden danger of data leakage and enhances the flexibility and security of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a system structure diagram of the present invention; Figure 2 It is a module architecture diagram of the multi-objective optimization module of the present invention; Figure 3 This is a module architecture diagram of the resource scheduling module of the present invention; Figure 4 This is a module architecture diagram of the reinforcement learning module of the present invention; Figure 5 The module architecture diagram of the Bayesian optimization module of the present invention; Figure 6 This is a module architecture diagram of the data processing module of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Please see attached Figure 1 -Attached Figure 6 , an embodiment of the present invention provides an intelligent detection AI model training cloud platform system, the system comprising: Multi-objective optimization module, which is used to optimize the resource allocation strategy of the cloud platform to achieve multi-objective optimization of resource utilization, training time and energy consumption; This module is used to solve the contradiction between resource utilization, training time and energy consumption, and ensure the efficient operation of the cloud platform. It converts these goals into a comprehensive optimization problem and adopts reasonable strategies to solve it, ultimately achieving the maximum resource utilization, the minimum training time and the reduction of energy consumption.
[0022] When training AI models on a cloud platform, the system needs to consider multiple optimization objectives at the same time. These objectives include: on the one hand, improving the utilization efficiency of computing resources; on the other hand, reducing the completion time of training tasks, while also controlling the energy consumption of the system when executing tasks. Therefore, how to achieve the best balance between these three objectives is the core problem that the multi-objective optimization module needs to solve.
[0023] In the workflow of the multi-objective optimization module, the weight of each objective is first evaluated according to the task requirements and resource status, and then the weighted summation method is used to merge multiple objective functions into a comprehensive optimization objective. Through continuous optimization, the resource allocation method gradually approaches the optimal state, thereby ensuring the smooth completion of the task while improving resource utilization efficiency and reducing energy consumption.
[0024] In this embodiment, the multi-objective optimization module combines the three objective functions of resource utilization, training time and energy consumption into an overall optimization function through the weighted summation method. The goal of this function is to achieve optimization by adjusting the resource allocation strategy. The specific implementation method is as follows: In the present invention, the objective function is combined in a weighted summation manner. Its basic form is: ; in, is a resource allocation vector, which indicates the resource allocation of each computing node in the cloud platform (such as CPU, GPU resource usage, memory, etc.); It is a function of resource utilization, which indicates the resource utilization efficiency of the cloud platform under the current resource allocation; is a function of training time, indicating the training time of the task; It is a function of energy consumption, which indicates the total energy consumption of the task during execution.
[0025] Weight coefficient , , It is a weight coefficient assigned to each goal based on actual needs to ensure a balance between multiple goals. In general, the weight coefficient meets the following conditions: ; in, : represents the weight coefficient of resource utilization; : represents the weight coefficient of training time; : represents the weight coefficient of energy consumption; In the specific implementation, the resource utilization function It is mainly used to measure the efficiency of cloud platform resource utilization. The overall resource utilization of the platform is obtained by calculating the resource utilization rate of each computing node of the cloud platform.
[0026] In a possible implementation, the resource utilization rate is calculated as follows: ; in, : Resource utilization function; :Indicates the The resource allocation of each computing node; : Indicates the total number of computing nodes in the cloud platform; : Indicates the total resources of the cloud platform; Training time function It is mainly used to measure the time required to complete a training task under a given resource configuration. The training time is inversely proportional to the allocation of computing resources, and more computing resources can reduce the training time.
[0027] In the specific implementation process, the training time is calculated as follows: ; in: Indicates that in resource configuration The time required to complete the training task; is the theoretical maximum time of the training task, usually the time required for the training task under the worst resource configuration; the goal of the training time function is to minimize , that is, shorten the training time as much as possible while ensuring the completion of the task.
[0028] Energy consumption function Used to measure the energy consumed during training. Reasonable resource allocation can not only improve training efficiency, but also reduce unnecessary energy consumption.
[0029] In this embodiment, the energy consumption function is in the form of: ; in: Indicates the cloud platform The energy consumption per unit resource of a computing node is usually related to the performance and resource usage of the hardware device; Indicates The resource allocation of each node.
[0030] The purpose of this function is to reduce energy consumption by optimizing resource allocation. During the training process, excessive allocation of computing resources will lead to unnecessary energy consumption, so the optimization of energy consumption is an important goal in the present invention.
[0031] In the implementation of the present invention, the optimization process solves the optimal resource allocation strategy through an algorithm. First, the value of each objective function is calculated, and then they are weighted and summed according to the set weight coefficient to obtain a comprehensive optimization objective function. Finally, an optimization algorithm (such as particle swarm optimization, genetic algorithm, gradient descent, etc.) is used to find a resource allocation solution that can minimize the optimization objective function.
[0032] This module achieves a dynamic balance of computing resources by comprehensively optimizing resource utilization, training time, and energy consumption. Through this optimization, the system can find the best compromise between multiple goals, thereby improving resource utilization efficiency, shortening training time, and reducing system energy consumption.
[0033] In this embodiment, the multi-objective optimization module uses a weighted summation method to achieve balanced optimization of resource utilization, training time and energy consumption. Through the precise calculation of each objective function, the system can flexibly adjust resource allocation according to the requirements of different tasks to ensure the efficient operation of the cloud platform; Resource scheduling module, which is used to dynamically schedule computing resources according to the Markov decision process model to ensure that resources are reasonably allocated according to task requirements; The resource scheduling module dynamically adjusts the computing resources in the cloud platform according to the current resource status of the system, task requirements, and feedback from historical tasks. By using the Markov decision process (MDP) model and reinforcement learning algorithm, the system can flexibly schedule computing resources according to the needs of real-time tasks to ensure that tasks can be completed on time and efficiently.
[0034] The design and implementation of the resource scheduling module are closely connected with the multi-objective optimization module and the reinforcement learning module. In the aforementioned multi-objective optimization module, a comprehensive optimization function including resource utilization, training time, and energy consumption has been established, and the resource scheduling module dynamically adjusts resource configuration based on this optimization goal. Specifically, the resource scheduling module will select the optimal resource allocation strategy based on the results of the objective function and the actual resource usage of the cloud platform, thereby optimizing the execution efficiency of the task.
[0035] In addition, the implementation of this module also makes full use of the Markov decision process (MDP) model to more accurately predict and select the next resource configuration. Through the MDP model framework, the resource scheduling module can select a suitable resource configuration action at each moment according to the current state of the cloud platform, evaluate the effect of the action, and then continuously adjust the scheduling strategy.
[0036] In this embodiment, the resource scheduling module models resource scheduling based on the Markov decision process (MDP). The MDP model is mainly composed of a state space, an action space, a transition probability, and a reward function, and these elements will be described in detail one by one below.
[0037] State Space Represents the resource usage status of the cloud platform at a certain moment. The resource allocation of each cloud platform node (such as CPU, GPU, memory, etc.) constitutes part of the state vector. Specifically, the system monitors the resource usage of each computing node and integrates this information into a complete state representation. In some embodiments, the state space It can be expressed as follows: ; in, : Indicates at time Next, the state vector of the system; : Indicates the current time step; : Indicates at time Next, Resource usage of each computing node; : Indicates the total number of computing nodes in the cloud platform; Action Space Indicates at time Resource adjustment actions taken. Each action Represents the adjustment of computing resources in the current state. The action may include allocating more resources to a computing node, or reducing the resource allocation of a node. The system can choose an action from multiple possible resource adjustment schemes to update the resource configuration of the platform. The action space can be represented as: ; in, : Indicates at time The resource scheduling action vector executed by the system; : indicates the current time step; : Indicates at time , for Resource adjustment actions performed by each computing node; : Indicates the total number of computing nodes in the cloud platform; Transition probability Indicates at time Take Action After that, the cloud platform changes from status Transfer to next state +1. This transition probability is usually calculated based on historical experience and resource scheduling rules to describe the impact of different resource allocation schemes on task execution. During implementation, the transition probability can be estimated through historical task data to infer the system state changes brought about by different resource allocation schemes.
[0038] Reward Function It is the most critical part of the resource scheduling module. It is used to evaluate the effect of taking a specific resource scheduling action under a given resource state. In this invention, the reward function is designed by comprehensively considering factors such as task training time, resource utilization and energy consumption. The specific function form is: ; in: : Reward value, measuring the state Next action The optimization effect brought about; : Indicates the change in training time after taking resource adjustment actions; : represents the change in energy consumption. A reduction in energy consumption will bring positive rewards. : Indicates the change in resource utilization. Improving resource utilization will result in higher rewards. , , : is a coefficient used to weigh training time, energy consumption and resource utilization, usually satisfying + + =1.
[0039] The purpose of the reward function is to guide the system to choose the best resource configuration by evaluating each resource scheduling action. Generally, the system will update the parameters in the reward function based on historical experience so that the scheduling strategy can be continuously optimized.
[0040] In the optimization process of the resource scheduling module, the present invention adopts a combination of deep Q learning (DQN) and policy gradient method to automatically learn and optimize the resource scheduling strategy. Specifically, the resource scheduling module will and historical data, and continuously adjust resource allocation plans through reinforcement learning algorithms.
[0041] Deep Q-learning (DQN) is used to solve resource scheduling problems in high-dimensional and complex environments. The system approximates the Q-value function through a neural network, calculates the expected reward for each possible action, and selects the optimal resource scheduling strategy. The Q-value update formula is as follows: ; in, : Q value function, indicating that in state Next select action The estimated value of the long-term cumulative reward that can be obtained. The Q value is used to guide the system to choose the optimal action; : Current system status, indicating the cloud platform at the moment the allocation of resources; : Current action, indicating that the state Specific adjustments to resource allocations made under the : Learning rate, which controls the weighted ratio of new information to old information when updating the Q value. The range is 0< ≤1; : Immediate reward, indicating taking action Finally, the system calculates the reward value based on resource utilization, training time and energy consumption; : Discount factor, which indicates the influence of future rewards on current decisions, ranging from 0≤ ≤1. The closer the value is to 1, the more the system focuses on long-term rewards; +1: next system state, from current state and actions transferred from : In state +1 to select the best action The corresponding maximum Q value is used to guide the next decision; In the policy gradient method, the policy is optimized by maximizing the long-term reward To achieve: ; in: are the parameters of the policy function; represents the gradient of the policy; is the advantage function, which is used to measure the action the good or bad.
[0042] After taking an action, the system Update to new status +1. This state transition is determined by the actual effect of resource adjustments.
[0043] In one possible implementation, the transition probability can be estimated by historical data or the performance of simulated training tasks. For example, when a node increases GPU resources, the training time may decrease, and the system will record this change to adjust the future transition probability.
[0044] The resource scheduling module will cycle through state evaluation, action selection, reward feedback, and strategy optimization until the optimal resource allocation for the training task is achieved.
[0045] In this embodiment, the resource scheduling module realizes the dynamic allocation of cloud platform resources through state evaluation, action selection, reward feedback and strategy optimization.
[0046] Reinforcement learning module, used to automatically learn the optimal resource scheduling strategy through deep Q learning or policy gradient algorithm; The reinforcement learning module achieves efficient allocation of cloud platform resources by learning the optimal resource scheduling strategy. The close relationship with the aforementioned resource scheduling module is that the reinforcement learning module enables the resource scheduling strategy to adapt to the dynamic changes of task requirements through continuous interaction and optimization, thus improving the intelligence and accuracy of scheduling.
[0047] In general, the resource scheduling module will select specific resource allocation actions based on the current task status, but this choice often depends on the prediction of future task results. The reinforcement learning module can find the best resource scheduling strategy in a large amount of task history data and real-time feedback by introducing technologies such as deep Q learning (DQN) and policy gradient algorithm. As an option, the reinforcement learning module can optimize scheduling decisions based on reward signals to ensure the maximization of long-term benefits.
[0048] In this embodiment, the reinforcement learning module implements its functions through the following steps, including: strategy representation, Q value update, reward feedback calculation and strategy optimization. These steps will be described in detail below.
[0049] In one possible implementation, the strategy It is the core of the reinforcement learning module. The strategy is expressed in the state Take action In order to make the strategy accurately reflect the dynamic changes of the task, this module uses a deep neural network to represent the strategy.
[0050] Specifically, the input is the current state , the output is for each possible action The network structure can be adjusted according to the complexity of the task, for example, a multi-layer perceptron (MLP) or a convolutional neural network (CNN) can be used. In general, the parameters of the network Through continuous optimization, the strategy can adapt to changes in tasks.
[0051] Q-value function in reinforcement learning module Indicates in status Next select action The expected cumulative reward that can be obtained. The Q value update rule is based on the Bellman equation, which is implemented by deep Q learning (DQN) in this embodiment.
[0052] In this embodiment, the reward function Used to measure the current action In general, the reward function will take into account multiple factors such as task completion efficiency, resource utilization, and energy consumption; The reinforcement learning module optimizes the policy function To improve long-term benefits. As an option, this module uses the policy gradient method to achieve optimization, and the optimization goal is to maximize the long-term cumulative reward ; In some embodiments, in order to improve the stability and convergence speed of the strategy, this module introduces the following extended optimization techniques: Specifically, the module can combine the proximal policy optimization (PPO) method to limit the amplitude of policy updates, thereby avoiding instability caused by excessive policy changes.
[0053] As an option, the module can also be combined with dual Q learning to reduce the problem of overestimation of Q values by introducing two independent Q value networks for estimation.
[0054] In this embodiment, the reinforcement learning module realizes the learning and optimization of the cloud platform resource scheduling strategy through four main links: strategy representation, Q value update, reward feedback and strategy optimization.
[0055] Bayesian optimization module, which is used to optimize the hyperparameters in resource scheduling to further improve the scheduling efficiency of the system; This module introduces the Bayesian optimization method to find the optimal parameter configuration while significantly reducing the number of experiments, thereby improving the overall efficiency of the system. Compared with the strategy optimization of the reinforcement learning module, the Bayesian optimization module focuses more on parameter selection in resource scheduling, such as GPU allocation ratio, training task batch size and other specific configuration parameters.
[0056] Generally speaking, when cloud platforms perform resource allocation or task scheduling, multiple hyperparameters with different value ranges may be involved. The interaction between these hyperparameters is complex, and traditional grid search or random search methods are often difficult to solve efficiently. As an alternative, Bayesian optimization introduces a probabilistic model (such as Gaussian process) to describe the posterior distribution of the objective function, combined with the acquisition function to guide the optimization direction, thereby achieving efficient parameter optimization.
[0057] In this embodiment, the Bayesian optimization module mainly includes the following steps: objective function modeling, posterior distribution estimation, acquisition function calculation, and parameter optimization. These steps work together to quickly converge to the optimal configuration of hyperparameters.
[0058] In the Bayesian optimization module, the objective function It is a performance indicator of resource scheduling or task execution. The goals that need to be optimized usually include task training time, energy consumption or resource utilization. Specifically, the objective function is a function of hyperparameters. The non-convex function may not be directly solved by analytical methods.
[0059] In one possible implementation, the objective function can be expressed as: ; in: Represents a vector of hyperparameters, such as batch size, learning rate, or GPU allocation ratio; Indicates the task training time under the current parameter configuration, in seconds; is the theoretical maximum value of training time, used to normalize time indicators; Indicates the total energy consumption under the current parameter configuration, in joules (J); It is the theoretical maximum value of energy consumption and is used to normalize the energy consumption index; Indicates the resource utilization under the current parameter configuration, without unit; It is the maximum value of resource utilization and is used to normalize resource utilization; , , are the weight coefficients of training time, energy consumption and resource utilization, which are used to balance the priorities of different optimization objectives. The purpose of objective function modeling is to find an optimal hyperparameter configuration in a given parameter space so that the objective function Take the minimum value.
[0060] After the objective function is modeled, the Bayesian optimization module estimates the posterior distribution of the objective function through the Gaussian process (GP). The Gaussian process is a non-parametric probability model that can model the objective function based on existing observation data.
[0061] Specifically, the Gaussian process uses a mean function and covariance function Represents the posterior distribution of the objective function: ; in: Indicates that the objective function is in the hyperparameter Expected value at Indicates that the objective function is at two parameter points and The covariance between represents the probability distribution of a Gaussian process.
[0062] In one possible implementation, the covariance function The radial basis kernel function or the Mahalanobis kernel function can be used. These kernel functions can capture the local variation characteristics of the target function, thereby improving the fitting accuracy of the posterior distribution.
[0063] After the Gaussian process completes the modeling of the objective function, the Bayesian optimization module selects the next hyperparameters through the acquisition function The acquisition function is used to balance the relationship between the currently known optimal solution and the exploration of new parameter space.
[0064] As an option, the acquisition function can use methods such as expected improvement or probability improvement. The formula for expected improvement is: ; in: Represents the currently known optimal value of the objective function; Represents the calculation of the expected value of the posterior distribution of a Gaussian process.
[0065] Specifically, maximizing the acquisition function can ensure that each step of optimization moves in the optimal direction while avoiding falling into a local optimal solution.
[0066] After the acquisition function selects the next step parameters, the Bayesian optimization module evaluates the objective function value through actual experiments and adds new observation points to the existing data. The new data points will update the posterior distribution of the Gaussian process, thereby further improving the fitting accuracy of the objective function.
[0067] In one possible implementation, the process of parameter optimization is as follows: Initialize the parameter space and initial observation points, and perform preliminary experiments with randomly selected hyperparameters.
[0068] The objective function is fitted through a Gaussian process and the posterior distribution is calculated.
[0069] Select new hyperparameters based on the acquisition function .
[0070] Update the observed data and refit the objective function, and repeat the above steps until the optimization termination condition is met (such as reaching the set maximum number of iterations or the convergence of the objective function value).
[0071] In this embodiment, the Bayesian optimization module realizes hyperparameter optimization in resource scheduling and task configuration through objective function modeling, posterior distribution estimation, acquisition function calculation and parameter optimization.
[0072] Data processing module, used to optimize the transmission and processing of large-scale data, reduce data bottlenecks, and improve data flow efficiency; The data processing module is used to efficiently manage and optimize the transmission and processing of large-scale data on the cloud platform. Its role is to provide fast data flow and storage support for AI model training tasks to avoid the decline in training efficiency due to data bottlenecks. The association with the Bayesian optimization module and the resource scheduling module is reflected in that this module optimizes the data transmission path and distributed storage structure to ensure that the required data can provide support for resource scheduling in a timely and accurate manner.
[0073] In general, data processing in cloud platforms needs to face the problem of multi-node transmission in a distributed environment and efficient scheduling of large-scale data flows. Traditional methods rely on fixed transmission paths and centralized storage strategies, which are prone to performance degradation due to bandwidth occupation or storage bottlenecks. As an option, the present invention adopts distributed storage and distributed computing technology, combined with a dynamic data flow control mechanism, to achieve efficient data management and support data processing requirements in a multi-task concurrent environment.
[0074] In this embodiment, the data processing module mainly includes the following functional units: a distributed storage unit, a data flow control unit, and a parallel computing unit. These units work together to form an efficient and stable cloud data processing architecture.
[0075] In a possible implementation, the distributed storage unit uses a distributed file system to manage the data of the cloud platform. The core idea of distributed storage is to split the data into multiple blocks and store them on different computing nodes to support efficient parallel read and write operations.
[0076] Specifically, the data It will be divided into multiple shards: ; in: Represents the complete data set that needs to be stored; Representation dataset No. shards; Indicates the total number of data shards. The size of each shard can be dynamically adjusted according to task requirements.
[0077] Generally, each data shard is stored on a different node, and multiple copies are created for each shard to ensure high data availability. In one possible implementation, the storage location of each shard is allocated through a consistent hashing algorithm to achieve uniform distribution and fast retrieval.
[0078] The task of the data flow control unit is to manage the transmission path and speed of training data between different computing nodes to avoid training bottlenecks caused by uneven bandwidth occupancy or data transmission delays.
[0079] Specifically, this unit optimizes data transmission through a dynamic flow control algorithm. In one possible implementation, the system dynamically adjusts the priority and transmission speed of data streams according to the current network bandwidth and node load. The selection of the transmission path can be achieved through the shortest path algorithm.
[0080] For example, suppose there is currently nodes need to access the dataset simultaneously , the system defines the priority of data flow as: ; in: Indicates The data flow priority of each node; Indicates The current network transmission delay of each node in seconds; Indicates The load of a node, usually expressed as the node's CPU or memory usage, without units.
[0081] By dynamically adjusting priorities, the data flow control unit can ensure that nodes with limited resources obtain the required data first, thereby avoiding computing stagnation caused by data delays.
[0082] The main function of the parallel computing unit is to support the distributed processing of large-scale data to meet the needs of efficient computing in a multi-task concurrent environment. In one possible implementation, this unit uses a distributed computing framework based on Apache Spark or Flink to process data in parallel.
[0083] Specifically, data processing tasks It will be divided into several subtasks: ; in: Represents a complete data processing task, such as an Epoch of model training; Indicates the task No. subtasks; Indicates the total number of subtasks. The processing scope of each subtask can be divided according to data sharding. to adjust the size.
[0084] In general, each subtask is assigned to the same computing node as the corresponding data shard, thereby reducing the data transmission overhead. Stored in the node If It will also be allocated to nodes first. .
[0085] In some embodiments, in order to further improve the performance of the data processing module, the present invention introduces the following extended optimization strategy: Specifically, the module can be combined with edge computing technology to sink some data processing tasks to edge nodes close to the data source to reduce the traffic pressure on the core network.
[0086] As an option, the module can also reduce the access latency to the distributed storage system by introducing a cache-based acceleration mechanism (such as Redis or Memcached) to store hot data in a high-speed cache.
[0087] In this embodiment, the data processing module realizes efficient management and processing of large-scale data through three functional units: distributed storage, data flow control and parallel computing.
[0088] Security module, used to ensure privacy and security during data transmission and prevent data leakage and illegal access; The main function of the security assurance module is to ensure the security and privacy of data during storage, transmission and processing. After completing the distributed storage and transmission optimization of large-scale data in the data processing module, the security assurance module uses a variety of security technologies to protect user data and system resources, avoiding security risks such as data leakage, tampering and illegal access.
[0089] In general, data transmission and storage on cloud platforms have potential security risks, especially in multi-task concurrent and distributed environments, where these risks are more significant. As an option, the present invention combines end-to-end encryption, blockchain auditing, and zero-trust access control technology to build a comprehensive security protection system to ensure that the entire process from data transmission to processing is under control.
[0090] In this embodiment, the security assurance module includes the following functional units: an end-to-end encryption unit, a blockchain audit unit, and a zero-trust access control unit. These units cooperate with each other to provide multi-level assurance for the safe operation of the system.
[0091] In one possible implementation, the end-to-end encryption unit uses the Advanced Encryption Standard (AES-256) to encrypt data to ensure that the data cannot be stolen or tampered with by an unauthorized third party during transmission.
[0092] Specifically, the encryption process involves two stages: encryption and decryption of data: Encryption process: Let the original data be , the key is , encrypted data It is expressed as: ; in, is the encryption algorithm function, It is ciphertext; In one possible implementation, the encryption key It is generated using an asymmetric encryption algorithm (such as RSA) and distributed to legitimate users by the authorization management module to ensure the security of the key.
[0093] Generally, the encryption key length is 256 bits. This high-strength encryption algorithm can effectively prevent brute force cracking.
[0094] The blockchain audit unit is used to record logs of all data access operations in the system and ensure that these log data cannot be tampered with, thereby providing highly transparent and traceable security.
[0095] In this embodiment, blockchain technology implements data auditing through distributed ledgers and smart contracts. Specifically, each data access generates a transaction record. : ; in: Indicates the user ID that performs the operation; Indicates specific access behavior, such as read, write, or modify; Indicates the timestamp of the operation; Represents the hash value of the previous transaction record.
[0096] Transaction Records It will be added to the blockchain to form a chain structure. The tamper-proof nature of the blockchain can effectively prevent log records from being maliciously modified.
[0097] Specifically, smart contracts can also be used to control data access. For example, when a user attempts to access sensitive data, the system triggers a smart contract to check their access rights and record relevant operation logs.
[0098] The zero-trust access control unit is the core of the security module, which is used to strictly authenticate and manage access rights for each user and device. Unlike the traditional trust boundary model, the zero-trust architecture assumes that all users and devices are untrusted and each access must be verified.
[0099] Specifically, this unit implements access control in the following ways: Multi-factor authentication: Users need to provide multiple identity credentials, such as passwords, biometrics (such as fingerprints, facial recognition), and dynamic tokens when accessing data or resources.
[0100] Access rights classification: Each user's rights are limited to the minimum scope required for their work. The permission set is , the system will check Whether the access requirements are met.
[0101] Dynamic trust evaluation: The system dynamically adjusts the trust level based on the user's behavior pattern and device status. For example, if the user's behavior is abnormal (such as frequent access attempts or from an uncommon IP address), the system will lower their trust level or even temporarily freeze their permissions.
[0102] In one possible implementation, the zero trust unit uses policy-based access control (PBAC) to define user access rights through policy rules. For example: ; in: Represents a user; Represents a resource; Indicates access conditions; Indicates that access is allowed.
[0103] When a user attempts to access a resource, the system dynamically evaluates whether to allow access based on their identity and behavior.
[0104] In some embodiments, in order to further improve security, the security module can combine machine learning technology to detect abnormal user behavior. For example, by constructing a feature vector of the user behavior pattern , the system can quickly identify potential security threats.
[0105] As an option, this module can also use a distributed key management system to store keys on multiple nodes to reduce the risk of single point failure.
[0106] In this embodiment, the security assurance module provides comprehensive protection for the data security of the cloud platform through three functional units: end-to-end encryption, blockchain auditing, and zero-trust access control.
[0107] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A cloud platform system for training an intelligent detection AI model, characterized in that: The system comprises: Multi-objective optimization module, which is used to optimize the resource allocation strategy of the cloud platform to achieve multi-objective optimization of resource utilization, training time and energy consumption; Resource scheduling module, which is used to dynamically schedule computing resources according to the Markov decision process model to ensure that resources are reasonably allocated according to task requirements; Reinforcement learning module, used to automatically learn the optimal resource scheduling strategy through deep Q learning or policy gradient algorithm; Bayesian optimization module, which is used to optimize the hyperparameters in resource scheduling to further improve the scheduling efficiency of the system; Data processing module, used to optimize the transmission and processing of large-scale data, reduce data bottlenecks, and improve data flow efficiency; The security module is used to ensure the privacy and security of data during transmission and prevent data leakage and illegal access.
2. According to claim 1, the intelligent detection AI model training cloud platform system is characterized in that: The multi-objective optimization module includes: An objective function generating unit, used to generate an objective function according to resource utilization, training time and energy consumption of the cloud platform; The optimization calculation unit is used to optimize the objective function through the weighted sum method to solve the optimal resource allocation plan.
3. The intelligent detection AI model training cloud platform system according to claim 2 is characterized in that: The objective function generating unit comprises: A resource utilization calculation unit, used to calculate resource utilization according to resource allocation conditions of each node of the cloud platform; A training time calculation unit is used to estimate the time required to complete the training task; The energy consumption calculation unit is used to calculate the energy consumption under a given resource configuration.
4. The intelligent detection AI model training cloud platform system according to claim 1 is characterized in that: The resource scheduling module includes: A state evaluation unit, used to evaluate the current resource usage state of the cloud platform and generate a current state representation; An action selection unit, used to select resource allocation actions based on current status and historical experience; The state transition control unit is used to calculate and control the transition probability from the current state to the next state.
5. The intelligent detection AI model training cloud platform system according to claim 1 is characterized in that: The Markov decision process model includes: The state space definition unit is used to define the resource allocation state of each computing node in the cloud platform; An action space definition unit, used to define resource adjustment actions and generate possible resource allocation solutions; The transition probability calculation unit is used to calculate the state transition probability based on the current resource state and the action taken.
6. The intelligent detection AI model training cloud platform system according to claim 1 is characterized in that: The reinforcement learning module includes: A deep Q network training unit, used to update the Q value function and train the scheduling strategy through a deep Q learning algorithm; A policy gradient calculation unit, used to optimize the scheduling strategy through the policy gradient method; The reward feedback unit is used to calculate the immediate reward based on the feedback of task execution and optimize the strategy.
7. The intelligent detection AI model training cloud platform system according to claim 1 is characterized in that: The Bayesian optimization module includes: Gaussian process modeling unit, used to model the objective function of the scheduling system and estimate the posterior distribution; The acquisition function calculation unit is used to select the optimal resource allocation strategy according to the posterior distribution of the current resource configuration.
8. The intelligent detection AI model training cloud platform system according to claim 1 is characterized in that: The data processing module comprises: Data storage optimization unit, used to distribute and store training data in different computing nodes of the cloud platform, supporting distributed parallel processing; The data flow control unit is used to automatically adjust the data flow path based on the task priority and data volume to ensure efficient data transmission.
9. The intelligent detection AI model training cloud platform system according to claim 1, characterized in that: The security module comprises: End-to-end encryption unit, used to encrypt data during transmission to ensure data privacy; Blockchain audit unit, which is used to record data access logs to ensure the traceability of the data transmission process; Zero Trust authentication unit for authenticating and controlling access to users and devices.
10. The intelligent detection AI model training cloud platform system according to claim 1, characterized in that: The system further comprises: Resource monitoring unit, which is used to monitor the usage of cloud platform resources in real time and provide visual reports on resource usage; The adaptive optimization unit is used to automatically adjust the resource scheduling strategy according to the platform operation status.
Citation Information
Patent Citations
Cloud edge collaborative element reinforcement learning calculation unloading method based on wide attention mechanism
CN116009990A
Computer vision processing system based on machine learning
CN118115810A
Cloud resource automatic allocation system
CN118363765A
Station area intelligent fusion terminal data processing system based on edge calculation
CN119440800A
Deep-reinforcement-learning-based adaptive efficient resource allocation method for cloud data center
WO2023184939A1