Artificial intelligence acceleration control method, system and device and storage medium
By transmitting the to-process data to the edge device in the computing system for pre-processing, extracting features and analyzing models, and formulating an acceleration adjustment strategy, the problem of high power consumption in computing system when processing complex tasks is solved, and more efficient and lower power consumption computing performance is achieved.
Patent Information
- Application Number
- CN202510197208.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing computing systems consume a lot of power when dealing with complex computing tasks, which limits their promotion and use in a wider range of application scenarios.
By transmitting the pending data to the edge device for pre-processing, key features are extracted using feature extraction algorithms, feature sets are input to the pre-trained analysis model for comparison and analysis, acceleration requirements are obtained, acceleration adjustment strategies are formulated based on the current operating status of the computing device and task priority, artificial intelligence computing parameters are allocated, and acceleration strategies are optimized.
It significantly reduces the power consumption of the computing system, improves the efficiency and accuracy of data processing, optimizes resource utilization, ensures that tasks can be completed faster and more efficiently, and continuously improves system performance.
Smart Images

Figure CN120104330A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a control method, system, device and storage medium for artificial intelligence acceleration. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the demand for high-performance computing and large-scale data processing has increased. Although existing computing systems can meet these needs to a certain extent, there are still major bottlenecks in terms of real-time performance and energy efficiency. Especially in fields such as machine learning and deep learning, the demand for computing resources is higher, and traditional CPU architectures are no longer capable of complex computing tasks.
[0003] The above-mentioned existing technical solutions have the following defects: the existing acceleration technologies have the problem of high power consumption when processing complex computing tasks, which limits their promotion and use in a wider range of application scenarios, so there is room for improvement. Summary of the invention
[0004] In order to reduce the power consumption of a computing system, the present application provides a control method, system, device and storage medium for artificial intelligence acceleration.
[0005] The above-mentioned invention objective of the present application is achieved through the following technical solutions: A control method for artificial intelligence acceleration, the control method for artificial intelligence acceleration comprising: Acquire the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; Using a feature extraction algorithm to extract features from the preprocessed data to obtain a feature set of the data to be processed; Inputting the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements; Acquire the current running state and task priority of the computing device, analyze the acceleration demand based on the current running state and task priority of the computing device, and obtain an acceleration adjustment strategy; Adjusting artificial intelligence computing parameters according to the acceleration adjustment strategy; After the artificial intelligence computing parameters are adjusted, feedback information is collected based on real-time monitoring results and actual performance, and the acceleration adjustment strategy is iteratively optimized based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
[0006] By adopting the above technical solution, the efficiency and accuracy of subsequent data processing can be significantly improved by transmitting the data to be processed to the edge device for preprocessing. This step not only helps to reduce the bandwidth requirements for data transmission, but also prepares neater and more structured data for subsequent feature extraction. By using the feature extraction algorithm to analyze the preprocessed data, key features can be effectively extracted. Through this process, the dimension of the data is greatly compressed while retaining the most valuable information for analysis. By inputting the feature set into the pre-trained model for comparison and analysis, patterns or anomalies in the data can be quickly identified, thereby deriving acceleration requirements, which can help identify potential optimization space or bottlenecks. By analyzing the operating status and task priority of the current computing device, an effective acceleration adjustment strategy can be formulated to ensure the effective allocation of resources and optimize the execution of computing tasks. The allocation of artificial intelligence computing parameters according to the acceleration adjustment strategy can directly improve the computing performance and resource utilization of the system, ensuring that tasks can be completed faster and more efficiently. By collecting real-time monitoring results and feedback information on actual performance, the acceleration adjustment strategy can be continuously optimized, which helps to improve the effectiveness of the acceleration strategy and ensure continuous improvement of system performance.
[0007] In a preferred example, the present application can be further configured as follows: the acquisition of the data to be processed, the transmission of the data to be processed to the edge device for preprocessing, and the acquisition of the preprocessed data to be processed include: Extract data from various data sources in real time using ETL tools, and transform the extracted data to obtain the data to be processed; gRPC is used to transmit the data to be processed to the edge device, and the data to be processed is preprocessed on the edge device to obtain the preprocessed data to be processed, and the preprocessing includes data cleaning, deduplication and format conversion.
[0008] By adopting the above technical solution, data is extracted from various data sources in real time through ETL tools, and the data is transformed to finally obtain the processed data suitable for processing, ensuring the real-time update of the data and the unification of the format, laying the foundation for subsequent analysis; gRPC is used to transmit the processed data to the edge device, and data preprocessing steps are performed on the edge device. The preprocessing steps include data cleaning, deduplication and format conversion to ensure the accuracy and consistency of the data, thereby improving the efficiency and effectiveness of subsequent analysis.
[0009] In a preferred example, the present application may be further configured as follows: the feature extraction algorithm is used to extract features from the pre-processed data to be processed, and the feature set of the data to be processed is obtained, including: Using a multi-level neural network structure to perform data analysis and extraction on the pre-processed data to be processed, identifying the features and patterns of the pre-processed data to be processed, and obtaining a data feature set; The principal component analysis is applied to perform data dimension reduction on the data feature set to reduce the feature dimension and obtain the feature set of the data to be processed.
[0010] By adopting the above technical solution, the pre-processed data is analyzed through a multi-level neural network structure to identify its characteristics and patterns, thereby generating a rich set of data features, ensuring that the most valuable information in the data is extracted, and laying the foundation for further data dimensionality reduction and analysis; by applying principal component analysis to reduce the dimensionality of the data feature set and reducing the dimension of the features, a simplified feature set of the data to be processed is finally obtained, which optimizes the data structure and makes the subsequent data analysis and decision-making process more efficient and accurate.
[0011] In a preferred example, the present application may be further configured as follows: before inputting the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain the acceleration requirement, the control method for artificial intelligence acceleration also includes: Collect and store various types of data and corresponding data processing acceleration strategies, classify and annotate the various types of data and corresponding data processing acceleration strategies, and obtain a comprehensive analysis training set; The analysis model built based on the Transformer architecture is iteratively trained using the comprehensive analysis training set to obtain an iteratively trained analysis model, and the iteratively trained analysis model is fine-tuned using transfer learning technology to obtain the pre-trained analysis model.
[0012] By adopting the above technical solution, by collecting and storing various types of data and their corresponding data processing acceleration strategies, and classifying and labeling these data, a comprehensive analysis training set is obtained, which effectively organizes data resources and provides a solid foundation and high-quality input for subsequent model training; by using the comprehensive analysis training set to iteratively train the analysis model based on the Transformer architecture, the analysis ability and accuracy of the model are improved, ensuring that the model can fully understand and utilize the information in the data, thereby improving its prediction and classification performance; through the transfer learning technology, the analysis model after iterative training is fine-tuned, and finally a pre-trained analysis model is obtained, which enhances the adaptability and efficiency of the model in different application scenarios, enabling it to maintain high performance in a wider range of tasks.
[0013] In a preferred example, the present application can be further configured as follows: the feature set of the data to be processed is input into a pre-trained analysis model for comparison and analysis, and the acceleration requirements are obtained including: Performing similarity calculation on the feature set of the data to be processed and the key feature data in the pre-trained analysis model to obtain a similarity calculation score; When the similarity score reaches a preset threshold, an acceleration scheme corresponding to the key feature data is determined, and the acceleration requirement is obtained according to the acceleration scheme corresponding to the key feature data.
[0014] By adopting the above technical solution, by calculating the similarity between the feature set of the data to be processed and the key feature data in the pre-trained analysis model, a similarity calculation score is obtained, which helps to identify the degree of matching between the data and the model features, and provides a key basis for further selection of acceleration solutions; by determining the acceleration solution corresponding to the key feature data when the similarity score reaches a preset threshold, and then obtaining the corresponding acceleration requirements, it ensures that the optimal acceleration strategy is selected in data processing, improves processing efficiency and performance, and meets the specific needs of different tasks.
[0015] In a preferred example, the present application may be further configured as follows: the obtaining of the operating status and task priority of the current computing device, and analyzing the acceleration demand based on the operating status and task priority of the current computing device to obtain the acceleration adjustment strategy includes: Use the APM tool to monitor the working status of the current computing device in real time, obtain the running status of the current computing device and the real-time status of the current task, and obtain the task priority of the current computing device according to the real-time status of the current task. The running status of the computing device includes CPU / GPU utilization, memory usage, network bandwidth, and power consumption; Based on the operating status and task priority of the current computing device, the acceleration demand is analyzed in combination with a dynamic resource scheduling algorithm and a reinforcement learning algorithm to obtain the acceleration adjustment strategy.
[0016] By adopting the above technical solution, by using the APM tool to monitor the working status of the current computing device in real time, the operating status of the device and the real-time status of the current task can be obtained, ensuring a comprehensive understanding of the device performance and supporting the reasonable evaluation and adjustment of the priority of subsequent tasks; based on the current operating status of the computing device and the task priority, by combining the dynamic resource scheduling algorithm and the reinforcement learning algorithm, an in-depth analysis of the acceleration demand is carried out to obtain an optimized acceleration adjustment strategy, which effectively optimizes resource allocation, improves processing efficiency, and ensures that critical tasks can receive timely resource support and priority processing.
[0017] In a preferred example, the present application can be further configured as follows: the control method for artificial intelligence acceleration also includes: Analyze the feedback information, trigger an emergency strategy when an indicator in the feedback information is abnormal, and generate an alarm signal to transmit to a display terminal; The feedback information and the acceleration adjustment strategy record are stored in a database, and the feedback information and the acceleration adjustment strategy record are analyzed by using big data analysis and machine learning technology to obtain a strategy optimization rule; Collect and store the user's usage habits and behavior patterns, analyze the user's usage habits, the user's behavior patterns and the strategy optimization rules, and predict the changes in acceleration demand in the next stage; According to the change in acceleration demand in the next stage, a user-personalized acceleration adjustment plan is generated, and the user-personalized acceleration adjustment plan is stored in the database.
[0018] By adopting the above technical solution, by analyzing the feedback information, when an indicator is found to be abnormal, the emergency strategy is immediately triggered and an alarm signal is generated and transmitted to the display end, which can effectively monitor the system status and provide timely warnings and processing measures in abnormal situations; by recording the feedback information and acceleration adjustment strategy in the database, and using big data analysis and machine learning technology for analysis, the law of strategy optimization is obtained to provide data and trend support for the continuous optimization of the strategy; by collecting the user's usage habits and behavior patterns, combined with the strategy optimization law obtained by analysis, the acceleration demand changes in the next stage are predicted, helping the system to prepare for resource adjustment in advance to adapt to the user's personalized needs; according to the acceleration demand changes in the next stage, the user's personalized acceleration adjustment plan is generated and stored, so that the system can dynamically adapt to the user's personal needs and optimize resource utilization efficiency.
[0019] The second object of the invention is achieved by the following technical solutions: An artificial intelligence accelerated control system, the artificial intelligence accelerated control system comprising: A data preprocessing module is used to obtain the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; A feature extraction module is used to extract features from the pre-processed data to be processed using a feature extraction algorithm to obtain a feature set of the data to be processed; A model analysis module is used to input the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements; A determination strategy module is used to obtain the current operating state and task priority of the computing device, analyze the acceleration demand based on the current operating state and task priority of the computing device, and obtain an acceleration adjustment strategy; A deployment module, used to deploy artificial intelligence computing parameters according to the acceleration adjustment strategy; The model optimization module is used to collect feedback information based on real-time monitoring results and actual performance after adjusting the artificial intelligence calculation parameters, and iteratively optimize the acceleration adjustment strategy based on the feedback information, wherein the feedback information includes acceleration effect, resource usage, and task completion time.
[0020] By adopting the above technical solution, the efficiency and accuracy of subsequent data processing can be significantly improved by transmitting the data to be processed to the edge device for preprocessing. This step not only helps to reduce the bandwidth requirements for data transmission, but also prepares neater and more structured data for subsequent feature extraction. By using the feature extraction algorithm to analyze the preprocessed data, key features can be effectively extracted. Through this process, the dimension of the data is greatly compressed while retaining the most valuable information for analysis. By inputting the feature set into the pre-trained model for comparison and analysis, patterns or anomalies in the data can be quickly identified, thereby deriving acceleration requirements, which can help identify potential optimization space or bottlenecks. By analyzing the operating status and task priority of the current computing device, an effective acceleration adjustment strategy can be formulated to ensure the effective allocation of resources and optimize the execution of computing tasks. The allocation of artificial intelligence computing parameters according to the acceleration adjustment strategy can directly improve the computing performance and resource utilization of the system, ensuring that tasks can be completed faster and more efficiently. By collecting real-time monitoring results and feedback information on actual performance, the acceleration adjustment strategy can be continuously optimized, which helps to improve the effectiveness of the acceleration strategy and ensure continuous improvement of system performance.
[0021] The third objective of the present application is achieved through the following technical solutions: A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned artificial intelligence accelerated control method when executing the computer program.
[0022] The fourth objective of the present application is achieved through the following technical solutions: A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned artificial intelligence accelerated control method.
[0023] In summary, the present application includes at least one of the following beneficial technical effects: 1. By transferring the data to be processed to the edge device for preprocessing, the efficiency and accuracy of subsequent data processing can be significantly improved. This step not only helps to reduce the bandwidth requirements for data transmission, but also prepares neater and more structured data for subsequent feature extraction. Using feature extraction algorithms to analyze the preprocessed data can effectively extract key features. Through this process, the dimensions of the data are greatly compressed while retaining the most valuable information for analysis. 2. By inputting the feature set into a pre-trained model for comparative analysis, patterns or anomalies in the data can be quickly identified, thereby deriving acceleration requirements and helping to identify potential optimization space or bottlenecks; by analyzing the current operating status and task priority of the computing device, an effective acceleration adjustment strategy can be formulated to ensure the effective allocation of resources and optimize the execution of computing tasks; by adjusting the artificial intelligence computing parameters according to the acceleration adjustment strategy, the computing performance and resource utilization of the system can be directly improved, ensuring that tasks can be completed faster and more efficiently; by collecting real-time monitoring results and feedback information on actual performance, the acceleration adjustment strategy can be continuously optimized, which helps to improve the effectiveness of the acceleration strategy and ensure continuous improvement of system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of an artificial intelligence accelerated control method in one embodiment of the present application; Figure 2 This is a flowchart for implementing step S10 in an artificial intelligence accelerated control method in one embodiment of the present application; Figure 3 This is a flowchart for implementing step S20 in an artificial intelligence accelerated control method in one embodiment of the present application; Figure 4 This is another implementation flow chart of step S30 in an artificial intelligence accelerated control method in one embodiment of the present application; Figure 5 This is a flowchart for implementing step S30 in an artificial intelligence accelerated control method in one embodiment of the present application; Figure 6 This is a flowchart for implementing step S40 in an artificial intelligence accelerated control method in one embodiment of the present application; Figure 7 This is another implementation flow chart of an artificial intelligence accelerated control method in one embodiment of the present application; Figure 8 This is a principle block diagram of an artificial intelligence accelerated control system in one embodiment of the present application; Fig. 9 It is a schematic diagram of a device in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The present application is further described in detail below in conjunction with the accompanying drawings.
[0026] In one embodiment, if Figure 1 As shown, the present application discloses a control method for artificial intelligence acceleration, which specifically includes the following steps: S10: Acquire the data to be processed, and transmit the data to be processed to the edge device for preprocessing to obtain the preprocessed data to be processed.
[0027] Specifically, the system first collects the data to be processed from the source, such as sensors, user devices, network traffic, etc. These data may include video, audio, text or other sensor data, and transmits the collected data to the location at the edge of the network close to the data source, that is, edge devices such as smartphones, IoT devices, edge servers, etc., to reduce latency and bandwidth consumption, and performs preliminary processing on the data on the edge device, such as data cleaning, denoising, format conversion, compression and other operations to improve the efficiency and accuracy of subsequent processing, and finally obtain the pre-processed data to be processed.
[0028] S20: extracting features from the preprocessed data to be processed using a feature extraction algorithm to obtain a feature set of the data to be processed.
[0029] Specifically, feature extraction algorithms are used to extract useful features from preprocessed data. For example, for image data, features such as edges and color histograms can be extracted; for time series data, features such as trends and periodicity can be extracted. The extracted features are organized into a structured feature set, i.e., the feature set of the data to be processed, to facilitate subsequent analysis and modeling.
[0030] S30: Input the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements.
[0031] Specifically, a pre-trained machine learning or deep learning model is used to analyze the feature set. The analysis model infers the input features, identifies data patterns or predicts results. Based on the output of the model, it is determined whether the current task needs to be accelerated. For example, if the analysis model predicts that a task will consume a lot of resources or time, the system may need to be adjusted to optimize performance and ultimately meet the acceleration requirements.
[0032] S40: Obtain the current operating status and task priority of the computing device, analyze the acceleration demand based on the current operating status and task priority of the computing device, and obtain an acceleration adjustment strategy.
[0033] Specifically, the operating status of the computing device is obtained in real time, including CPU / GPU utilization, memory usage, network bandwidth, storage status, etc., the priority of the currently running or pending tasks is identified, and it is determined which tasks need to be prioritized. In combination with the operating status of the device and the task priority, the acceleration requirements are analyzed to analyze which tasks need to be accelerated and how to accelerate them. For example, under high load, high-priority tasks are prioritized, or resource allocation is dynamically adjusted to finally obtain an acceleration adjustment strategy.
[0034] S50: Adjusting artificial intelligence computing parameters according to the acceleration adjustment strategy.
[0035] Specifically, based on the acceleration adjustment strategy, adjust the computing parameters of the AI model, such as adjusting the CPU / GPU usage ratio, allocating more memory or storage resources; adjust the number of layers and nodes of the neural network, or change the model's operating mode, switching from high-precision mode to low-precision mode to increase the speed; rearrange the execution order or concurrency of tasks to optimize resource utilization and task completion time.
[0036] S60: After adjusting the AI computing parameters, feedback information is collected based on real-time monitoring results and actual performance, and the acceleration adjustment strategy is iteratively optimized based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
[0037] Specifically, the performance of the adjusted system is continuously monitored, including indicators such as acceleration effect, resource usage, and task completion time. The above monitoring data is collected as feedback information to evaluate the effectiveness of the adjustment strategy. Based on the feedback information, it is determined whether the current adjustment strategy has achieved the expected effect. If the effect is not good, the strategy parameters are adjusted or new optimization measures are taken. For example, further adjust resource allocation, optimize model parameters, improve task scheduling algorithms, etc. The monitoring, feedback, and optimization cycles are continuously repeated to achieve adaptive optimization of the system.
[0038] By adopting the above technical solution, the efficiency and accuracy of subsequent data processing can be significantly improved by transmitting the data to be processed to the edge device for preprocessing. This step not only helps to reduce the bandwidth requirements for data transmission, but also prepares neater and more structured data for subsequent feature extraction. By using the feature extraction algorithm to analyze the preprocessed data, key features can be effectively extracted. Through this process, the dimension of the data is greatly compressed while retaining the most valuable information for analysis. By inputting the feature set into the pre-trained model for comparison and analysis, patterns or anomalies in the data can be quickly identified, thereby deriving acceleration requirements, which can help identify potential optimization space or bottlenecks. By analyzing the operating status and task priority of the current computing device, an effective acceleration adjustment strategy can be formulated to ensure the effective allocation of resources and optimize the execution of computing tasks. The allocation of artificial intelligence computing parameters according to the acceleration adjustment strategy can directly improve the computing performance and resource utilization of the system, ensuring that tasks can be completed faster and more efficiently. By collecting real-time monitoring results and feedback information on actual performance, the acceleration adjustment strategy can be continuously optimized, which helps to improve the effectiveness of the acceleration strategy and ensure continuous improvement of system performance.
[0039] In one embodiment, if Figure 2 As shown, in step S10, the data to be processed is obtained, and the data to be processed is transmitted to the edge device for preprocessing to obtain the preprocessed data to be processed, which specifically includes: S11: Use ETL tools to extract data from various data sources in real time, and transform the extracted data to obtain data to be processed.
[0040] Specifically, ETL tools are used to extract data from various data sources in real time. ETL tools connect and access various data sources, continuously and with low latency obtain the latest data from the data sources, and transform the extracted data. During the transformation process, data from different data sources are integrated, the data format and structure are unified, and the data is converted into predefined formats and standards to obtain the data to be processed.
[0041] S12: gRPC is used to transmit the data to be processed to the edge device, and the data to be processed is preprocessed on the edge device to obtain the preprocessed data to be processed. The preprocessing includes data cleaning, deduplication and format conversion.
[0042] Specifically, first use Protocol Buffers to write a .proto file to define the data transmission interface and the message format of the data to be transmitted. The ETL tool establishes a connection with the edge device through gRPC, transmits the data to be processed in real time, and preprocesses the data to be processed on the edge device. The preprocessing includes data cleaning, deduplication, and format conversion to obtain the preprocessed data to be processed.
[0043] In one embodiment, if Figure 3 As shown, in step S20, feature extraction is performed on the pre-processed data to be processed using a feature extraction algorithm to obtain a feature set of the data to be processed, which specifically includes: S21: Use a multi-level neural network structure to perform data analysis and extraction on the preprocessed data to be processed, identify the features and patterns of the preprocessed data to be processed, and obtain a data feature set.
[0044] Specifically, a multi-level neural network structure is used to perform data analysis and extraction on the preprocessed data to be processed, and the preprocessed data to be processed is passed through a specific feature extraction layer to identify the features and patterns of the preprocessed data to be processed, and finally the extracted data is integrated to obtain a data feature set.
[0045] S22: Apply principal component analysis to perform data dimensionality reduction on the data feature set, reduce the feature dimension, and obtain a feature set of the data to be processed.
[0046] Specifically, principal component analysis is applied to reduce the dimensionality of the data feature set, calculate the covariance matrix of the data, perform eigendecomposition on the covariance matrix to obtain a set of eigenvalues and corresponding eigenvectors, sort the eigenvectors according to the size of the eigenvalues, select the eigenvectors corresponding to the first k largest eigenvalues, use the selected principal component eigenvectors as column vectors, construct a projection matrix W, and map the original data set to the new feature space through the projection matrix W, that is, calculate the linear combination of the original data and these principal components, and finally obtain the feature set of the data to be processed.
[0047] In one embodiment, if Figure 4 As shown, before step S30, that is, before inputting the feature set of the data to be processed into the pre-trained analysis model for comparison and analysis to obtain the acceleration requirement, the control method for artificial intelligence acceleration also includes: S301: Collect and store various types of data and corresponding data processing acceleration strategies, classify and label various types of data and corresponding data processing acceleration strategies, and obtain a comprehensive analysis training set.
[0048] Specifically, various types of data are collected from internal databases, public data sets, partners or web crawlers, and detailed information on various data processing acceleration strategies is collected through literature research, technical documents, industry reports, and expert interviews. The collected data of various types and the corresponding data processing acceleration strategies are stored in the database, and the data are classified according to dimensions such as data type, application scenario, and structure. The data and acceleration strategies are labeled using natural language processing technology and rule engines, and the classified and labeled data are paired with the corresponding acceleration strategies to form training samples. Each sample should contain information such as data type, data characteristics, acceleration strategy characteristics, and their effects.
[0049] S302: Iteratively train the analysis model built based on the Transformer architecture using the comprehensive analysis training set to obtain an iteratively trained analysis model, and fine-tune the iteratively trained analysis model using transfer learning technology to obtain a pre-trained analysis model.
[0050] Specifically, select a Transformer variant model suitable for the task, such as BERT, GPT, RoBERTa, etc., or design a customized Transformer architecture according to specific needs, use a comprehensive analysis training set to iteratively train the analysis model built on the Transformer architecture, use a gradient descent optimization algorithm, perform multiple rounds of training, monitor the training loss and evaluation indicators, and finally obtain the iteratively trained analysis model, and fine-tune the iteratively trained analysis model through transfer learning technology, freeze certain layers of the pre-trained model as needed, retain the high layers for feature extraction of the target task, add a layer suitable for the target task on top of the pre-trained model, perform fine-tuning training on the training set of the target task, adjust the parameters of the pre-trained model to adapt to the new task, and finally obtain the pre-trained analysis model.
[0051] In one embodiment, if Figure 5 As shown, in step S30, the feature set of the data to be processed is input into a pre-trained analysis model for comparison and analysis to obtain the acceleration requirements, which specifically include: S31: Calculate the similarity between the feature set of the data to be processed and the key feature data in the pre-trained analysis model to obtain a similarity calculation score.
[0052] Specifically, the similarity between the feature set of the data to be processed and the key feature data in the pre-trained analysis model is calculated. According to the selected similarity measurement method, the similarity score between the feature set of the data to be processed and the key feature data of the model is calculated, such as cosine similarity, Euclidean distance, Manhattan distance, etc., and finally the similarity calculation score is obtained.
[0053] S32: When the similarity score reaches a preset threshold, an acceleration scheme corresponding to the key feature data is determined, and an acceleration requirement is obtained according to the acceleration scheme corresponding to the key feature data.
[0054] Specifically, various acceleration schemes are planned and stored in advance, and each scheme corresponds to a specific combination of key features. For example, different hardware optimization strategies, such as GPU acceleration, parallel computing, quantized models, etc., can be used as different acceleration schemes. When the similarity score reaches a preset threshold, the selection of the acceleration scheme is triggered. According to the matched key feature combination, the most suitable acceleration scheme is selected. The specific content and requirements of the acceleration scheme are understood, such as the required hardware resources, optimization methods, adjustment parameters, etc. Based on the acceleration scheme, the specific acceleration requirements are clarified, and finally the acceleration requirements are obtained.
[0055] In one embodiment, if Figure 6 As shown, in step S40, the operating state and task priority of the current computing device are obtained, and based on the operating state and task priority of the current computing device, the acceleration demand is analyzed to obtain an acceleration adjustment strategy, which specifically includes: S41: Use the APM tool to monitor the working status of the current computing device in real time, obtain the running status of the current computing device and the real-time status of the current task, and obtain the task priority of the current computing device based on the real-time status of the current task. The running status of the computing device includes CPU / GPU utilization, memory usage, network bandwidth, and power consumption.
[0056] Specifically, use APM tools to monitor the working status of the current computing device in real time, track the currently running tasks, obtain the status, resource consumption and execution progress of each task, obtain the current operating status of the computing device and the real-time status of the current task, define priority standards according to the importance, urgency, resource requirements and business rules of the task, and use the priority scheduling algorithm to calculate the priority of each task based on the real-time task status and predefined standards. The operating status of the computing device includes CPU / GPU utilization, memory usage, network bandwidth, and power consumption.
[0057] S42: Based on the current operating status and task priority of the computing device, the acceleration demand is analyzed in combination with the dynamic resource scheduling algorithm and the reinforcement learning algorithm to obtain an acceleration adjustment strategy.
[0058] Specifically, the dynamic resource scheduling algorithm dynamically adjusts resource allocation according to the current device operating status and task priority. For example, preemptive scheduling or non-preemptive scheduling can be used to optimize resource utilization and task response time. The reinforcement learning algorithm adjusts the strategy according to the system status and action results through continuous trial and error. Combining the results of the dynamic resource scheduling algorithm and the reinforcement learning algorithm, the system can develop a set of acceleration adjustment strategies to optimize resource allocation, namely, the acceleration adjustment strategy.
[0059] In one embodiment, if Figure 7 As shown, the artificial intelligence accelerated control method also includes: S70: Analyze the feedback information, and when an indicator in the feedback information is abnormal, trigger an emergency strategy, and generate an alarm signal to be transmitted to the display end.
[0060] Specifically, the system continuously collects various types of feedback information, uses monitoring tools or custom algorithms to perform real-time or periodic analysis on the collected feedback information, sets thresholds or uses machine learning models to identify abnormal indicators, such as a sudden surge in CPU usage, prolonged response time, increased error rate, etc. When abnormal indicators are detected, the system automatically calls predefined emergency strategies, such as automatic resource expansion, service restart, traffic limiting, etc., to respond to emergencies. After the emergency strategy is triggered, the system generates an alarm signal and transmits the alarm information to the relevant display terminal via email, SMS, dashboard display or other notification methods to ensure that the operation and maintenance personnel respond in a timely manner.
[0061] S80: The feedback information and the acceleration adjustment strategy records are stored in a database, and the feedback information and the acceleration adjustment strategy records are analyzed using big data analysis and machine learning technology to obtain strategy optimization rules.
[0062] Specifically, all feedback information and triggered acceleration adjustment strategies are recorded in detail and stored in a database. Big data analysis and machine learning techniques are used to analyze the feedback information and acceleration adjustment strategy records, analyze the relationship between the feedback information and the adjustment strategy, and discover optimization rules. For example, regression analysis, decision trees, neural networks and other models are used to identify which strategies are most effective under what circumstances, and analyze the strategy optimization rules.
[0063] S90: Collect and store users’ usage habits and behavior patterns, analyze users’ usage habits, behavior patterns, and strategy optimization rules, and predict changes in acceleration demand in the next stage.
[0064] Specifically, through user behavior tracking tools or custom logging systems, we collect users' usage habits and behavior patterns, such as user access frequency, usage duration, operation paths, function usage preferences, etc., and use time series analysis, regression models, deep learning and other machine learning technologies, combined with analysis of user usage habits, user behavior patterns and strategy optimization rules, to predict the next stage of acceleration demand changes.
[0065] S100: Generate a user-personalized acceleration adjustment plan according to the change in acceleration demand in the next stage, and store the user-personalized acceleration adjustment plan in a database.
[0066] Specifically, based on the predicted changes in acceleration demand and combined with the user's usage habits and behavior patterns, a specific acceleration adjustment plan is generated, such as adjusting resource allocation, optimizing network bandwidth, adjusting the number of service instances, etc., and the user's personalized acceleration adjustment plan is stored in the database.
[0067] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0068] In one embodiment, a control system for artificial intelligence acceleration is provided, and the control system for artificial intelligence acceleration corresponds one-to-one to the control method for artificial intelligence acceleration in the above embodiment. Figure 8 As shown, the artificial intelligence accelerated control system includes a data preprocessing module, a feature extraction module, a model analysis module, a strategy determination module, a deployment module and a model optimization module. The detailed description of each functional module is as follows: A data preprocessing module is used to obtain the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; A feature extraction module is used to extract features from the pre-processed data to be processed using a feature extraction algorithm to obtain a feature set of the data to be processed; The model analysis module is used to input the feature set of the data to be processed into the pre-trained analysis model for comparison and analysis to obtain the acceleration requirements; A determination strategy module is used to obtain the current operating status and task priority of the computing device, analyze the acceleration demand based on the current operating status and task priority of the computing device, and obtain an acceleration adjustment strategy; A deployment module is used to deploy AI computing parameters according to the acceleration adjustment strategy; The model optimization module is used to collect feedback information based on real-time monitoring results and actual performance after adjusting the artificial intelligence computing parameters, and iteratively optimize the acceleration adjustment strategy based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
[0069] Optionally, the data preprocessing module includes: The data conversion submodule is used to extract data from various data sources in real time using ETL tools and convert the extracted data to obtain data to be processed; The preprocessing submodule is used to transmit the data to be processed to the edge device using gRPC, and preprocess the data to be processed on the edge device to obtain the preprocessed data to be processed. The preprocessing includes data cleaning, deduplication and format conversion.
[0070] Optionally, the feature extraction module includes: The analysis and extraction submodule is used to use a multi-level neural network structure to perform data analysis and extraction on the pre-processed data to be processed, identify the characteristics and patterns of the pre-processed data to be processed, and obtain a data feature set; The dimensionality reduction submodule is used to apply principal component analysis to perform data dimensionality reduction on the data feature set, reduce the feature dimension, and obtain the feature set of the data to be processed.
[0071] Optionally, the model analysis module includes: The similarity calculation submodule is used to calculate the similarity between the feature set of the data to be processed and the key feature data in the pre-trained analysis model to obtain a similarity calculation score; The demand determination submodule is used to determine the acceleration scheme corresponding to the key feature data when the similarity score reaches a preset threshold, and obtain the acceleration demand based on the acceleration scheme corresponding to the key feature data.
[0072] Optionally, determining the strategy module includes: The monitoring submodule is used to use the APM tool to monitor the working status of the current computing device in real time, obtain the running status of the current computing device and the real-time status of the current task, and obtain the task priority of the current computing device according to the real-time status of the current task. The running status of the computing device includes CPU / GPU utilization, memory usage, network bandwidth, and power consumption; The combined analysis submodule is used to analyze the acceleration demand based on the current operating status and task priority of the computing device, combined with the dynamic resource scheduling algorithm and the reinforcement learning algorithm, to obtain the acceleration adjustment strategy.
[0073] Optionally, the artificial intelligence accelerated control system further includes: A training set acquisition module is used to collect and store various types of data and corresponding data processing acceleration strategies, classify and annotate various types of data and corresponding data processing acceleration strategies, and obtain a comprehensive analysis training set; Iterative training module, used to iteratively train the analysis model built based on the Transformer architecture using the comprehensive analysis training set to obtain the iteratively trained analysis model, and fine-tune the iteratively trained analysis model through transfer learning technology to obtain a pre-trained analysis model; The feedback analysis module is used to analyze the feedback information. When an indicator in the feedback information is abnormal, an emergency strategy is triggered and an alarm signal is generated and transmitted to the display end; A storage module is used to store the feedback information and the acceleration adjustment strategy records in a database, and analyze the feedback information and the acceleration adjustment strategy records using big data analysis and machine learning technology to obtain strategy optimization rules; The prediction module is used to collect and store users' usage habits and behavior patterns, and analyze users' usage habits, user behavior patterns, and strategy optimization rules to predict the changes in acceleration demand in the next stage; The personalized solution generation module is used to generate a user-personalized acceleration adjustment solution according to the acceleration demand changes in the next stage, and store the user-personalized acceleration adjustment solution in a database.
[0074] For the specific definition of an artificial intelligence accelerated control system, please refer to the definition of an artificial intelligence accelerated control method above, which will not be repeated here. Each module in the above-mentioned artificial intelligence accelerated control system can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0075] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store feedback information and acceleration adjustment strategies. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a control method for artificial intelligence acceleration is implemented.
[0076] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: Acquire the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; Using a feature extraction algorithm to extract features from the preprocessed data to obtain a feature set of the data to be processed; Input the feature set of the data to be processed into the pre-trained analysis model for comparison and analysis to obtain the acceleration requirements; Obtaining the current running status and task priority of the computing device, analyzing the acceleration demand based on the current running status and task priority of the computing device, and obtaining an acceleration adjustment strategy; Adjust AI computing parameters based on acceleration strategies; After adjusting the AI computing parameters, feedback information is collected based on real-time monitoring results and actual performance, and the acceleration adjustment strategy is iteratively optimized based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
[0077] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Acquire the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; Using a feature extraction algorithm to extract features from the preprocessed data to obtain a feature set of the data to be processed; Input the feature set of the data to be processed into the pre-trained analysis model for comparison and analysis to obtain the acceleration requirements; Obtaining the current running status and task priority of the computing device, analyzing the acceleration demand based on the current running status and task priority of the computing device, and obtaining an acceleration adjustment strategy; Adjust AI computing parameters based on acceleration strategies; After adjusting the AI computing parameters, feedback information is collected based on real-time monitoring results and actual performance, and the acceleration adjustment strategy is iteratively optimized based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
[0078] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0079] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0080] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A control method for artificial intelligence acceleration, characterized in that: The control method for artificial intelligence acceleration includes: Acquire the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; Using a feature extraction algorithm to extract features from the preprocessed data to obtain a feature set of the data to be processed; Inputting the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements; Acquire the current running state and task priority of the computing device, analyze the acceleration demand based on the current running state and task priority of the computing device, and obtain an acceleration adjustment strategy; Adjusting artificial intelligence computing parameters according to the acceleration adjustment strategy; After the artificial intelligence computing parameters are adjusted, feedback information is collected based on real-time monitoring results and actual performance, and the acceleration adjustment strategy is iteratively optimized based on the feedback information. The feedback information includes acceleration effect, resource usage, and task completion time.
2. The control method for artificial intelligence acceleration according to claim 1, characterized in that: The obtaining of the data to be processed and transmitting the data to be processed to the edge device for preprocessing to obtain the preprocessed data to be processed includes: Extract data from various data sources in real time using ETL tools, and transform the extracted data to obtain the data to be processed; gRPC is used to transmit the data to be processed to the edge device, and the data to be processed is preprocessed on the edge device to obtain the preprocessed data to be processed, and the preprocessing includes data cleaning, deduplication and format conversion.
3. The control method for artificial intelligence acceleration according to claim 1, characterized in that: The feature extraction algorithm is used to extract features from the preprocessed data to obtain a feature set of the data to be processed, which includes: Using a multi-level neural network structure to perform data analysis and extraction on the pre-processed data to be processed, identifying the features and patterns of the pre-processed data to be processed, and obtaining a data feature set; The principal component analysis is applied to perform data dimension reduction on the data feature set to reduce the feature dimension and obtain the feature set of the data to be processed.
4. The artificial intelligence acceleration control method according to claim 1, characterized in that: Before inputting the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements, the control method for artificial intelligence acceleration also includes: Collect and store various types of data and corresponding data processing acceleration strategies, classify and annotate the various types of data and corresponding data processing acceleration strategies, and obtain a comprehensive analysis training set; The analysis model built based on the Transformer architecture is iteratively trained using the comprehensive analysis training set to obtain an iteratively trained analysis model, and the iteratively trained analysis model is fine-tuned using transfer learning technology to obtain the pre-trained analysis model.
5. The artificial intelligence acceleration control method according to claim 1, characterized in that: The feature set of the data to be processed is input into a pre-trained analysis model for comparison and analysis to obtain the acceleration requirements, including: Performing similarity calculation on the feature set of the data to be processed and the key feature data in the pre-trained analysis model to obtain a similarity calculation score; When the similarity score reaches a preset threshold, an acceleration scheme corresponding to the key feature data is determined, and the acceleration requirement is obtained according to the acceleration scheme corresponding to the key feature data.
6. The artificial intelligence acceleration control method according to claim 1, characterized in that: The obtaining of the current operating state and task priority of the computing device, and analyzing the acceleration demand based on the current operating state and task priority of the computing device to obtain the acceleration adjustment strategy includes: Use the APM tool to monitor the working status of the current computing device in real time, obtain the running status of the current computing device and the real-time status of the current task, and obtain the task priority of the current computing device according to the real-time status of the current task. The running status of the computing device includes CPU / GPU utilization, memory usage, network bandwidth, and power consumption; Based on the operating status and task priority of the current computing device, the acceleration demand is analyzed in combination with a dynamic resource scheduling algorithm and a reinforcement learning algorithm to obtain the acceleration adjustment strategy.
7. The artificial intelligence acceleration control method according to claim 1, characterized in that: The artificial intelligence accelerated control method further includes: Analyze the feedback information, trigger an emergency strategy when an indicator in the feedback information is abnormal, and generate an alarm signal to transmit to a display terminal; The feedback information and the acceleration adjustment strategy record are stored in a database, and the feedback information and the acceleration adjustment strategy record are analyzed by using big data analysis and machine learning technology to obtain a strategy optimization rule; Collect and store the user's usage habits and behavior patterns, analyze the user's usage habits, the user's behavior patterns and the strategy optimization rules, and predict the changes in acceleration demand in the next stage; According to the change in acceleration demand in the next stage, a user-personalized acceleration adjustment plan is generated, and the user-personalized acceleration adjustment plan is stored in the database.
8. An artificial intelligence accelerated control system, characterized in that: The artificial intelligence accelerated control system comprises: A data preprocessing module is used to obtain the data to be processed, transmit the data to be processed to the edge device for preprocessing, and obtain the preprocessed data to be processed; A feature extraction module is used to extract features from the pre-processed data to be processed using a feature extraction algorithm to obtain a feature set of the data to be processed; A model analysis module is used to input the feature set of the data to be processed into a pre-trained analysis model for comparison and analysis to obtain acceleration requirements; A determination strategy module is used to obtain the current operating state and task priority of the computing device, analyze the acceleration demand based on the current operating state and task priority of the computing device, and obtain an acceleration adjustment strategy; A deployment module, used to deploy artificial intelligence computing parameters according to the acceleration adjustment strategy; The model optimization module is used to collect feedback information based on real-time monitoring results and actual performance after adjusting the artificial intelligence calculation parameters, and iteratively optimize the acceleration adjustment strategy based on the feedback information, wherein the feedback information includes acceleration effect, resource usage, and task completion time.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the artificial intelligence accelerated control method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of an artificial intelligence accelerated control method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
AI network acceleration method of heterogeneous network and secure path decision-making system
CN120512726A
An ai network acceleration method and a secure path decision system of a heterogeneous network
CN120512726B