AI model training method and device, equipment and storage medium
By selecting some model layers from the large language model to process training samples and adjust parameters, the problem of long computing time is solved and faster training and inference speeds are achieved.
Patent Information
- Application Number
- CN202410302011.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-19
AI Technical Summary
The existing large language models take too long to compute during inference, and there is a need for acceleration.
Select M model layers from the N model layers of the AI model, skip the remaining model layers to process training and prediction samples, and adjust the model parameters based on the processing results.
It speeds up the processing of training samples while ensuring the stability and efficiency of the model's reasoning capabilities.
Smart Images

Figure CN120670539A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a training method, apparatus, device, and storage medium for an AI (Artificial Intelligence) model. Background Art
[0002] With the continuous development of science and technology, research in natural language processing (NLP) has attracted considerable attention. Large language models, in particular, have attracted widespread attention across various fields. As large language models are developed, their reasoning capabilities are becoming increasingly powerful.
[0003] In related technologies, as the reasoning capabilities of large language models become increasingly powerful, the computational time required to perform reasoning using these models is also increasing. Therefore, there is a need to accelerate the reasoning speed of large language models. Summary of the Invention
[0004] The present application provides an AI model training method, apparatus, device, and storage medium. The technical solutions provided by the present application are as follows:
[0005] According to one aspect of an embodiment of the present application, a method for training an AI model is provided, the method comprising:
[0006] Obtain a training sample set of the AI model, where the AI model includes N model layers, the training sample set includes multiple training samples, and N is an integer greater than or equal to 3;
[0007] Determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the training samples, and M is an integer greater than or equal to 2 and less than N;
[0008] Using the M model layers retained in the AI model, each of the training samples is processed separately to obtain a processing result;
[0009] The parameters of the AI model are adjusted based on the processing results to obtain a trained AI model.
[0010] According to one aspect of an embodiment of the present application, a prediction method of an AI model is provided, the method comprising:
[0011] Obtain a prediction sample set of the AI model, where the AI model includes N model layers, the prediction sample set includes a plurality of prediction samples, and N is an integer greater than or equal to 3;
[0012] Determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the prediction samples, and M is an integer greater than or equal to 2 and less than N;
[0013] The M model layers retained in the AI model are used to process each of the prediction samples separately to obtain a processing result.
[0014] According to one aspect of an embodiment of the present application, a training device for an AI model is provided, the device comprising:
[0015] A training sample acquisition module, configured to acquire a training sample set for the AI model, wherein the AI model includes N model layers, and the training sample set includes a plurality of training samples, where N is an integer greater than or equal to 3;
[0016] A training model layer determination module, configured to determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the training samples, and M is an integer greater than or equal to 2 and less than N;
[0017] A training sample processing module, configured to process each of the training samples separately using the M model layers retained in the AI model to obtain a processing result;
[0018] A model adjustment module is used to adjust the parameters of the AI model based on the processing results to obtain a trained AI model.
[0019] According to one aspect of an embodiment of the present application, a prediction device for an AI model is provided, the device comprising:
[0020] A prediction sample acquisition module, configured to acquire a prediction sample set of the AI model, wherein the AI model includes N model layers, and the prediction sample set includes a plurality of prediction samples, where N is an integer greater than or equal to 3;
[0021] A prediction model layer determination module, configured to determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the prediction samples, and M is an integer greater than or equal to 2 and less than N;
[0022] The prediction sample processing module is used to use the M model layers retained in the AI model to process each of the prediction samples separately to obtain a processing result.
[0023] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned AI model training method or AI model prediction method.
[0024] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned AI model training method or AI model prediction method.
[0025] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned AI model training method or AI model prediction method.
[0026] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:
[0027] By determining the retained M model layers from the N model layers of the AI model, and then using the retained M model layers to process multiple training samples, corresponding processing results are obtained, where M is less than N. Finally, the parameters of the above-mentioned AI model are adjusted based on the processing results to obtain the trained AI model; since some model layers in the AI model are skipped, the processing speed of the training samples is accelerated, and it is ensured that the retained model layers in the trained AI model are used to process other samples while still having good reasoning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;
[0029] Figure 2 This is a flowchart of a method for training an AI model provided by one embodiment of the present application;
[0030] Figure 3 This is a flowchart of determining a retained model layer provided by an embodiment of the present application;
[0031] Figure 4 This is a comparison chart of the computational depth of inference timing of different methods provided in one embodiment of the present application;
[0032] Figure 5 This is a flowchart of a prediction method of an AI model provided by one embodiment of the present application;
[0033] Figure 6 This is a block diagram of an AI model training device provided by one embodiment of the present application;
[0034] Figure 7 This is a block diagram of a prediction device for an AI model provided by one embodiment of the present application;
[0035] Figure 8 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0037] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0038] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0039] Natural language processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; it also involves important technologies for model training in the fields of computer science, mathematics, and artificial intelligence. The pre-trained model is developed from the large language model (LLM) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.
[0040] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration. Pretrained models are the latest development in deep learning, integrating these techniques.
[0041] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (Artificial Intelligence Generated Content, referred to as AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0042] A pre-training model (PTM), also known as a cornerstone model or a large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on massive amounts of unlabeled data, leveraging the function approximation capabilities of large-parameter DNNs to enable the PTM to extract common features from the data. Through fine-tuning, parameter-efficient fine-tuning (PEFT), prompt-tuning, and other techniques, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be divided into language models, visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multimodal models (ViBERT, CLIP, Flamingo, Gato), etc., according to the data modality processed. Among them, a multimodal model refers to a model that establishes feature representations of two or more data modalities. Pre-training models are important tools for outputting AI-generated content, and can also serve as a general interface for connecting multiple specific task models.
[0043] Distributed training refers to splitting the workload of training models and sharing them among multiple microprocessors. Large models have large parameters and training data, which exceeds the capacity of a single machine, so distributed parallelism is needed to speed up. Parallel mechanisms include data parallelism (DP), model parallelism (MP), pipeline parallelism (PP), and hybrid parallelism (HP). Structural designs include those based on parameter servers, reduce protocols, and message passing interfaces (MPI).
[0044] Model compression and quantization uses compression and quantization techniques to reduce model size and accelerate model inference, thereby lowering model storage and computational costs. Model compression typically includes pruning, low-rank decomposition, and knowledge distillation. Model quantization converts floating-point parameters in the model to fixed-point or integer parameters, reducing model size and accelerating model inference.
[0045] Model parallel computing involves distributing model computational tasks across multiple computing devices (such as central processing units, graphics processing units, and tensor processing units) to perform computations simultaneously, thereby accelerating model training and inference. Model parallel computing can effectively utilize computing resources, improving model computational efficiency and training speed.
[0046] The solutions provided in the embodiments of this application involve technologies such as natural language processing and machine learning of artificial intelligence, which are specifically introduced and explained through the following embodiments.
[0047] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment may include: a model training device 10 and a model use device 20.
[0048] The model training device 10 is an electronic device with data calculation, processing and storage functions. The model training device 10 can be either a terminal device or a server. The model training device 10 can include but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality (AR) devices, virtual reality (VR) devices and other electronic devices. The model training device 10 is used to train the AI model.
[0049] In the present application, an AI model is an AI model having multiple model layers. For example, the AI model may be a large language model, or a model constructed based on a convolutional neural network. Optionally, the model training device 10 may use a machine learning method to train the above-mentioned AI model, so that the trained AI model has higher computational efficiency than the pre-trained AI model while ensuring that there is no significant loss.
[0050] In some embodiments, the training process of the AI model is as follows (this is only a brief description, please refer to the following embodiments for the specific training process): obtain a training sample set for AI model training, determine the retained model layer from the multiple model layers of the AI model, and use the above-mentioned retained model layer to process each training sample in the training sample set. During the processing, skip the execution of the multiple model layers of the above-mentioned AI model except the retained model layer to obtain a processing result, and adjust the parameters of the AI model based on the processing result to obtain a trained AI model.
[0051] In some embodiments, the prediction process of the AI model is as follows (this is only a brief description, please refer to the following embodiments for the specific training process): obtain a prediction sample set for prediction, determine the retained model layer from the multiple model layers of the trained AI model, and use the retained model layer in the trained AI model to process each prediction sample in the prediction sample set separately. During the processing, skip the execution of the multiple model layers of the trained AI model except the retained model layer to obtain the processing result.
[0052] The model-using device 20 is an electronic device with data calculation, processing and storage functions. The model-using device 20 can be either a terminal device or a server. The model-using device 20 can include but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality devices, virtual reality devices, cloud technology platforms, intelligent robots, smart transportation terminal systems, and driving control systems. The model-using device 20 uses the trained AI model to process multiple prediction samples. When processing each prediction sample, the retained M model layers are activated to process the above-mentioned multiple prediction samples to obtain processing results of the multiple prediction samples.
[0053] The model training device 10 and the model using device 20 can be two independent devices or the same device.
[0054] For example, the server mentioned above can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but is not limited to these.
[0055] Please refer to Figure 2 , which shows a flow chart of the AI model training method provided by one embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be Figure 1 The model training device 10 in the solution implementation environment shown. The method may include at least one of the following steps (210-240):
[0056] Step 210: Obtain a training sample set of an AI model, where the AI model includes N model layers, and the training sample set includes multiple training samples, where N is an integer greater than or equal to 3.
[0057] In some embodiments, the above-mentioned AI model is a neural network model for specific tasks built using machine learning, deep learning and other technologies, and can be applied to various fields, such as natural language processing, computer vision, speech recognition, etc.
[0058] In some embodiments, the above-mentioned AI model may include but is not limited to the following models: a model based on the Transform architecture, a model based on the Convolutional Neural Network (CNN), a model based on the Recurrent Neural Network (RNN), etc.
[0059] For example, in the field of natural language processing, the AI model can be an AI model based on the Transform architecture, such as a large language model, which is used to process natural language processing tasks such as text generation, text classification, and semantic understanding.
[0060] For example, in the field of computer vision, the AI model can be an AI model based on a convolutional neural network, such as a target detection model, which is used to process computer vision tasks such as target recognition, image classification, semantic segmentation, and face recognition.
[0061] For example, in the field of speech recognition, the AI model can be an AI model based on a recurrent neural network, such as a speech recognition model, which is used to process sequence data tasks such as time series analysis, speech recognition, and text analysis.
[0062] Model layers are components of a neural network model, used to process input data and generate output data. Model layers may include, but are not limited to, the following types: input layer, fully connected layer, convolutional layer, pooling layer, recurrent layer, embedding layer, normalization layer, activation function layer, attention mechanism layer, activation function layer, output layer, etc. Alternatively, skilled artisans can compose a neural network model based on different model layers as needed.
[0063] Exemplarily, the AI model based on the Transform architecture includes an input embedding layer (Input Embedding), a multi-head attention mechanism layer (Multi-Head Attention), a masked multi-head attention mechanism layer (Masked Multi-Head Attention), a normalization layer (Add&Norm), a feed forward layer (Feed Forward), a linear layer (Linear), an activation function layer (Softmax) and an output embedding layer (Output Embedding).
[0064] A training sample set refers to a sample set of training samples used to train an AI model. Optionally, a training sample includes: sample data and label data. For example, when an AI model performs supervised learning, each training sample used includes sample data and corresponding label data. The AI model learns the relationship between the sample data and the corresponding label data, enabling the AI model to better predict other input data. Optionally, a training sample includes sample data. For example, when an AI model performs unsupervised learning, each training sample used only includes sample data. The AI model learns the potential structure in the sample data, enabling the AI model to better predict other input data based on the learned potential structure.
[0065] Step 220: Determine the retained M model layers from the N model layers, wherein the remaining model layers except the M model layers in the N model layers are skipped when the AI model is used to process each training sample, and M is an integer greater than or equal to 2 and less than N.
[0066] In related technologies, when a complete AI model is used to process training samples, after the sample data is input into the AI model, the sample data flows based on the composition structure of its model layer, that is, the sample data is transmitted between each model layer according to the composition structure of the model layer.
[0067] In this application, when an AI model is used to process each training sample, for any training sample, after the sample data is input into the AI model, the AI model only activates the above-mentioned retained M model layers to process the training sample, that is, when the sample data flows between the N model layers of the AI model, if the sample data is located in one of the above-mentioned retained M model layers, the model layer processes the input data and passes the intermediate result to the next model layer; if the sample data is located in the remaining model layers other than the M model layers in the N model layers, the model layer does not process the input data and passes the data to the next model layer. It should be noted that the above-mentioned "input data" refers to the data passed to the current model layer after being processed by the previous model layer.
[0068] In some embodiments, the above-mentioned M model layers are evenly distributed in the AI model.
[0069] In this application, uniform distribution means that the above-mentioned M model layers are relatively evenly distributed in the AI model, that is, the above-mentioned M model layers will not be concentrated in a certain part of the AI model.
[0070] Through the above method, since different model layers have different functions and importance, by making the distribution of the model layers retained in the AI model relatively balanced, the negative impact of skipping the execution of the remaining model layers on the AI model can be minimized.
[0071] In some embodiments, the number of retained model layers M is determined based on the total number of model layers N of the AI model and a pre-set acceleration ratio, where the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model; based on the number of retained model layers M, M model layers are selected from the N model layers as the retained M model layers.
[0072] In some embodiments, when the AI model is a model based on the Transform architecture, the total number of model layers N of the AI model refers to the number of attention mechanism layers included in the AI model.
[0073] For example, if the AI model is a model based on the Transform architecture and includes 12 attention mechanism layers, the total number of model layers N of the AI model is 12.
[0074] In some embodiments, when the AI model is a convolutional neural network-based model, the total number of model layers N of the AI model refers to the number of convolutional layers included in the AI model.
[0075] For example, if the AI model is a convolutional neural network-based model and includes 18 convolutional layers, the total number N of convolutional layers of the AI model is 18.
[0076] In some embodiments, when the AI model is a model based on a recurrent neural network, the total number of model layers of the AI model refers to the number of recurrent layers included in the AI model.
[0077] For example, if the AI model is a model based on a recurrent neural network and includes 5 recurrent layers, the total number of model layers N of the AI model is 5.
[0078] In some embodiments, the preset acceleration ratio, which may also be referred to as a target acceleration ratio, is a positive integer greater than or equal to 1. Optionally, the preset acceleration ratio is set by relevant technicians according to needs, and this application does not limit this.
[0079] It should be noted that the greater the acceleration ratio, the fewer model layers are retained, and the greater the negative impact on the AI model due to skipping the remaining model layers.
[0080] Table 1 Performance of related methods and the method of this application
[0081]
[0082]
[0083] Please refer to Table 1, which compares this method with two related methods: the Early Exit strategy and the Skip Decode strategy. These two related methods are computational acceleration strategies suitable for large language models, so this method is also applied to large language models. Using BLOOMZ-7B as the base model, the performance of these related methods and this method on translation and text summarization tasks was evaluated. The Bilingual Evaluation Understudy (BLEU) and COMET metrics were used to evaluate the performance of the above methods on translation tasks, and RG-1 (Rouge-1), RG-2 (Rouge-2), and RG-L (Rouge-L) were used to evaluate the performance of the above methods on text summarization tasks.
[0084] It can be seen from Table 1 that as the target speedup ratio increases, the performance of this method in translation tasks and text summarization tasks becomes lower. However, compared with related methods, this method performs better in translation tasks and text summarization tasks.
[0085] Among them, the early exit strategy uses various carefully designed heuristic rules or an additional layer-by-layer classifier. For different tokens of an input sample, this method determines whether to exit the prediction layer early based on the prediction results of different tokens at the model layer. In other words, different tokens are often assigned different computational depths. Therefore, the early exit strategy is sensitive to the input sample. Although intuitive, it limits its value in practical application scenarios. Specifically, in existing commonly used batch decoding techniques, the early exit strategy cannot guarantee that different input samples within the same batch have the same exit point and activate the same sub-network structure, and therefore cannot support batch decoding. Similarly, with the KV Caching technology commonly used in autoregressive prediction models, if the execution depth of a subsequent token exceeds the execution depth of the token stored in the buffer, the current buffer contents need to be recalculated, which can easily lead to inefficient and wasteful computation.
[0086] The skip-layer decoding strategy predefines the exit points of all input samples and ensures that the number of model layers executed by the subsequent tokens must be less than the number of model layers executed by the previous tokens. This method can alleviate the problems of computational waste and inefficiency to a certain extent, but it still has several limitations. First, because in the actual processing process, due to the cumulative effect of errors, the prediction difficulty of the later positions is often greater than that of the previous positions, but the skip-layer decoding strategy allocates less prediction cost to these more difficult token predictions, which is inconsistent with the observations and conclusions of related technologies. Secondly, when actually processing input samples, it is impossible to know in advance the length of the final prediction result. Therefore, in most cases, the AI model will still execute more depth at a relatively early position according to the above strategy, resulting in limited acceleration and inability to be stable and controllable.
[0087] BLEU is a method for measuring the similarity between texts and is used as an indicator for evaluating the quality of machine translation. COMET is a machine translation quality evaluation indicator based on neural network semantic similarity calculation. Recall-Oriented Understudy for Gisting Evaluation (ROUGE) is an indicator used to evaluate the quality of machine text. RG-1 is a similarity evaluation indicator based on a single word, which only considers whether the text summary contains the same words as the keywords in the reference summary; RG-2 is a similarity evaluation indicator based on two words, which considers not only whether the text summary contains keywords, but also the order and combination of keywords; RG-L is a similarity evaluation indicator based on the longest common subsequence, which finds the longest common subsequence between the reference summary and the text summary and calculates its similarity score.
[0088] Through the above method, the total number of retained model layers is determined according to the acceleration ratio, so that the influence of input samples is reduced when determining the model layers retained by the AI model, thereby expanding the application scenarios of the method provided in this application.
[0089] In some embodiments, the numbers of the N model layers are determined based on their positional relationships in the AI model; and M model layers are selected from the N model layers as the retained M model layers based on the numbers and acceleration ratios of the N model layers.
[0090] In some embodiments, the numbers of the N model layers are determined based on the order in which the training samples are passed through each model layer after being input into the AI model.
[0091] In some embodiments, the total number of model layers N is divided by the acceleration ratio and then rounded to the nearest integer to obtain the number of model layers M retained.
[0092] For example, the number of model layers M retained can be obtained by the following formula:
[0093]
[0094] Where N represents the total number of model layers of the AI model, and r represents the pre-set acceleration ratio. Indicates rounding down to the nearest integer.
[0095] In some embodiments, the numbers of the M model layers are pre-set based on the number of retained model layers M; based on the numbering of the M model layers, M model layers are determined from the N model layers as the retained M model layers. It should be noted that the numbering of the M model layers is set by relevant technicians and is not limited in this application.
[0096] In some embodiments, based on a balanced discrete algorithm, the numbers of the M model layers are determined according to the number M of retained model layers, wherein the balanced discrete algorithm is used to generate a series of uniformly distributed numerical values; based on the numbers of the M model layers, M model layers are determined from the N model layers as the retained M model layers. For example, the balanced discrete algorithm can be an algorithm based on a Bayesian network, which generates a series of uniformly distributed numerical values based on the Bayesian network.
[0097] In some embodiments, according to the positional relationship of the N model layers in the AI model, the N model layers are numbered starting from 0 in order of their positions from front to back, and the numbers of the N model layers are determined.
[0098] For example, suppose the AI model includes three model layers, namely model layer 1, model layer 2, and model layer 3. Model layer 3 is in front of model layer 1 and model layer 2, and model layer 1 is in front of model layer 2. Therefore, model layer 1, model layer 2, and model layer 3 are numbered 2, 3, and 1 respectively.
[0099] In some embodiments, a retention set is initialized, and the retention set is used to record the numbers of the retained model layers; the numbers of the first model layer and the last model layer in the N model layers are added to the retention set; when the number of numbers included in the retention set is less than M, the number of the i-th model layer in the AI model is obtained, and the initial value of i is 2; if the number of the i-th model layer is divisible by the acceleration ratio, the number of the i-th model layer is added to the retention set; if the number of the i-th model layer is not divisible by the acceleration ratio, the value of i+1 is assigned to i, and the execution is started again from the step of obtaining the number of the i-th model layer in the AI model; when the number of numbers included in the retention set is equal to M, the model layers corresponding to the M numbers included in the retention set are determined as the M retained model layers.
[0100] For example, please refer to Figure 3 , which shows a flowchart for determining the retained model layers provided by an embodiment of the present application. The number M of the retained model layers is determined according to the acceleration ratio. When determining the M retained model layers from the N model layers of the AI model, the retained set that records the numbers of the retained model layers is first initialized, and the number of the first model layer and the number of the last model layer of the N model layers are added to the retained set. Then, it is determined in turn whether the numbers of the remaining N-2 model layers can be divided by the acceleration ratio. If they can be divided by the acceleration ratio and the number of model layers included in the retained set is less than M, the number of the model layer is added to the retained set. Finally, the model layers corresponding to the M numbers contained in the retained set are determined as the M retained model layers.
[0101] In the above manner, the retained model layers are determined from multiple model layers of the AI model by determining whether the model layer number is divisible by the acceleration ratio, thereby ensuring that the retained model layers are evenly and discretely distributed in the AI model.
[0102] In some embodiments, the retained model layer is determined from the N model layers according to the predefined model layer numbers. It should be noted that the above-mentioned predefined model layer numbers are identified or indicated by relevant technicians and are not limited in this application.
[0103] Step 230: Use the M model layers retained in the AI model to process each training sample separately to obtain a processing result.
[0104] In some embodiments, corresponding to any training sample, the M model layers retained in the AI model are activated; the sample data included in the training sample is processed using the activated M model layers to obtain the model output result of the training sample.
[0105] In some embodiments, the processing results include model output data corresponding to each training sample.
[0106] In some embodiments, a batch of samples is determined from a training sample set, where the batch of samples includes at least two training samples; the M model layers retained in the AI model are used to process at least two training samples contained in the batch of samples in parallel to obtain processing results.
[0107] In some embodiments, the M model layers retained in the AI model are activated; the training samples in the above-mentioned batch of samples are distributed to different computing nodes; the above-mentioned different computing nodes use the activated M model layers to process their respective training samples to obtain model output results corresponding to their respective training samples; and the processing results are obtained based on the model output results corresponding to different computing nodes.
[0108] Table 2 Throughput improvement effects of related methods and this method
[0109]
[0110]
[0111] As shown in Table 2, as the number of training samples in a batch increases, the method provided by this application significantly improves actual throughput at the same target speedup ratio. Since the early exit strategy is not applicable to batch processing of training samples, it is not compared with it. However, at the same target speedup ratio, this method also achieves better throughput improvement than the skip-layer decoding strategy.
[0112] Through the above method, the AI model processes multiple training samples at the same time, which is conducive to improving the processing efficiency of the AI model.
[0113] In some embodiments, when the retained M model layers include an attention mechanism layer, when the AI model performs the j-th round of inference on the training sample, the training sample is inferred based on the intermediate calculation results of the M model layers of the j-1-th round saved in the buffer to obtain the intermediate calculation results of the M model layers of the j-1-th round, where j is an integer greater than 1; wherein the buffer is used to store the intermediate calculation results of the M model layers.
[0114] The intermediate calculation results of the model layer refer to the key value (Key) and the corresponding value (Value) calculated by the attention mechanism layer on the input of the attention mechanism layer. For example, the input of the attention mechanism layer is multiplied by the weight vector corresponding to the Key and the weight vector corresponding to the Value respectively to obtain the intermediate calculation results of the attention mechanism layer.
[0115] In some embodiments, when the AI model is a large language model, the training samples include prompt information for the model processing task. When the AI model is used to process each training sample, M model layers are used to predict the prompt processing results of each training sample. The prompt information is used to guide the AI model to perform the model processing task, and the prompt processing result is obtained by processing the prompt information of the training sample using the model layer used to process the prompt information among the N model layers of the AI model. That is, in the stage of calculating the prompt information, the AI model uses the full depth of the model layer used to process the prompt information in the AI model for calculation, and in the stage of inferring the training sample, the retained M model layers are used for calculation.
[0116] For example, if the AI model is a large language model based on the Encoder-Decoder architecture, the Encoder part includes 6 Encoder layers and 6 Decoder layers. After the training sample is input to the AI model, the AI model uses the 6 Encoder layers in the Encoder part to process the input sequence to obtain the input feature sequence. Then, the retained Decoder layers determined in the Decoder part are used to infer and predict the input feature sequence to obtain the final model output result. The Encoder-Decoder architecture is a type of Transform architecture.
[0117] Please refer to Figure 4 , which shows a comparison diagram of the computational depth of the inference timing of different methods provided by an embodiment of the present application, wherein Figure 4 Subgraph (a) is the computational depth graph of the inference sequence of the early exit strategy. Figure 4 Sub-graph (b) is the computational depth graph of the inference sequence of the skip-layer decoding strategy. Figure 4 Subgraph (c) of is the calculation depth graph of the reasoning sequence of the method provided by this application. The above three methods are applied to the AI model based on Decoder-only. Figure 4 From the sub-graphs (a), (b) and (c), we can see that the time series x1~x m In the input data processing stage, after the sample is input into the model, all three methods use the complete AI model to process the prompt information of the sample. n For the prediction and reasoning stage, Figure 4As can be seen from the subgraph (a) of , the early exit strategy uses different computational depths for prediction and reasoning of different tokens, and is more sensitive to samples. Figure 4 As can be seen from the subgraph (b) of , the skip-layer strategy performs more computational depth when inferring and predicting relatively early tokens; Figure 4 As can be seen from sub-figure (c), the method provided in this application uses the same computing depth in the inference and prediction stage, and the model layers used are evenly and uniformly distributed in the AI model.
[0118] Through the above method, by saving the intermediate calculation results of the model layer retained by the AI model for each round of reasoning on the training samples, the calculation overhead is reduced, thereby speeding up the calculation speed of the AI model on the training samples.
[0119] Given a given total model layer count and acceleration ratio, all samples retain the same model layers and activate the same sub-layer network structure, enabling the model to easily support modern decoding acceleration strategies, such as batch decoding and KV caching. Furthermore, the sample-independent layer skipping strategy ensures that the model is unaffected by sample dynamics during actual inference and decoding, resulting in a stable, precise, and controllable acceleration effect.
[0120] Step 240: Adjust the parameters of the AI model based on the processing results to obtain a trained AI model.
[0121] In some embodiments, the parameters of all N model layers of the AI model are adjusted based on the processing results to obtain a trained AI model.
[0122] In some embodiments, the parameters of the retained M model layers are adjusted based on the processing results to obtain a trained AI model.
[0123] In some embodiments, the loss function value of the AI model is calculated based on the model output data corresponding to each training sample and the label data included in each training sample; the parameters of the AI model are adjusted based on the loss function value to obtain a trained AI model.
[0124] In some embodiments, the loss function value of the AI model can be calculated using loss function formulas such as the cross entropy loss function, the mean absolute error loss function (MAL), and the mean square error loss function (MSE).
[0125] In some embodiments, the parameters of the AI model can be adjusted using gradient descent.
[0126] It should be noted that there are many ways to adjust the parameters of the AI model based on the processing results. The adjustment method can be selected based on the type of AI model and the needs of the actual application. This application does not limit this.
[0127] Through the above method, it is ensured that the AI model still has good model reasoning capabilities when skipping the execution of some model layers.
[0128] To sum up, the technical solution provided by the embodiment of the present application determines the retained M model layers from the N model layers of the AI model, and then uses the retained M model layers to process multiple training samples to obtain corresponding processing results, where M is less than N. Finally, the parameters of the above-mentioned AI model are adjusted based on the processing results to obtain the trained AI model; since some model layers in the AI model are skipped, the processing speed of the training samples is accelerated, and it is ensured that the retained model layers in the trained AI model are used to process other samples while still having good reasoning capabilities.
[0129] The following describes the prediction process using the AI model. The AI model usage process corresponds to the method steps of the training process. For details not described in detail in the embodiments corresponding to the training process, please refer to the description of the embodiments corresponding to the usage process. Similarly, for details not described in detail in the embodiments corresponding to the usage process, please refer to the description of the embodiments corresponding to the training process.
[0130] Please refer to Figure 5 , which shows a flow chart of the prediction method of the AI model provided by one embodiment of the present application. The execution subject of each step of the method can be a computer device, for example, the computer device can be Figure 1 The model in the solution implementation environment shown uses a device 20. The method may include at least one of the following steps (510-530):
[0131] Step 510: Obtain a prediction sample set of the AI model. The AI model includes N model layers. The prediction sample set includes multiple prediction samples, where N is an integer greater than or equal to 3.
[0132] In some embodiments, the above-mentioned AI model is obtained by training using the AI model training method provided in the previous embodiment.
[0133] The prediction sample set is a sample set that requires the use of an AI model for predictive reasoning.
[0134] Step 520: Determine the retained M model layers from the N model layers, wherein the remaining model layers except the M model layers in the N model layers are skipped when the AI model is used to process each prediction sample, and M is an integer greater than or equal to 2 and less than N.
[0135] In some embodiments, the reserved M model layers are determined from the N model layers according to the reserved set. The reserved set is used to indicate the reserved M model layers.
[0136] Optionally, the reserved set can be obtained during the training process of the AI model, or can be pre-labeled by relevant technical personnel.
[0137] In some embodiments, during the prediction process, the model utilization device 20 may obtain model information related to the AI model, where the model information includes information indicating the M retained model layers. For example, the model information may include the serial numbers of the M retained model layers. The model utilization device 20 determines the M retained model layers from the N model layers based on the model information.
[0138] In some embodiments, the M model layers are evenly distributed in the AI model.
[0139] In some embodiments, the M model layers are determined during the training phase of the AI model based on the total number N of model layers of the AI model and a pre-set acceleration ratio, where the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model.
[0140] Step 530: Use the M model layers retained in the AI model to process each prediction sample separately to obtain a processing result.
[0141] In some embodiments, a batch of samples is determined from the predicted samples, the batch of samples including at least two of the training samples; the at least two training samples included in the batch of samples are processed in parallel using the M model layers retained in the aforementioned AI model to obtain the processing results. For a detailed description of this, please refer to the corresponding section of the model training method embodiment.
[0142] In some embodiments, when the retained M model layers include an attention mechanism layer, when the AI model performs the jth round of inference on a prediction sample, the prediction sample is inferred based on the intermediate calculation results of the M model layers in the j-1th round to obtain the intermediate calculation results of the M model layers in the jth round, where the initial value of j is 2; the intermediate calculation results of the M model layers in the jth round are added to a buffer, which is used to store the intermediate calculation results of the M model layers. For a detailed description of this, please refer to the corresponding section in the model training method embodiment.
[0143] To sum up, according to the technical solution provided in the embodiment of the present application, by using part of the model layer of the AI model to process the prediction sample set, the processing results corresponding to the prediction sample set are obtained, which is conducive to improving the calculation speed of the AI model.
[0144] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0145] Please refer to Figure 6 , which shows a block diagram of an AI model training device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned AI model training method, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the model training device 6 introduced above, or it can be set in the model training device 6. Figure 6 As shown, the apparatus 600 may include a training sample acquisition module 66 , a training model layer determination module 620 , a training sample processing module 630 and a model adjustment module 640 .
[0146] The training sample acquisition module 610 is used to obtain a training sample set of the AI model, where the AI model includes N model layers, and the training sample set includes multiple training samples, where N is an integer greater than or equal to 3.
[0147] The training model layer determination module 620 is used to determine the retained M model layers from the N model layers, wherein the remaining model layers except the M model layers in the N model layers are skipped when the AI model is used to process each of the training samples, and M is an integer greater than or equal to 2 and less than N.
[0148] The training sample processing module 630 is used to use the M model layers retained in the AI model to process each of the training samples separately to obtain a processing result.
[0149] The model adjustment module 640 is used to adjust the parameters of the AI model based on the processing results to obtain a trained AI model.
[0150] In some embodiments, the M model layers are evenly distributed in the AI model.
[0151] In some embodiments, the training model layer determination module 620 includes a layer number determination submodule and a model layer determination submodule (in Figure 6 not shown).
[0152] The layer number determination submodule is used to determine the number of retained model layers M based on the total number of model layers N of the AI model and a pre-set acceleration ratio, where the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model.
[0153] The model layer determination submodule is used to select M model layers from the N model layers according to the number M of retained model layers as the retained M model layers.
[0154] In some embodiments, the layer number determination submodule is used to divide the total number of model layers N by the acceleration ratio and then round it up to obtain the retained number of model layers M.
[0155] In some embodiments, the model layer determination submodule includes a layer numbering unit and a layer determination unit (in Figure 6 not shown).
[0156] The layer numbering unit is used to determine the numbers of the N model layers according to the positional relationship of the N model layers in the AI model.
[0157] A layer determining unit is configured to select M model layers from the N model layers as the retained M model layers according to respective numbers of the N model layers and the acceleration ratio.
[0158] In some embodiments, the layer numbering unit is used to assign numbers starting from 0 to the N model layers according to the positional relationship of the N model layers in the AI model, in order from front to back, to determine the numbers of each of the N model layers.
[0159] In some embodiments, the layer determination unit is used to initialize a retention set, which is used to record the numbers of the retained model layers; the numbers of the first model layer and the last model layer in the N model layers are added to the retention set; when the number of numbers contained in the retention set is less than M, the number of the i-th model layer in the AI model is obtained, and the initial value of i is 2; if the number of the i-th model layer is divisible by the acceleration ratio, the number of the i-th model layer is added to the retention set; if the number of the i-th model layer is not divisible by the acceleration ratio, the value of i+1 is assigned to i, and the step of obtaining the number of the i-th model layer in the AI model is started again; when the number of numbers contained in the retention set is equal to M, the model layers corresponding to the M numbers contained in the retention set are determined as the retained M model layers.
[0160] In some embodiments, the training sample processing module 630 is used to determine a batch of samples from the training sample set, where the batch of samples includes at least two of the training samples; and the M model layers retained in the AI model are used to process at least two of the training samples contained in the batch of samples in parallel to obtain the processing results.
[0161] In some embodiments, the training sample processing module 630 is also used to, when the retained M model layers include an attention mechanism layer, perform inference on the training sample based on the intermediate calculation results of the M model layers of the j-1th round stored in the buffer, to obtain the intermediate calculation results of the M model layers of the jth round, where j is an integer greater than 1; wherein the buffer is used to store the intermediate calculation results of the M model layers.
[0162] To sum up, the technical solution provided by the embodiment of the present application determines the retained M model layers from the N model layers of the AI model, and then uses the retained M model layers to process multiple training samples to obtain corresponding processing results, where M is less than N. Finally, the parameters of the above-mentioned AI model are adjusted based on the processing results to obtain the trained AI model; since some model layers in the AI model are skipped, the processing speed of the training samples is accelerated, and it is ensured that the retained model layers in the trained AI model are used to process other samples while still having good reasoning capabilities.
[0163] Please refer to Figure 7 , which shows a block diagram of a prediction device for an AI model provided by an embodiment of the present application. The device has the function of implementing the prediction method of the above-mentioned AI model, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the model using device 20 introduced above, or it can be set in the model using device 20. Figure 7 As shown, the apparatus 700 may include a prediction sample acquisition module 710 , a prediction model layer determination module 720 , and a prediction sample processing module 730 .
[0164] The prediction sample acquisition module 710 is used to obtain a prediction sample set of the AI model, where the AI model includes N model layers, and the prediction sample set includes multiple prediction samples, where N is an integer greater than or equal to 3.
[0165] The prediction model layer determination module 720 is used to determine the retained M model layers from the N model layers, wherein the remaining model layers except the M model layers in the N model layers are skipped when the AI model is used to process each of the prediction samples, and M is an integer greater than or equal to 2 and less than N.
[0166] The prediction sample processing module 730 is used to use the M model layers retained in the AI model to process each of the prediction samples separately to obtain a processing result.
[0167] In some embodiments, the M model layers are evenly distributed in the AI model.
[0168] In some embodiments, the M model layers are determined during the training phase of the AI model based on the total number N of model layers of the AI model and a pre-set acceleration ratio, wherein the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model.
[0169] To sum up, according to the technical solution provided in the embodiment of the present application, by using part of the model layer of the AI model to process the prediction sample set, the processing results corresponding to the prediction sample set are obtained, which is conducive to improving the calculation speed of the AI model.
[0170] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0171] Please refer to Figure 8 , which shows a block diagram of a computer device 800 provided in one embodiment of the present application. The computer device 800 may be Figure 1 The model training device 10 in the implementation environment shown can also be Figure 1 The model using device 20 in the illustrated implementation environment is used to implement the AI model training method or AI model prediction method provided in the above embodiments. Specifically:
[0172] Typically, the computer device 800 includes a processor 810 and a memory 820 .
[0173] The processor 810 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 810 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 810 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 810 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 810 may also include an AI processor for processing computing operations related to machine learning.
[0174] The memory 820 may include one or more computer-readable storage media, which may be non-transitory. The memory 820 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 820 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned AI model training method or AI model prediction method.
[0175] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the computer device 800, and the computer device 800 may include more or fewer components than shown in the figure, or combine some components, or adopt a different arrangement of components.
[0176] In an exemplary embodiment, a computer-readable storage medium is further provided, wherein a computer program is stored in the storage medium, and when the computer program is executed by the processor, the training method of the above-mentioned AI model or the prediction method of the AI model is implemented. Optionally, the computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD) or an optical disc, etc. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM).
[0177] In an exemplary embodiment, a computer program product is further provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described AI model training method or AI model prediction method.
[0178] It should be noted that the collection and processing of relevant data (such as training samples and prediction samples) in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in real cases, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0179] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.
[0180] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A training method for an artificial intelligence (AI) model, characterized in that: The method comprises: Obtain a training sample set of the AI model, where the AI model includes N model layers, the training sample set includes multiple training samples, and N is an integer greater than or equal to 3; Determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the training samples, and M is an integer greater than or equal to 2 and less than N; Using the M model layers retained in the AI model, each of the training samples is processed separately to obtain a processing result; The parameters of the AI model are adjusted based on the processing results to obtain a trained AI model.
2. The method according to claim 1, characterized in that The M model layers are evenly distributed in the AI model.
3. The method according to claim 1, characterized in that Determining the retained M model layers from the N model layers includes: Determining the number of model layers M to be retained based on the total number N of model layers of the AI model and a preset acceleration ratio, wherein the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model; According to the number M of retained model layers, M model layers are selected from the N model layers as the retained M model layers.
4. The method according to claim 3, characterized in that The step of determining the number of model layers M to be retained based on the total number N of model layers of the AI model and a preset acceleration ratio includes: The total number of model layers N is divided by the acceleration ratio and then rounded to the nearest integer to obtain the number of retained model layers M.
5. The method according to claim 3, characterized in that The selecting M model layers from the N model layers as the retained M model layers according to the retained number M of model layers includes: Determining numbers of the N model layers according to positional relationships of the N model layers in the AI model; According to the respective numbers of the N model layers and the acceleration ratio, M model layers are selected from the N model layers as the retained M model layers.
6. The method according to claim 5, characterized in that Determining the numbers of the N model layers according to the positional relationship of the N model layers in the AI model includes: According to the positional relationship of the N model layers in the AI model, the N model layers are numbered starting from 0 in order of their positions, and the numbers of the respective N model layers are determined.
7. The method according to claim 5, characterized in that The selecting M model layers from the N model layers as the retained M model layers according to the respective numbers of the N model layers and the acceleration ratios includes: Initialize a reserved set, where the reserved set is used to record the numbers of the reserved model layers; Adding the number of the first model layer and the number of the last model layer in the N model layers to the retained set; When the number of numbers included in the reserved set is less than M, obtaining the number of the i-th model layer in the AI model, where the initial value of i is 2; If the number of the i-th model layer is divisible by the acceleration ratio, then the number of the i-th model layer is added to the retained set; If the number of the i-th model layer is not divisible by the acceleration ratio, assign a value of i+1 to i, and start again from the step of obtaining the number of the i-th model layer in the AI model; When the number of numbers included in the reserved set is equal to M, the model layers corresponding to the M numbers included in the reserved set are determined as the reserved M model layers.
8. The method according to any one of claims 1 to 7, characterized in that The M model layers retained in the AI model are used to process each of the training samples separately to obtain processing results, including: Determine a batch of samples from the training sample set, wherein the batch of samples includes at least two of the training samples; The M model layers retained in the AI model are used to process at least two of the training samples included in the batch of samples in parallel to obtain the processing result.
9. The method according to any one of claims 1 to 7, characterized in that In the case that the retained M model layers include an attention mechanism layer, when the AI model performs the j-th round of inference on the training sample, the training sample is inferred based on the intermediate calculation results of the M model layers of the j-1-th round saved in the buffer to obtain the intermediate calculation results of the M model layers of the j-th round, where j is an integer greater than 1; wherein the buffer is used to store the intermediate calculation results of the M model layers.
10. A prediction method for an artificial intelligence (AI) model, characterized in that: The method comprises: Obtain a prediction sample set of the AI model, where the AI model includes N model layers, the prediction sample set includes a plurality of prediction samples, and N is an integer greater than or equal to 3; Determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the prediction samples, and M is an integer greater than or equal to 2 and less than N; The M model layers retained in the AI model are used to process each of the prediction samples separately to obtain a processing result.
11. The method according to claim 10, characterized in that The M model layers are evenly distributed in the AI model.
12. The method according to claim 10, characterized in that The M model layers are determined during the training phase of the AI model based on the total number N of model layers of the AI model and a pre-set acceleration ratio, wherein the acceleration ratio is used to indicate the expected degree of improvement in the computing speed of the AI model.
13. A training device for an artificial intelligence (AI) model, characterized in that: The device comprises: A training sample acquisition module, configured to acquire a training sample set for the AI model, wherein the AI model includes N model layers, and the training sample set includes a plurality of training samples, where N is an integer greater than or equal to 3; A training model layer determination module, configured to determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the training samples, and M is an integer greater than or equal to 2 and less than N; A training sample processing module, configured to process each of the training samples separately using the M model layers retained in the AI model to obtain a processing result; A model adjustment module is used to adjust the parameters of the AI model based on the processing results to obtain a trained AI model.
14. A prediction device for an artificial intelligence (AI) model, characterized in that: The device comprises: A prediction sample acquisition module, configured to acquire a prediction sample set of the AI model, wherein the AI model includes N model layers, and the prediction sample set includes a plurality of prediction samples, where N is an integer greater than or equal to 3; A prediction model layer determination module, configured to determine M retained model layers from the N model layers, wherein the remaining model layers in the N model layers except the M model layers are skipped when the AI model is used to process each of the prediction samples, and M is an integer greater than or equal to 2 and less than N; The prediction sample processing module is used to use the M model layers retained in the AI model to process each of the prediction samples separately to obtain a processing result.
15. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 12.
16. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which is used to be executed by a processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 12.
17. A computer program product, characterized in that The computer program product comprises a computer program, and the computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 12.