Training method of vertical field visual model of general large model
By targeted training of the parameters of the general big model, using backpropagation algorithm and gradient optimization methods, the problem of low accuracy in the application of general models in vertical fields is solved, and high accuracy recognition in fields such as medical imaging is achieved.
Patent Information
- Application Number
- CN202510741763.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The difficulty of universal large models to directly capture and utilize the specificity of vertical field data, resulting in a significant reduction in application accuracy in fields such as medical imaging.
By obtaining the general large model and vertical domain image training set, using backpropagation algorithm and gradient optimization method, the model parameters are updated until the predicted classification results are consistent with the real results, and a small vertical domain visual model is trained.
It improves the recognition accuracy in vertical fields and reduces the risk of misdiagnosis and missed diagnosis, especially in the field of medical imaging, which can accurately identify subtle features of the lesions.
Smart Images

Figure CN120259792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a training method for a vertical domain visual model of a general large model. Background Art
[0002] With the rapid development of computer technology, computer vision technology has made remarkable progress. Models based on deep neural networks have shown excellent performance in tasks such as image recognition, object detection, and image segmentation. General large visual models, such as those trained on large-scale image datasets (e.g., ImageNet), have been able to make relatively accurate judgments in a wide range of visual scenarios; Although general models have strong generalization capabilities, data in vertical domains often have unique distributions and characteristics. For example, in the field of medical imaging, the imaging principles, image contents, and visual patterns of imaging data such as X-rays, CTs, and MRIs are very different from natural images. General models are difficult to directly capture and utilize the specificities of these vertical domain data, resulting in a significant reduction in accuracy in vertical domain applications. Summary of the Invention
[0003] The purpose of the present invention is to provide a training method for a vertical domain visual model of a general large model to solve the technical problem that the general model has weak data processing capabilities in vertical domains. For example, in the field of medical imaging, the imaging principles, image contents, and visual patterns of imaging data such as X-rays, CTs, and MRIs are very different from natural images. General models are difficult to directly capture and utilize the specificities of these vertical domain data, resulting in a significant reduction in accuracy in vertical domain applications.
[0004] To achieve the above purpose, the present invention provides the following technical solutions: A training method for a vertical domain visual model of a general large model, comprising: Obtain a general large model and a vertical domain image training set, extract vertical domain image feature data from the vertical domain image training set according to the general large model, and output prediction classification result information according to the vertical domain image feature data; Obtain the true classification result information of the vertical domain image training set; Obtain a vertical domain prediction error rate evaluation value according to the true classification result and the prediction classification result information; Obtain the gradient of each parameter in the general large model through the backpropagation algorithm based on the vertical domain prediction error rate evaluation value; Update the key parameters of the general large model based on the gradient until the prediction classification result information is equal to the true classification result information, and obtain a trained vertical domain visual small model.
[0005] Preferably, the step of extracting vertical domain image feature data from the vertical domain image training set according to the general large model and outputting prediction classification result information includes: Obtain the vertical domain image data of the vertical domain image training set; Preprocess the vertical domain image data; Construct the convolutional neural network of the general large model; Extract vertical domain image feature data through the convolutional layer of the convolutional neural network; Fuse the vertical domain image feature data through the fully connected layer of the convolutional neural network to obtain a comprehensive feature vector; Perform classification calculation on the comprehensive feature vector according to the Softmax classifier to obtain the prediction classification result information.
[0006] Preferably, the step of obtaining the vertical domain prediction error rate evaluation value according to the true classification result and the prediction classification result information includes: Obtain the total number of categories of the prediction classification result information; Obtain the true result data of each category in the true classification result information; Obtain the prediction result data of each category in the prediction classification result information; Obtain the weight of the true result data of each category; Obtain the total number of samples in the vertical domain image training set; Obtain the vertical domain prediction error rate evaluation value according to the total number of categories, the true result data, the prediction result data and the weight of the true result data. The calculation formula is: ; In the formula, the represents the vertical domain prediction error rate evaluation value, represents the total number of samples, represents the total number of categories, represents the th weight of the category, represents the th true result data of the sample, represents the th prediction result data of the sample.
[0007] Preferably, obtain the gradient of each parameter in the general large model through the backpropagation algorithm based on the vertical domain prediction error rate evaluation value: Obtain the activation value of each neuron in the general large model; Obtain the input values of each neuron in the general large model; Obtain the partial derivative of the loss function in the general large model; Obtain the partial derivative of the weights in the general large model; Obtain the gradient of each parameter in the general large model according to the activation value, the input value, the partial derivative of the loss function, and the partial derivative of the weights. The calculation formula is: ; In the formula, represents the gradient of each parameter in the general large model, represents the partial derivative of the loss function, represents the th activation value of the th neuron in the th layer, represents the input value of the th neuron in the
[0008] Preferably, the step of updating the key parameters of the general large model based on the gradient includes: Obtain the gradient; Based on the gradient update, obtain the first-order moment estimation mean and the second-order moment estimation non-central variance; Correct the first-order moment estimation mean and the second-order moment estimation non-central variance; Obtain the key parameters of the general large model according to the corrected first-order moment estimation mean and the second-order moment estimation non-central variance.
[0009] Preferably, the step of obtaining the first-order moment estimation mean based on the gradient update includes: Obtain the gradient; Obtain the first hyperparameter; Obtain the first-order moment estimation mean according to the first hyperparameter and the gradient. The calculation formula is: ; In the formula, represents the updated first-order moment estimation mean, represents the first-order moment estimation mean before update, represents the gradient of each parameter in the general large model.
[0010] Preferably, the step of obtaining the second-order moment estimation non-central variance based on the gradient update includes: Obtain the gradient; Obtain the second hyperparameter; Obtain the first - moment estimate mean according to the second hyperparameter and the gradient, and its calculation formula is: ; In the formula, represents the updated second - moment estimate non - central variance, represents the second - moment estimate non - central variance before update, represents the second hyperparameter, represents the gradient of each parameter in the general large - model.
[0011] Preferably, the step of correcting the first - moment estimate mean includes: Obtain the initial first - moment estimate mean; Obtain the first hyperparameter; Obtain the first - moment estimate mean according to the initial first - moment estimate mean and the first hyperparameter, and the calculation formula is: ; In the formula, represents the first - moment estimate mean, the first hyperparameter of the t - th iteration, represents the initial first - moment estimate mean.
[0012] Preferably, the step of correcting the second - moment estimate non - central variance includes: Obtain the initial second - moment estimate non - central variance; Obtain the second hyperparameter; Obtain the second - moment estimate non - central variance according to the initial second - moment estimate non - central variance and the second hyperparameter, and the calculation formula is: ; In the formula, represents the second - moment estimate non - central variance, the second hyperparameter of the t - th iteration, represents the initial second - moment estimate non - central variance.
[0013] Preferably, the step of obtaining the key parameters of the general large - model according to the corrected first - moment estimate mean and the second - moment estimate non - central variance includes: Obtain the corrected first - moment estimate mean; Obtain the corrected second - moment estimate non - central variance; Obtain the key parameter update step size; Obtain the key parameters of the general large - model according to the first - moment estimate mean, the second - moment estimate non - central variance and the key parameter update step size, and its calculation formula is: ; In the formula, represents the key parameter of the general large model, represents the non-central variance of the second moment estimation, represents the mean of the first moment estimation, represents the key parameter update step size, represents the positive number stability value.
[0014] The beneficial effects of this application are as follows: The training method of this application is specifically designed for the data characteristics and visual tasks in the vertical field. Through targeted training on the vertical field dataset, the model can accurately capture the unique image features and patterns in this field. For example, in the field of medical imaging, it can accurately identify the subtle features of lesions, and there is a significant improvement in the accuracy rate compared to the application of the general model in the vertical field, reducing the risk of misdiagnosis and missed diagnosis. Brief Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of the method for an embodiment of this application.
[0016] The realization, functional characteristics, and advantages of the purpose of this application will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments
[0017] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] As Figure 1 shown, this application provides a training method for a vertical field visual model of a general large model, including: Obtain a general large model and a vertical field image training set, extract vertical field image feature data from the vertical field image training set according to the general large model, and output prediction classification result information according to the vertical field image feature data; Obtain the true classification result information of the vertical field image training set; Obtain a vertical field prediction error rate evaluation value according to the true classification result and the prediction classification result information; Based on the vertical field prediction error rate evaluation value, obtain the gradient of each parameter in the general large model through the backpropagation algorithm; Update the key parameters of the general large model based on the gradient until the prediction classification result information is equal to the true classification result information, and obtain a trained vertical field visual small model.
[0019] The step of extracting vertical field image feature data from the vertical field image training set according to the general large model and outputting prediction classification result information according to the vertical field image feature data includes: Obtain the vertical domain image data of the vertical domain image training set; Preprocess the vertical domain image data; Construct the convolutional neural network of the general large model; Extract the vertical domain image feature data through the convolutional layer of the convolutional neural network; Fuse the vertical domain image feature data through the fully connected layer of the convolutional neural network to obtain a comprehensive feature vector; Perform classification calculation on the comprehensive feature vector according to the Softmax classifier to obtain the predicted classification result information.
[0020] The vertical domain image training set includes industrial images, medical images, etc. Taking lung images in medical images as an example, by training with a general large model, it can be trained using a large-scale general image dataset and has learned a lot of general image features, such as the edges and textures of lung diseases. By loading these pre-trained parameters, the training convergence speed of the small model in the vertical domain can be accelerated. According to the Softmax classifier, classification calculation is performed on the comprehensive feature vector to obtain the predicted classification result information. For example, we have a dataset for diagnosing lung diseases, which contains lung CT images. These images are divided into three categories: normal lungs, lungs of pneumonia patients, and lungs of lung cancer patients. First, these images are preprocessed by normalizing them to the same size (such as 256×256 pixels) and normalizing the pixel values in the images so that their range is between 0 and 1. Then, a convolutional neural network (CNN) is constructed. The network structure includes multiple convolutional layers, pooling layers, and fully connected layers. The first layer is a convolutional layer with a convolutional kernel size of 3×3 and a quantity of 16. This layer can extract low-level features in lung images, such as edges and textures. Then, there is a 2×2 max pooling layer to reduce the data dimension while retaining important features. The second convolutional layer has a convolutional kernel size of 3×3 and a quantity of 32, further extracting more complex features. Then, there is another 2×2 max pooling layer. Finally, there is a fully connected layer to fuse and classify the features extracted previously. When a lung image is input into the network, the first convolutional layer starts to work. For example, for a lung CT image of a pneumonia patient, the convolutional layer will detect some local features in the lung image. It may detect the edge features of the lung consolidation area. These edge features are low-level features related to shape features because pneumonia causes local consolidation in the lungs, and there is an obvious boundary between the edge of the consolidated area and normal lung tissue. At the same time, the convolutional layer will also detect the texture features of uneven density inside the consolidated area, which belong to low-level features related to density features. As the network deepens, the second convolutional layer will further extract more advanced features on the basis of the features extracted by the first layer. For example, it may combine some local consolidation edge features and density texture features to form a pattern that can better represent the characteristics of pneumonia, obtaining high-level features related to shape features and density features. Feature fusion is performed in the last fully connected layer of the network;After passing through the previous convolutional layer and pooling layer, shape feature vectors and density feature vectors are obtained. The shape feature vectors may contain information about the shape complexity of the lung consolidation regions (such as features obtained by calculating the perimeter, area, etc. of the consolidated regions), and the density feature vectors may contain information such as the variance of the internal density of the consolidated regions and the gray histogram features. In the fully connected layer, these two feature vectors are concatenated. For example, if the length of the shape feature vector is 10 and the length of the density feature vector is 10, a combined feature vector with a length of 20 is obtained after concatenation. This combined feature vector is input into a Softmax classifier to determine whether the lung image belongs to normal lungs, pneumonia, or lung cancer.
[0021] The step of obtaining the vertical domain prediction error rate evaluation value according to the true classification result and the prediction classification result information includes: Obtain the total number of categories of the prediction classification result information; Obtain the true result data of each category in the true classification result information; Obtain the prediction result data of each category in the prediction classification result information; Obtain the weight of the true result data of each category; Obtain the total number of samples in the vertical domain image training set; Obtain the vertical domain prediction error rate evaluation value according to the total number of categories, the true result data, the prediction result data, and the weight of the true result data. The calculation formula is: ; In the formula, the represents the vertical domain prediction error rate evaluation value, represents the total number of samples, represents the total number of categories, represents the th category weight, represents the th sample's true result data, represents the th sample's prediction result data.
[0022] Taking lung image detection as an example, for instance, the number of samples of lesions (such as tumors, nodules, etc.) is often much less than that of normal samples. This will cause the model to tend to predict as normal samples. By assigning different weights to different classes, the impact of this data imbalance can be alleviated, and a more accurate evaluation value of the vertical domain prediction error rate can be obtained. When the probability of the model predicting the correct class is higher, the evaluation value of the vertical domain prediction error rate is smaller; conversely, when the probability of predicting the correct class is lower, the evaluation value of the vertical domain prediction error rate is larger. This encourages the model to continuously adjust parameters during training, making the predicted probability distribution closer to the probability distribution of the true labels, thereby improving the accuracy of classification.
[0023] Based on the evaluation value of the vertical domain prediction error rate, the gradient of each parameter in the general large model is obtained through the backpropagation algorithm: Obtain the activation value of each neuron in the general large model; Obtain the input value of each neuron in the general large model; Obtain the partial derivative of the loss function in the general large model; Obtain the partial derivative of the weight in the general large model; According to the activation value, the input value, the partial derivative of the loss function, and the partial derivative of the weight, the gradient of each parameter in the general large model is obtained. The calculation formula is: ; In the formula, represents the gradient of each parameter in the general large model, represents the partial derivative of the loss function, represents the th activation value of the th neuron in the th layer, represents the partial derivative of the weight.
[0024] The backpropagation algorithm is a method for calculating the gradients of the loss function with respect to the parameters (weights and biases) of each layer in a neural network. It is based on the chain rule and starts from the output layer of the neural network, calculating the gradients layer by layer in reverse, providing the information required for the optimization algorithm to update the parameters to minimize the loss function. Taking the lung disease detection model as an example, in the lung disease diagnosis model, assuming the cross-entropy loss function is used, for a specific training sample (lung image and its corresponding true result data), the gradients calculated through the backpropagation algorithm will tell the model in which direction each parameter should be adjusted to reduce the value of the loss function. For example, if the element corresponding to a certain convolutional kernel weight is positive, it means that increasing this weight will increase the loss function, so this weight should be reduced during parameter update.
[0025] The steps of updating the key parameters of the general large model based on the gradient include: Obtain the gradient; Based on the gradient, update to obtain the first moment estimation mean and the second moment estimation non - central variance; Correct the first moment estimation mean and the second moment estimation non - central variance; Obtain the key parameters of the general large model according to the corrected first moment estimation mean and the second moment estimation non - central variance.
[0026] This step can quickly and stably obtain the key parameters of the general large model. It adjusts the learning rate according to the historical gradient information of each parameter. For parameters that are frequently updated, the learning rate will gradually decrease. This characteristic enables the key parameters to be found faster, but in the long - term training process, it may lead to the learning rate being too small to continue optimizing the loss function. In contrast, this method combines the ideas of the momentum method and adaptive learning rate, performs well in terms of convergence speed and stability, can make the loss function reach a better value faster, and find the key parameters faster.
[0027] The steps of updating to obtain the first moment estimation mean based on the gradient include: Obtain the gradient; Obtain the first hyperparameter; Obtain the first moment estimation mean according to the first hyperparameter and the gradient, and its calculation formula is: ; In the formula, represents the updated first moment estimation mean, represents the first moment estimation mean before update, represents the first hyperparameter, represents the gradient of each parameter in the general large model.
[0028] Taking the lung disease detection model as an example, the first hyperparameter represents the number of layers of the neural network. When training the lung disease diagnosis model, it is first set to a common hyperparameter value, and the performance of the model on the validation set (such as accuracy, loss function value, etc.) is observed. Then, the hyperparameters are gradually adjusted. If the model is too simple (with fewer layers), it may not be able to fully learn the complex features in the lung images (such as shape features and density features); if the model is too complex (with too many layers), it may lead to overfitting, that is, it performs well on the training data but poorly on the test data. For example, a simple neural network may have 3 layers (including the input layer, hidden layer, and output layer), while a complex model may have more than 10 layers. When training the lung disease diagnosis model, start with some common hyperparameter values, observe the performance of the model on the validation set (such as accuracy, loss function value, etc.), and then gradually adjust the hyperparameters.
[0029] The step of obtaining the second-moment estimate non-central variance based on the gradient update includes: Obtain the gradient; Obtain the second hyperparameter; Obtain the first-moment estimate mean according to the second hyperparameter and the gradient, and its calculation formula is: ; In the formula, represents the updated second-moment estimate non-central variance, represents the second-moment estimate non-central variance before update, represents the second hyperparameter, represents the gradient of each parameter in the general large model.
[0030] The second hyperparameter is similar to the first hyperparameter and is usually close to 1, such as 0.999. At each iteration, the new second-moment estimate non-central variance is the old second-moment estimate non-central variance multiplied by the second hyperparameter plus the square of the current gradient multiplied by , which helps to estimate the change amplitude of the gradient.
[0031] The step of correcting the first-moment estimate mean includes: Obtain the initial first-moment estimate mean; Obtain the first hyperparameter; Obtain the first-moment estimate mean according to the initial first-moment estimate mean and the first hyperparameter, and the calculation formula is: ; In the formula, represents the first-moment estimate mean, the first hyperparameter at the t-th iteration, represents the initial first-moment estimate mean.
[0032] In the initial stage of training, since the mean of the first - moment estimate is initialized to, it will lead to a deviation in the initial estimate value. This formula is a deviation correction for the mean of the first - moment estimate. As the number of iterations increases, the influence of the deviation correction gradually decreases. Taking the lung disease recognition model as an example, in the initial stage of training the lung disease recognition model, this deviation correction can make the parameter update more reasonable.
[0033] The steps for correcting the non - central variance of the second - moment estimate include: Obtain the initial non - central variance of the second - moment estimate; Obtain the second hyperparameter; Obtain the non - central variance of the second - moment estimate according to the initial non - central variance of the second - moment estimate and the second hyperparameter. The calculation formula is: ; In the formula, represents the non - central variance of the second - moment estimate, is the second hyperparameter at the t - th iteration, represents the initial non - central variance of the second - moment estimate.
[0034] This formula is a deviation correction for the second - moment estimate, which ensures that in the early iterations, the learning rate adjustment based on the non - central variance of the second - moment estimate does not become unreasonable due to the initial value being 0.
[0035] The steps for obtaining the key parameters of the general large - model according to the corrected mean of the first - moment estimate and the non - central variance of the second - moment estimate include: Obtain the corrected mean of the first - moment estimate; Obtain the corrected non - central variance of the second - moment estimate; Obtain the key parameter update step size; Obtain the key parameters of the general large - model according to the mean of the first - moment estimate, the non - central variance of the second - moment estimate, and the key parameter update step size. The calculation formula is: ; In the formula, represents the key parameters of the general large - model, represents the non - central variance of the second - moment estimate, represents the mean of the first - moment estimate, represents the key parameter update step size, represents a positive stability value.
[0036] According to the corrected mean of the first - moment estimate and the non - central variance of the second - moment estimate to obtain the key parameters of the general large - model, is the key parameter update step size, which determines the step size of each parameter update. is a very small positive number (such as 1e - 8), used to prevent the denominator from being zero. Taking the identification of lung diseases as an example, in the lung disease identification model, the parameters are updated in this way in each iteration, enabling the model to gradually learn the relationship between lung image features and disease types, thereby minimizing the loss function and improving the accuracy of disease identification. For example, during the training process, when an input lung CT image is provided, the model first performs forward propagation to obtain the disease prediction result, then calculates the loss function between the prediction result and the true result data. Next, the gradient is calculated according to the formula of the above algorithm, and the first - order moment estimation mean and the second - order moment estimation non - central variance are updated. After bias correction, the model parameters are updated. After multiple iterations, the model can accurately identify the lung disease type and obtain a vertical - domain vision small model.
[0037] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A training method for a vertical domain visual model of a general large model, characterized in that, Including: Obtain a general large model and a vertical domain image training set, extract vertical domain image feature data from the vertical domain image training set according to the general large model, and output prediction classification result information according to the vertical domain image feature data; Obtain the true classification result information of the vertical domain image training set; Obtain a vertical domain prediction error rate evaluation value according to the true classification result and the prediction classification result information; Based on the vertical domain prediction error rate evaluation value, obtain the gradient of each parameter in the general large model through the backpropagation algorithm; Update the key parameters of the general large model based on the gradient until the prediction classification result information is equal to the true classification result information, and obtain a trained vertical domain vision small model.
2. The training method of the vertical domain visual model of the general large model according to claim 1, characterized in that, The step of extracting vertical domain image feature data from the vertical domain image training set according to the general large model and outputting prediction classification result information according to the vertical domain image feature data includes: Obtain the vertical domain image data of the vertical domain image training set; Preprocess the vertical domain image data; Construct a convolutional neural network of the general large model; Extract vertical domain image feature data through the convolutional layer of the convolutional neural network; Fuse the vertical domain image feature data through the fully connected layer of the convolutional neural network to obtain a comprehensive feature vector; Perform classification calculation on the comprehensive feature vector according to the Softmax classifier to obtain the prediction classification result information.
3. The training method of the vertical domain visual model of the general large model according to claim 1, characterized in that The step of obtaining a vertical domain prediction error rate evaluation value according to the true classification result and the prediction classification result information includes: Obtain the total number of categories of the prediction classification result information; Obtain the true result data of each category in the true classification result information; Obtain the prediction result data of each category in the prediction classification result information; Obtain the weight of the true result data of each category; Obtain the total number of samples in the vertical domain image training set; Obtain a vertical domain prediction error rate evaluation value according to the total number of categories, the true result data, the prediction result data, and the weight of the true result data.
4. The training method of the vertical domain visual model of the general large model according to claim 1, wherein Based on the vertical domain prediction error rate evaluation value, obtain the gradient of each parameter in the general large model through the backpropagation algorithm: Obtain the activation value of each neuron in the general large model; Obtain the input value of each neuron in the general large model; Obtain the partial derivative of the loss function in the general large model; Obtain the partial derivative of the weight in the general large model; Obtain the gradient of each parameter in the general large model according to the activation value, the input value, the partial derivative of the loss function, and the partial derivative of the weight.
5. The training method of the vertical domain visual model of the general large model according to claim 1, characterized in that The step of updating the key parameters of the general large model based on the gradient includes: Obtain the gradient; Update based on the gradient to obtain the first moment estimation mean and the second moment estimation non-central variance; Correct the first moment estimation mean and the second moment estimation non-central variance; Obtain the key parameters of the general large model according to the corrected first moment estimation mean and the second moment estimation non-central variance.
6. The training method of the vertical domain visual model of the general large model according to claim 5, characterized in that, The step of obtaining the first moment estimation mean by updating based on the gradient includes: Obtain the gradient; Obtain the first hyperparameter; Obtain the first moment estimate mean according to the first hyperparameter and the gradient.
7. The training method of the vertical domain visual model of the general large model according to claim 5, characterized in that, The step of obtaining the second moment estimate non-central variance based on the gradient update includes: Obtain the gradient; Obtain the second hyperparameter; Obtain the first moment estimate mean according to the second hyperparameter and the gradient.
8. The training method of the vertical domain visual model of the general large model according to claim 5, characterized in that The step of correcting the first moment estimate mean includes: Obtain the initial first moment estimate mean; Obtain the first hyperparameter; Obtain the first moment estimate mean according to the initial first moment estimate mean and the first hyperparameter.
9. The training method of the vertical domain visual model of the general large model according to claim 5, wherein The step of correcting the second moment estimate non-central variance includes: Obtain the initial second moment estimate non-central variance; Obtain the second hyperparameter; Obtain the second moment estimate non-central variance according to the initial second moment estimate non-central variance and the second hyperparameter.
10. The training method of the vertical domain visual model of the general large model according to claim 5, wherein The step of obtaining the key parameters of the general large model according to the corrected first moment estimate mean and the second moment estimate non-central variance includes: Obtain the corrected first moment estimate mean; Obtain the corrected second moment estimate non-central variance; Obtain the key parameter update step size; Obtain the key parameters of the general large model according to the first moment estimate mean, the second moment estimate non-central variance, and the key parameter update step size.
Citation Information
Patent Citations
CE-CNN-Adam-based medical record image automatic classification method
CN118710980A