Model Training Method, Apparatus, and Computer Device

By using the inference platform to extract feature and send it to the training platform for model training, the problem of inconsistent accuracy of the training and deployment end operators is solved, and the end-side inference accuracy of the model is improved.

CN119048868BActive Publication Date: 2025-06-27深圳市欧冶半导体有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411536767.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-06-27
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

The existing model training methods cannot ensure the consistency between the training and deployment ends in operator accuracy, resulting in a decrease in accuracy during model side inference.

Method used

By acquiring image preprocessing data, use the feature extraction model of the inference platform for feature extraction, and sending these features to the training platform for model training. After training, replace the feature extraction model of the inference platform to ensure model consistency in the training and inference stages.

Benefits of technology

By ensuring the consistency of the model in the training and inference stages, the accuracy reduction problem caused by operator accuracy differences is solved, and the accuracy and stability of the model during the end-side inference are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048868B_ABST
    Figure CN119048868B_ABST
Patent Text Reader

Abstract

The present application relates to a model training method, apparatus, and computer device. The method includes: obtaining image preprocessing data, extracting features from the image preprocessing data through a feature extraction model of an inference platform to obtain image features; sending the image features to a training platform through the inference platform; training a model to be trained on the training platform based on the image features to obtain a trained model that has been completed on the training platform, where the feature extraction model and the model to be trained have the same structure; replacing the feature extraction model of the inference platform with the trained model that has been completed to obtain a trained model that has been completed on the inference platform. Using this method can solve the problem of accuracy drop during model-side inference. And it provides a mechanism for continuous optimization of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence computing vision, and particularly to a model training method, apparatus, and computer device. Background Art

[0002] Training platforms and inference platforms are indispensable in realizing the verification and online deployment of artificial intelligence application models. Among them, the training platform is used for the training and optimization of models, emphasizing high-performance computing. The inference platform is used for the deployment and inference of models, emphasizing efficient operation and low power consumption. By separately constructing the training platform and the inference platform and giving full play to their respective advantages, it is ensured that the entire life cycle of the model from development to deployment can be carried out efficiently and stably.

[0003] However, although the inference platform and the training platform are consistent in the calculation and design logic of the forward operator of the model, since the inference platform usually adopts some optimization means in order to save power consumption and achieve hardware acceleration, these optimization means may lead to a decrease in the accuracy of the operator, thereby affecting the accuracy of the model during end-side inference. Therefore, the biggest defect of the existing model training method is that it cannot ensure the consistency of the operator accuracy between the training side and the deployment side, thus causing the problem of accuracy drop during the end-side inference of the model. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a model training method, apparatus, computer device, computer-readable storage medium, and computer program product that can eliminate the accuracy loss during inference caused by the differences between the inference platform and the training platform operators.

[0005] In a first aspect, the present application provides a model training method, including:

[0006] Obtain image preprocessing data, and perform feature extraction on the image preprocessing data through a feature extraction model of an inference platform to obtain image features;

[0007] Send the image features to a training platform through the inference platform;

[0008] Train a model to be trained on the training platform based on the image features to obtain a trained model on the training platform, where the feature extraction model and the model to be trained have the same structure;

[0009] Replace the feature extraction model of the inference platform with the trained model to obtain a trained model on the inference platform.

[0010] In one embodiment, the performing feature extraction on the image preprocessing data through a feature extraction model of an inference platform to obtain image features includes:

[0011] Through multiple processing layers in the feature extraction model of the inference platform, the image preprocessing data is successively subjected to feature extraction to obtain the image features output by each of the multiple processing layers.

[0012] In one embodiment, training the model to be trained on the training platform based on the image features to obtain the target feature extraction model of the training platform includes:

[0013] Obtain the image label corresponding to the image preprocessing data, as well as the parameters of each of the multiple processing layers; determine the gradients of each of the multiple processing layers according to the image label, the parameters of each of the multiple processing layers, and the image features output by each; update the parameters of the multiple processing layers in the model to be trained according to the gradients of each of the multiple processing layers to obtain the trained model of the training platform.

[0014] In one embodiment, determining the gradients of each of the multiple processing layers according to the image label, the parameters of each of the multiple processing layers, and the image features output by each includes:

[0015] Determine the loss function of the model to be trained according to the image label and the image features output by each of the multiple processing layers; determine the gradients of each of the multiple processing layers according to the loss function, the parameters of each of the multiple processing layers, and the image features output by each.

[0016] In one embodiment, obtaining the image preprocessing data includes:

[0017] The training platform sends the original image to the inference platform; the inference platform preprocesses the original image to obtain image preprocessing data; the data extraction tool on the inference platform extracts the image preprocessing data; the image preprocessing data extracted by the data extraction tool is sent to the training platform through a preset communication interface, and the preset communication interface is the data transmission interface between the training platform and the inference platform.

[0018] In one embodiment, preprocessing the original image by the inference platform to obtain image preprocessing data includes:

[0019] According to the brightness value range of the original image, perform normalization processing on the original image to obtain normalized data; determine quantization parameters according to the brightness value range of the original image; perform quantization processing on the normalized data according to the quantization parameters to obtain first target data; perform inverse quantization processing on the first target data according to the quantization parameters to obtain image preprocessing data.

[0020] In one embodiment, extracting the image preprocessing data by means of the data extraction tool on the inference platform to obtain the image preprocessing data includes:

[0021] Adding the image preprocessing data and a preset value through an addition operator to obtain a second target data; quantizing the second target data according to the quantization parameter to obtain a third target data; and dequantizing the third target data according to the quantization parameter to obtain the image preprocessing data.

[0022] In one embodiment, obtaining the image preprocessing data includes:

[0023] Preprocessing the original image through a preprocessing tool deployed on the training platform to obtain the image preprocessing data.

[0024] In a second aspect, the present application further provides a model training device, including:

[0025] An extraction module, configured to obtain image preprocessing data, and extract features of the image preprocessing data through a feature extraction model of the inference platform to obtain image features;

[0026] A sending module, configured to send the image features to the training platform through the inference platform;

[0027] A training module, configured to train a model to be trained on the training platform based on the image features to obtain a trained model completed on the training platform, where the feature extraction model and the model to be trained have the same structure;

[0028] A replacement module, configured to replace the feature extraction model of the inference platform with the trained model to obtain a trained model completed on the inference platform.

[0029] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0030] Obtaining image preprocessing data, and extracting features of the image preprocessing data through a feature extraction model of the inference platform to obtain image features;

[0031] Sending the image features to the training platform through the inference platform;

[0032] Training a model to be trained on the training platform based on the image features to obtain a trained model completed on the training platform, where the feature extraction model and the model to be trained have the same structure;

[0033] Replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0034] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0035] Obtain image preprocessing data, and perform feature extraction on the image preprocessing data through the feature extraction model of the inference platform to obtain image features;

[0036] Send the image features to the training platform through the inference platform;

[0037] Train the model to be trained on the training platform based on the image features to obtain the trained model of the training platform. The feature extraction model and the model to be trained have the same structure;

[0038] Replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0039] Fifthly, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0040] Obtain image preprocessing data, and perform feature extraction on the image preprocessing data through the feature extraction model of the inference platform to obtain image features;

[0041] Send the image features to the training platform through the inference platform;

[0042] Train the model to be trained on the training platform based on the image features to obtain the trained model of the training platform. The feature extraction model and the model to be trained have the same structure;

[0043] Replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0044] The above model training method, device, computer device, computer-readable storage medium, and computer program product obtain image preprocessing data, extract features from the image preprocessing data through a feature extraction model of an inference platform to obtain image features; through the preliminary processing of the inference platform, the influence of invalid or low-quality data on the training process can be effectively reduced, and the training efficiency can be improved. The image features are sent to the training platform through the inference platform; only the preliminarily processed image features rather than the original images are transmitted, which can significantly reduce the data transmission volume and improve the overall efficiency of the system. Based on the image features, the model to be trained on the training platform is trained to obtain the trained model on the training platform. The feature extraction model and the model to be trained have the same structure; by using a model with the same structure as the inference platform for training, the consistency of the models in the training stage and the inference stage is ensured, which helps to reduce the performance gap caused by the model structure difference. And training based on the image features processed by the inference platform can enable the model to better adapt to the actual application environment and improve the practicability and accuracy of the model. The feature extraction model of the inference platform is replaced with the trained model to obtain the trained model of the inference platform. Replacing the original inference model with a new model optimized by the training platform can directly improve the performance of the inference platform, especially in terms of accuracy, which helps to solve the problem of accuracy drop during model-side inference. And it provides a continuous optimization mechanism for the model. As the training data increases and technology advances, the inference model can be continuously updated and iterated to maintain the best state. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0046] Figure 1 It is an application environment diagram of the model training method in an embodiment;

[0047] Figure 2 It is a schematic flowchart of the model training method in an embodiment;

[0048] Figure 3 It is a schematic diagram of the principle of obtaining image preprocessing data in an embodiment;

[0049] Figure 4 It is a schematic diagram of the principle of the image preprocessing data extraction tool on the inference platform in an embodiment;

[0050] Figure 5 Schematic diagram of the principle of image preprocessing on a training platform in another embodiment;

[0051] Figure 6 Schematic diagram of the principle of a model training method in one embodiment;

[0052] Figure 7 Schematic flowchart of a model training method in another embodiment;

[0053] Figure 8 Structural block diagram of a model training device in one embodiment;

[0054] Figure 9 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0055] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] The model training method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The terminal 102 generates a model training request and sends the model training request to the server 104, so that the server 104 replaces the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, a smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0057] In an exemplary embodiment, as shown in Figure 2 , a model training method is provided. Taking the method applied to the server 104 in Figure 1 as an example, the method includes the following steps 202 to 208. Among them:

[0058] Step 202: Obtain the preprocessed image data, and extract features from the preprocessed image data through the feature extraction model of the inference platform to obtain image features.

[0059] Among them, the preprocessed image data refers to the image data that has been preprocessed and is used for training or testing a machine learning model. These images can come from different scenarios, such as surveillance video frames, medical images, natural landscape photos, etc., depending on the application scenario of the model.

[0060] The inference platform refers to the place where the model actually runs for inference after model training is completed. The inference platform is usually designed with dedicated hardware acceleration modules and optimized model operators (operators) to improve the inference performance of the model and reduce power consumption.

[0061] The feature extraction model refers to the model deployed on the inference platform, which is responsible for performing preliminary feature extraction on the input preprocessed image data. Optionally, the feature extraction model on the inference platform has the same structure as the model to be trained on the training platform, ensuring the consistency between the training and inference stages.

[0062] Image features refer to the useful information extracted from the preprocessed image data after being trained by the feature extraction model. These information usually exist in the form of vectors or tensors, reflecting certain characteristics or patterns of the image.

[0063] Specifically, first collect the preprocessed image data from the actual application scenario. These preprocessed image data can come from different sources, such as photos taken by cameras, network pictures, etc. And preprocess the collected preprocessed image data to obtain the preprocessed preprocessed image data. Then transmit the preprocessed preprocessed image data to the inference platform. Use the feature extraction model deployed on the inference platform to process the preprocessed image data and extract the features of the image. Optionally, the structure of this model is the same as the model to be trained on the training platform to ensure the logical consistency of feature extraction between the two. Finally, after the feature extraction model finishes training the preprocessed image data, it outputs a feature vector or tensor. Store the extracted features for subsequent transmission to the training platform.

[0064] Step 204: Send the image features to the training platform through the inference platform.

[0065] Among them, the training platform refers to the computing environment for model training and optimization, usually equipped with high-performance computing resources, such as GPUs (Graphics Processing Units), TPUs (Tensor Processing Units), etc., to support large-scale data processing and complex model training tasks.

[0066] Specifically, first store the extracted image features in the memory or cache of the inference platform to ensure the integrity and accuracy of the data. Then ensure that there is a stable communication channel between the inference platform and the training platform. For example, establish a communication interface between the inference platform and the training platform, such as through a network protocol (such as TCP / IP) or other dedicated communication protocols. Conduct a connection test to ensure that the communication channel is unobstructed and that data will not be lost or damaged during transmission. Next, package the image features stored in the inference platform into a data format suitable for transmission. Through the established communication interface, send the packaged image features from the inference platform to the training platform. After the training platform receives the data, return a confirmation message to ensure the successful transmission of the data. Finally, after the training platform receives the data, unpack the data and restore it to the image features. Verify the unpacked image features to ensure the integrity and accuracy of the data. Store the verified image features in the memory or storage device of the training platform for use in model training.

[0067] Step 206: Train the model to be trained on the training platform based on the image features to obtain the trained model on the training platform. The feature extraction model and the model to be trained have the same structure.

[0068] Among them, the model to be trained refers to the model deployed on the training platform, which is used to receive the image features transmitted from the inference platform and perform training based on these features.

[0069] The trained model refers to the finally obtained feature extraction model after being optimized by the training platform. This model will be deployed back to the inference platform to replace the original feature extraction model.

[0070] Specifically, first define a model to be trained with the same structure as the feature extraction model on the training platform. Ensure that the parameters such as the network structure, number of layers, and activation function of the two are exactly the same. Initialize the parameters of the model to be trained, which can be randomly initialized or the parameters of a pre-trained model. Then load the stored image features into the training environment. Configure the training parameters of the model to be trained, such as the learning rate, loss function, optimizer, etc. Write or call a training script to define the training process, including steps such as forward propagation, backward propagation, and parameter update. Next, input the image features into the model to be trained for forward propagation and calculate the output of the model. Calculate the value of the loss function based on the output of the model and the label (if any). Through the backpropagation algorithm, calculate the gradients of each layer of the model. Update the parameters of the model according to the calculated gradients to minimize the loss function. Repeat the steps of forward propagation, loss calculation, backward propagation, and parameter update until the model converges or reaches the predetermined number of training epochs. Finally, save the parameters and structure of the trained model to be trained to form the trained model. Verify the saved trained model to ensure that its performance is consistent with that on the training platform.

[0071] Step 208: Replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0072] Specifically, first on the training platform, save the parameters and structure of the trained model to form a transferable model file. Verify the saved trained model to ensure the integrity and accuracy of the model file.

[0073] Then ensure there is a stable communication channel between the training platform and the inference platform. For example, establish a communication interface between the training platform and the inference platform, such as through a network protocol (e.g., TCP / IP) or other dedicated communication protocols. Conduct connection tests to ensure the communication channel is unobstructed and data will not be lost or damaged during transmission.

[0074] Next, package the saved trained model file into a data format suitable for transmission. Through the established communication interface, send the packaged trained model from the training platform to the inference platform. After the inference platform receives the data, return a confirmation message to ensure successful data transmission.

[0075] After the inference platform receives the data, unpack the data and restore it to the trained model file. Verify the unpacked trained model to ensure the integrity and accuracy of the data. Store the verified trained model in the memory or storage device of the inference platform for model replacement.

[0076] Finally, deactivate the feature extraction model on the inference platform to ensure it no longer participates in the inference task. Load the received trained model into the memory of the inference platform to ensure it can run properly. Verify the loaded trained model to ensure its performance on the inference platform is consistent with that on the training platform. Activate the trained model to make it participate in the inference task, and continuously optimize the model based on the feedback in actual applications to further improve its performance.

[0077] In one embodiment, through multiple processing layers in the feature extraction model of the inference platform, sequentially perform feature extraction on the preprocessed image data to obtain image features output by each processing layer.

[0078] Among them, the processing layer refers to a functional module or neural network layer in the feature extraction model, which is responsible for performing specific processing or transformation on the input data. Each processing layer is responsible for performing specific processing or transformation on the input data, and gradually extracts high-level features.

[0079] Specifically, first collect image preprocessing data from the actual application scenario to ensure the representativeness of the samples. Perform necessary preprocessing on the image preprocessing data, such as normalization, size adjustment, etc., to make it meet the requirements of the model input. Then load the feature extraction model into the memory of the inference platform to ensure the normal operation of the model. Verify the loaded model to ensure the integrity and accuracy of the model. Next, send the preprocessed image preprocessing data into the input layer of the feature extraction model. Start the forward propagation process of the model, and pass the image preprocessing data layer by layer to each processing layer of the model. Finally, at the output nodes of each processing layer, extract the output features of that layer. Store the output features of each processing layer for subsequent model training.

[0080] Since the multi-level features of the image are gradually extracted through multiple processing layers, the richness and diversity of the features are ensured. Each processing layer can capture features at different scales and abstraction levels, from low-level edge and texture features to high-level semantic features, which helps the model to understand the image content more comprehensively.

[0081] In one embodiment, obtain the image label corresponding to the image preprocessing data, and the parameters of each of the multiple processing layers; determine the gradients of each of the multiple processing layers according to the image label, the parameters of each of the multiple processing layers, and the image features output by each of them; update the parameters of the multiple processing layers in the model to be trained according to the gradients of each of the multiple processing layers to obtain the trained model of the training platform.

[0082] Among them, the image label refers to the label information corresponding to the image preprocessing data, which represents the content or category of the image. For example, in an object detection task, the image label may contain the position information and category information of each object in the image; in a classification task, the image label may be the category to which the image belongs. These labels serve as supervision information during the model training process to help the model learn how to correctly predict the labels of the input data.

[0083] The parameter refers to the initial value of the parameters (weights and biases) in each processing layer (such as convolutional layer, fully connected layer, etc.) before the model starts training.

[0084] The gradient refers to the derivative of the loss function with respect to the model parameters. It indicates the increasing direction of the loss function in the parameter space. During the training process, by calculating the gradient of the loss function with respect to the model parameters, it can be known how to adjust these parameters to reduce the value of the loss function, that is, to improve the performance of the model.

[0085] Specifically, first, prepare a set of image pre - processed data and their corresponding labels. These labels are used for supervised learning to guide the model to learn the correct mapping relationship. Before the model starts training, each processing layer (such as convolutional layer, fully - connected layer, etc.) needs to initialize its parameters (weights and biases).

[0086] Then, input the image pre - processed data into the model, and perform forward propagation through each processing layer in the model (each layer uses its parameters) to calculate the image features output by each processing layer. This process will generate a series of intermediate outputs until the last layer of the model, generating the final prediction output.

[0087] Next, compare the prediction output of the model with the true label of the image and calculate the value of the loss function. The loss function measures the gap between the model prediction and the actual label. The loss function can be cross - entropy loss, mean - squared error loss, etc., which can be specifically set according to actual needs.

[0088] Then, use the backpropagation algorithm to calculate the gradient of each processing layer according to the derivative of the loss function with respect to the model parameters. The gradient represents the increasing direction of the loss function in the parameter space and is used to guide the update direction of the parameters. Using the calculated gradient information, update the parameters of each processing layer through gradient descent or its variants (such as Adam, RMSprop, etc.). The update rule is usually to adjust the parameter values along the negative direction of the gradient at a certain learning rate to reduce the value of the loss function.

[0089] Finally, continuously execute the loop of forward propagation, calculating loss, calculating gradient, and parameter update until the performance metric of the model (such as accuracy on the validation set) no longer improves significantly or reaches the preset maximum number of iterations. Eventually, obtain a trained model that is trained well on the training platform.

[0090] Since the model inference is implemented on the inference platform and the intermediate results of the inference (including the output of the last layer) are transmitted to the training platform for gradient update, the problem of precision loss caused by the operator differences between the training platform and the inference platform is solved. This method ensures that even if there are subtle differences in the operator implementation, the model training can be based on the actual behavior of the inference platform, thereby improving the robustness and accuracy of the model. Through the above steps, the model training not only considers the characteristics of the training platform but also fully considers the characteristics of the inference platform, making the training process closer to the actual application scenario. This helps to improve the performance of the model in the actual deployment environment and reduce the precision degradation caused by environmental differences.

[0091] In one embodiment, a loss function of the model to be trained is determined according to the image tags and the image features respectively output by multiple processing layers; according to the loss function, the parameters of each of the multiple processing layers, and the image features output by each of them, the gradients of each of the multiple processing layers are determined.

[0092] Among them, the loss function is a mathematical function that takes the predicted output of the model and the true label as inputs and outputs a scalar value, which reflects the quality of the model's prediction.

[0093] Specifically, first, the preprocessed image data is input into the model to be trained, and forward propagation is performed through each processing layer in the model to calculate the image features output by each processing layer. This process generates a series of intermediate outputs until the last layer of the model, generating the final predicted output.

[0094] Then, according to the final output of the model to be trained (i.e., the output of the last processing layer) and the true label of the image, the value of the loss function is calculated. The loss function measures the difference between the prediction of the model to be trained and the true label. Commonly used loss functions include cross-entropy loss, mean squared error loss, etc. The specific calculation formula is as follows:

[0095] L = loss(y*, y)

[0096] Among them, y* is the predicted output of the model to be trained, y is the true label of the preprocessed image data, and loss is the loss function.

[0097] Next, using the backpropagation algorithm, the gradient of each processing layer is calculated according to the derivative of the loss function with respect to the model parameters. The gradient represents the increasing direction of the loss function in the parameter space and is used to guide the update direction of the parameters. The specific calculation process is as follows:

[0098] Starting from the loss function, the gradient of each processing layer is calculated layer by layer forward. This is usually achieved through the Chain Rule, that is: & L \& wi = (& L \& zi ) × (& zi \& wi )

[0099] Among them, wi is the parameter of the i-th processing layer, zi is the output of the i-th processing layer, & L \& zi is the gradient of the loss function with respect to the output of the i-th processing layer, and &zi\&wi is the gradient of the output of the i-th processing layer with respect to the parameter.

[0100] Finally, using the calculated gradient information, update the parameters of each processing layer through gradient descent or its variants (such as Adam, RMSprop, etc.). The update rule is usually to adjust the parameter values along the negative direction of the gradient by a certain learning rate to reduce the value of the loss function. Continuously execute the loop of forward propagation, calculating the loss function, calculating the gradient, and parameter update until the performance metrics of the model (such as the accuracy on the validation set) no longer improve significantly or reach the preset maximum number of iterations. Finally, obtain the trained model that has been trained on the training platform. The specific update formula is as follows:

[0101] wi = wi - μ × (∂L / ∂wi)

[0102] where μ is the learning rate, and ∂L / ∂wi is the gradient of the i-th processing layer.

[0103] Since through the precise loss function, the calculation of the gradient is more accurate, which can more effectively guide the update of the model parameters, thereby improving the training accuracy of the model. And the model can converge within fewer iterations, improving the training efficiency. Also, by reasonably utilizing the hardware acceleration capabilities of the inference platform, the efficiency of data preprocessing and model inference is improved, thus accelerating the overall training process.

[0104] In one embodiment, the original image is sent to the inference platform through the training platform; the original image is preprocessed by the inference platform to obtain image preprocessing data; the image preprocessing data is extracted by the data extraction tool on the inference platform; the image preprocessing data extracted by the data extraction tool is sent to the training platform through a preset communication interface, and the preset communication interface is the data transmission interface between the training platform and the inference platform.

[0105] Among them, the original image refers to the original input image without any preprocessing. These images are the starting point of model training and are usually the original picture files (such as JPEG, PNG, etc.) obtained from the dataset

[0106] The preset communication interface refers to the interface used for data transmission between the training platform and the inference platform. This interface is responsible for transmitting data between the two platforms to ensure that the data can be efficiently and accurately transmitted back from the inference platform to the training platform.

[0107] Specifically, first, the training platform sends the original image to the inference platform through a preset communication interface. Then, the inference platform performs a series of preprocessing operations on the original image, including but not limited to: decoding the image file format (such as JPEG, PNG) into pixel data. Adjusting the size of the image to ensure it meets the requirements of the model input (such as cropping, scaling). Normalizing the preprocessed image data to a specific range (such as between 0 and 1) for better model processing. After the preprocessing, the preprocessed image data is ready to be input into the model. Next, through a data extraction tool, which is a model running on the inference platform, the preprocessed image data is extracted from the inference platform. Understandably, the extracted preprocessed image data will be used for model training and optimization. Finally, the preprocessed image data extracted by the data extraction tool is sent back to the training platform through the preset communication interface. After receiving these preprocessed image data, the training platform can use them for model training and optimization.

[0108] In one example, referring to Figure 3 , by establishing an efficient communication interface, namely the preset communication interface, a stable link between the training platform and the inference platform can be achieved. Deploy the already designed model extraction tool, namely the data extraction tool, to the inference platform. When the training platform has sufficient storage space, the preprocessing results of all data can be saved before the training starts. In this case, in the first epoch of training, the inference platform will participate in the data preprocessing process, and in the subsequent training process, the training platform can directly load the previously saved preprocessing results for model training. If the storage space of the training platform is not sufficient to save all the preprocessed data, then continuous cooperation between the training platform and the inference platform is required. In this case, during each training iteration, that is, each batch, the training platform will send the original data to the inference platform for preprocessing, and then return the preprocessed data to the training platform for model training.

[0109] Since the inference platform usually has a dedicated hardware acceleration module and optimized preprocessing algorithms, it can process the preprocessed image data more precisely. This ensures higher quality of the preprocessed image data, which helps improve the accuracy of model training. And by introducing the data preprocessing and inference results of the inference platform during the training process, the collaborative work between the training platform and the inference platform is achieved. This not only improves the efficiency of model training but also simplifies the model deployment process. And it reduces the debugging cost caused by platform differences, making it easier for the model to achieve the expected performance during deployment.

[0110] In one embodiment, the original image is normalized according to the brightness value range of the original image to obtain normalized data; the quantization parameter is determined according to the brightness value range of the original image; the normalized data is quantized according to the quantization parameter to obtain first target data; and the first target data is dequantized according to the quantization parameter to obtain image preprocessing data.

[0111] Herein, the brightness value range refers to the value range of the pixel brightness values in the image. For a common RGB image, the brightness value of each pixel is usually between 0 and 255, where 0 represents black and 255 represents white. The normalized data refers to converting the brightness values of the original image to a specific range through a certain method, usually between 0 and 1. The quantization parameter is a key parameter used to adjust the value range during the quantization process. In this application, the quantization parameter is 1 / 255, which means that each channel value of the input data (such as the pixel value in an RGB image) will be multiplied by this coefficient for quantization processing. The first target data refers to the data after quantization processing. Quantization processing converts floating-point numbers to fixed-point or low-precision values, usually used to reduce the storage requirements and computational complexity of the data.

[0112] Specifically, first, the brightness values of the original image are normalized to between 0 and 1 to obtain normalized data A*, and the specific calculation formula is as follows:

[0113] A* = A / 255

[0114] where A* is the normalized data, A is the original image, and 255 is the brightness value range of the original image.

[0115] Then, according to the brightness value range, the quantization coefficient is determined to be 1 / 255, and the quantization range is from -128 to 127. It can be understood that the quantization parameter mainly includes two parts: the quantization coefficient and the quantization range.

[0116] Next, the normalized data A* is quantized to values between -128 and 127 to obtain first target data B, and the specific calculation formula is as follows:

[0117] B = round(A* × 255 - 128)

[0118] where B is the first target data, A* is the normalized data, and 1 / 255 is the quantization coefficient.

[0119] Finally, the quantized data B is restored to image preprocessing data C to ensure that data C is equal to A*, and the specific calculation formula is as follows:

[0120] C = (B + 128) / 255

[0121] Among them, B is the first target data, C is the image preprocessing data, and 1 / 255 is the quantization coefficient.

[0122] Since the normalized data is quantized to values between -128 and 127, it can significantly reduce the storage requirements and computational complexity of the data. This is particularly important for data processing and model inference on inference platforms with limited resources. Through the dequantization process, the quantized data is restored to the original normalized data, ensuring that the accuracy of the data is not lost during the quantization and dequantization processes. And by using the same normalization, quantization, and dequantization parameters between the inference platform and the training platform, the consistency of data processing is ensured, improving the robustness and reliability of the model.

[0123] In one embodiment, the image preprocessing data and a preset value are added through an addition operator to obtain second target data; according to the quantization parameter, the second target data is quantized to obtain third target data; according to the quantization parameter, the third target data is dequantized to obtain the image preprocessing data.

[0124] Among them, the addition operator refers to a basic mathematical operation used to add two numerical values. In deep learning frameworks, the addition operator is usually used to perform element-wise addition on tensors.

[0125] The preset value refers to a fixed numerical value, usually used as one of the operands in the addition operator. In this application, the preset value is 0. It can be understood that in this application, the preset data is 0. By adding the image preprocessing data and 0 through the addition operator, the image preprocessing data is actually not changed, but it can ensure the consistency and integrity of the data and avoid introducing other irrelevant information.

[0126] The second target data refers to the data after being processed by the addition operator. In this application, the second target data is the result obtained by adding the image preprocessing data and the preset value 0.

[0127] The third target data refers to the second target data after being quantized. In this application, the third target data is the second target data quantized to values between -128 and 127.

[0128] Specifically, first, the image preprocessing data C and the preset value 0 are added to obtain the second target data D. The specific calculation formula is as follows:

[0129] D = C + 0

[0130] Among them, D is the second target data, and C is the image preprocessing data.

[0131] Then, the second target data D is quantized to a value between -128 and 127 to obtain the third target data E. The specific calculation formula is as follows:

[0132] E = round(D × 255 - 128)

[0133] Where E is the third target data, D is the second target data, and 1 / 255 is the quantization coefficient.

[0134] Finally, the third target data E is restored to the image preprocessing data F, which is the original normalized data. The specific calculation formula is as follows:

[0135] F = (E + 128) / 255

[0136] Where F is the image preprocessing data, E is the third target data, and 1 / 255 is the quantization coefficient.

[0137] In one example, referring to Figure 4 , the process of extracting the image preprocessing data through the data extraction tool on the inference platform is as follows:

[0138] First, the input data A is already normalized data with a normalization coefficient of 1 / 255, which means the input data has been converted into floating-point numbers between 0 and 1. Then, these normalized data A are sent to the QuantizeLinear function (linear quantization function) for quantization processing. In this process, using the quantization coefficient 1 / 255, these data are converted into integer values between -128 and 127. This quantization process essentially converts floating-point numbers into integers with lower precision while keeping the information of the original data as little loss as possible. The quantized data is denoted as B, which is equivalent to adding an offset of -128 to the original RGB image values.

[0139] Then, the quantized data B is subsequently sent to the DequantizeLinear function (linear dequantization function), which is a dequantization process aiming to restore the quantized data to its original floating-point representation. By using the same quantization coefficient 1 / 255, the quantized data B is restored to the same normalized data C as the original input A, that is, C = A. This step ensures that the precision of the data will not be lost during the quantization and dequantization processes.

[0140] Next, the restored normalized data C is input into an Add operator (addition operator) and added to 0. Understandably, to avoid introducing other irrelevant information and ensure the purity of the data. By adding 0, we obtain the data D, and D is actually the same as C.

[0141] Then, send the data D into the QuantizeLinear function again for quantization using the same quantization coefficient 1 / 255. Since D is the same as C, and C is the restored version of the original input A, the result E of this quantization process should be the same as the result B of the first quantization process.

[0142] Finally, send the re-quantized data E into the DequantizeLinear function for dequantization using the same quantization coefficient 1 / 255. The dequantized data F is equal to the original input A, which proves the reversibility and consistency of the quantization and dequantization processes, ensuring the consistency and precision of the data at different processing stages.

[0143] Due to precise addition operators, quantization, and dequantization processes, the model can converge within fewer iterations, improving the training efficiency. And by reasonably utilizing the hardware acceleration capabilities of the inference platform, the efficiency of data processing and model inference is improved, thus accelerating the overall training process.

[0144] In one embodiment, the original image is preprocessed by a preprocessing tool deployed on the training platform to obtain preprocessed image data.

[0145] The preprocessing tool refers to a piece of program code, usually written in programming languages such as Python and C++, for performing a series of preprocessing operations on the original image. These preprocessing operations include, but are not limited to, decoding, size transformation, normalization, etc., with the aim of converting the original image into a data format suitable for model input.

[0146] Specifically, first deploy the preprocessing tool on the training platform. Then read the original image from the specified path through the preprocessing tool, and finally preprocess the original image to obtain the preprocessed image data. Optionally, the preprocessing process for the original image includes, but is not limited to: decoding the image file into pixel data, adjusting the size of the image to ensure it meets the requirements of model input, common size transformations including cropping and scaling, normalizing the preprocessed image data to a specific range, usually between 0 and 1, color space conversion, and data augmentation (such as rotation and flipping).

[0147] In one example, refer to Figure 5, first, the user deploys the software implementation of the data preprocessing module on the inference platform, i.e., the preprocessing tool (usually the software equivalent implementation of an optimized hardware acceleration module), to the training platform. This step ensures that the training platform can use the same data preprocessing logic as the inference platform. Then, on the training platform, the user needs to write or modify the training script to call the deployed software implementation of data preprocessing. This step ensures that during the model training process, the data preprocessing steps are consistent with those on the inference platform. Finally, the preprocessed data will be fed into the model for training. During the training process, the model will perform forward propagation and backward propagation based on the preprocessed data, gradually optimizing the model parameters until the model converges.

[0148] By using the preprocessing tool on the training platform, the consistency between the training data and the actual deployed data is ensured, simplifying the model deployment process and reducing the debugging cost caused by platform differences. Moreover, the preprocessed image preprocessing data can be directly used for model training, avoiding repeated preprocessing during training and inference, and improving the data transmission efficiency.

[0149] In one embodiment, refer to Figure 6 , first, the same model, i.e., the feature extraction model, is deployed on the training platform and the inference platform respectively. This ensures that the same model architecture and parameters are used for training and inference.

[0150] Then, on the inference platform, the data input in batch form is preprocessed, and then the preprocessed data is input into the feature extraction model for inference. During the inference process, the output results of each layer are saved. The output results of each layer (including the output of the last layer) saved on the inference platform are transmitted to the training platform. These intermediate results are used for subsequent loss calculation and gradient update.

[0151] Next, on the training platform, the output of the last layer of the model is compared with the ground truth (GT), i.e., the image label, to calculate the loss function (Loss). Using the calculated loss value, combined with the model parameters (such as weights) and the output results of each layer, the gradient is calculated through chain differentiation, and the model parameters are updated.

[0152] Finally, the updated model parameters are redeployed to the inference platform to ensure that the model parameters on the inference platform and the training platform are consistent. Repeat the above steps until the model converges on the training platform to meet the accuracy requirements.

[0153] In an exemplary embodiment, as Figure 7 shown, it includes steps 702 to 706. Among them:

[0154] Step 702: Preprocess the original image through a preprocessing tool deployed on the training platform to obtain preprocessed image data, or send the original image to the inference platform through the training platform; normalize the original image according to the brightness value range of the original image to obtain normalized data; determine quantization parameters according to the brightness value range of the original image; perform quantization processing on the normalized data according to the quantization parameters to obtain first target data; perform inverse quantization processing on the first target data according to the quantization parameters to obtain preprocessed image data; perform an addition operation on the preprocessed image data and a preset value through an addition operator to obtain second target data; perform quantization processing on the second target data according to the quantization parameters to obtain third target data; perform inverse quantization processing on the third target data according to the quantization parameters to obtain preprocessed image data; send the extracted preprocessed image data to the training platform through a preset communication interface, and the preset communication interface is the data transmission interface between the training platform and the inference platform;

[0155] Successively extract features from the preprocessed image data through multiple processing layers in the feature extraction model of the inference platform to obtain image features output by each of the multiple processing layers;

[0156] Step 704: Send the image features to the training platform through the inference platform;

[0157] Step 706: Obtain the image label corresponding to the preprocessed image data and the parameters of each of the multiple processing layers; determine the loss function of the model to be trained according to the image label and the image features output by each of the multiple processing layers; determine the gradients of each of the multiple processing layers according to the loss function, the parameters of each of the multiple processing layers, and the image features output by each of them; update the parameters of the multiple processing layers in the model to be trained according to the gradients of each of the multiple processing layers to obtain the trained model of the training platform. The feature extraction model and the model to be trained have the same structure;

[0158] Step 708: Replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0159] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0160] Based on the same inventive concept, an embodiment of the present application further provides a model training device for implementing the model training method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model training device provided below can refer to the limitations on the model training method in the above text, and will not be repeated here.

[0161] In an exemplary embodiment, as Figure 8 shown, a model training device 800 is provided, including: an extraction module 802, a sending module 804, a training module 806, and a replacement module 808, where:

[0162] The extraction module 802 is configured to obtain image preprocessing data, and perform feature extraction on the image preprocessing data through the feature extraction model of the inference platform to obtain image features;

[0163] The sending module 804 is configured to send the image features to the training platform through the inference platform;

[0164] The training module 806 is configured to train the model to be trained on the training platform based on the image features to obtain the trained model on the training platform. The feature extraction model and the model to be trained have the same structure;

[0165] The replacement module 808 is configured to replace the feature extraction model of the inference platform with the trained model to obtain the trained model of the inference platform.

[0166] In one of the embodiments, the extraction module 802 is configured to perform feature extraction on the image preprocessing data through multiple processing layers in the feature extraction model of the inference platform to obtain image features output by each of the multiple processing layers.

[0167] In one embodiment, the training module 806 is configured to obtain the image labels corresponding to the preprocessed image data and the parameters of each of the multiple processing layers; determine the gradients of each of the multiple processing layers according to the image labels, the parameters of each of the multiple processing layers, and the image features output by each of them; and update the parameters of the multiple processing layers in the model to be trained according to the gradients of each of the multiple processing layers, so as to obtain the trained model of the training platform.

[0168] In one embodiment, the training module 806 is configured to determine the loss function of the model to be trained according to the image labels and the image features output by each of the multiple processing layers; and determine the gradients of each of the multiple processing layers according to the loss function, the parameters of each of the multiple processing layers, and the image features output by each of them.

[0169] In one embodiment, the extraction module 802 is configured to send the original image to the inference platform through the training platform; preprocess the original image through the inference platform to obtain preprocessed image data; extract the preprocessed image data through the data extraction tool on the inference platform; and send the preprocessed image data extracted by the data extraction tool to the training platform through a preset communication interface, where the preset communication interface is the data transmission interface between the training platform and the inference platform.

[0170] In one embodiment, the extraction module 802 is configured to normalize the original image according to the brightness value range of the original image to obtain normalized data; determine quantization parameters according to the brightness value range of the original image; quantize the normalized data according to the quantization parameters to obtain first target data; and dequantize the first target data according to the quantization parameters to obtain preprocessed image data.

[0171] In one embodiment, the extraction module 802 is configured to add the preprocessed image data and a preset value through an addition operator to obtain second target data; quantize the second target data according to the quantization parameters to obtain third target data; and dequantize the third target data according to the quantization parameters to obtain preprocessed image data.

[0172] In one embodiment, the extraction module 802 is configured to preprocess the original image through a preprocessing tool deployed on the training platform to obtain preprocessed image data.

[0173] Each module in the above model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0174] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 9 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to model training. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a model training method.

[0175] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0176] In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0177] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0178] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0179] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0180] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0181] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0182] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A model training method, characterized in that: The method comprises: When the storage space of the training platform is sufficient, all original images are sent to the inference platform through the training platform. When the storage space of the training platform is insufficient, all original images are divided into multiple batches of original images. For each training iteration, the original images of the batches are sent to the inference platform through the training platform. The inference platform is provided with a hardware acceleration module and an optimization model operator. The hardware acceleration module and the optimization model operator are used to specifically improve the inference performance and preprocessing performance of the model. According to the brightness value range of the original image, the original image is normalized to obtain normalized data; according to the brightness value range of the original image, a quantization parameter is determined, and the quantization parameter is used to reduce the storage requirement and calculation complexity of the data; according to the quantization parameter, the normalized data is quantized to obtain first target data; according to the quantization parameter, the first target data is dequantized to obtain image preprocessing data; Extracting the image preprocessing data by using a data extraction tool on the inference platform; The image preprocessing data extracted by the data extraction tool is sent to the training platform through a preset communication interface, where the preset communication interface is a data transmission interface between the training platform and the reasoning platform, and feature extraction is performed on the image preprocessing data through a feature extraction model of the reasoning platform to obtain image features; Sending the image features to the training platform via the inference platform; Training the model to be trained of the training platform based on the image features to obtain a trained model of the training platform, wherein the feature extraction model and the model to be trained have the same structure; The feature extraction model of the reasoning platform is replaced with the trained model to obtain the trained model of the reasoning platform.

2. The method according to claim 1, characterized in that The feature extraction model of the inference platform is used to extract features from the image preprocessing data to obtain image features, including: Through the multiple processing layers in the feature extraction model of the inference platform, feature extraction is performed on the image preprocessing data in turn to obtain the image features output by each of the multiple processing layers.

3. The method according to claim 2, characterized in that The step of training the to-be-trained model of the training platform based on the image features to obtain the trained model of the training platform includes: Obtaining an image label corresponding to the image preprocessing data and respective parameters of the multiple processing layers; Determining the gradients of the multiple processing layers according to the image label, the parameters of the multiple processing layers and the image features of the respective outputs; According to the respective gradients of the multiple processing layers, the parameters of the multiple processing layers in the model to be trained are updated to obtain a trained model of the training platform.

4. The method according to claim 3, characterized in that Determining the gradients of the multiple processing layers according to the image label, the parameters of the multiple processing layers, and the image features of the outputs of the multiple processing layers comprises: Determining a loss function of the model to be trained according to the image label and the image features output by each of the multiple processing layers; The gradients of the multiple processing layers are determined according to the loss function, the parameters of the multiple processing layers, and the image features of the respective outputs.

5. The method according to claim 1, characterized in that The extracting the image preprocessing data by using a data extraction tool on the inference platform to obtain the image preprocessing data includes: The image preprocessing data and the preset value are added together by an addition operator to obtain second target data; quantizing the second target data according to the quantization parameter to obtain third target data; According to the quantization parameter, the third target data is subjected to inverse quantization processing to obtain image preprocessing data.

6. The method according to claim 1, characterized in that The original image is an original input image that has not been preprocessed.

7. The method according to claim 1, characterized in that The method for obtaining image preprocessing data also includes: The original image is preprocessed by the preprocessing tool deployed on the training platform to obtain image preprocessing data.

8. A model training device, characterized in that: The device comprises: The extraction module is used to send all original images to the inference platform through the training platform when the storage space of the training platform is sufficient. The inference platform is provided with a hardware acceleration module and an optimization model operator. When the storage space of the training platform is insufficient, all original images are divided into multiple batches of original images. For each training iteration, the original images of the batch are sent to the inference platform through the training platform. The hardware acceleration module and the optimization model operator are used to specifically improve the inference performance and preprocessing performance of the model; the original images are normalized according to the brightness value range of the original images to obtain normalized data; the quantization parameters are determined according to the brightness value range of the original images. The quantization parameters are The parameters are used to reduce the storage requirements and computational complexity of data; according to the quantization parameters, the normalized data is quantized to obtain first target data; according to the quantization parameters, the first target data is dequantized to obtain image preprocessing data; the image preprocessing data is extracted by the data extraction tool on the inference platform; the image preprocessing data extracted by the data extraction tool is sent to the training platform through a preset communication interface, the preset communication interface is a data transmission interface between the training platform and the inference platform, and the image preprocessing data is feature extracted by a feature extraction model of the inference platform to obtain image features; A sending module, used for sending the image features to the training platform through the inference platform; A training module, used for training the model to be trained of the training platform based on the image features to obtain a trained model of the training platform, wherein the feature extraction model and the model to be trained have the same structure; A replacement module is used to replace the feature extraction model of the reasoning platform with the trained model to obtain the trained model of the reasoning platform.

9. The device according to claim 8, characterized in that The extraction module is used to extract features from the image preprocessing data in sequence through multiple processing layers in the feature extraction model of the inference platform to obtain image features output by each of the multiple processing layers.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Embedded object cognition system based on image processing

    CN113158968A