Model training method and device, classification method and device, medium and program product

By introducing an uncertainty estimation module in model training and adjusting the loss weight, the problem that unlabeled data noise in semi-supervised learning affects the model accuracy is solved, and more efficient model training and noise processing is achieved.

CN120217034APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311837928.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the prior art uses semi-supervised learning method to train models, the noise in the unlabeled data leads to low model training accuracy, and due to the huge number of unlabeled data and complex noise, a single noise removal method is difficult to effectively eliminate noise.

Method used

A model training method is proposed. Through the coordinated training of feature extraction module, classification module and uncertainty estimation module, the pseudo-label and prediction tag of unlabeled data are used to calculate the loss, and the loss weight is adjusted based on the uncertainty value output by the uncertainty estimation module to better rely on unlabeled data with high certainty for training.

Benefits of technology

By adjusting the loss weight, the model can more effectively utilize unlabeled data with high certainty, thereby improving the model training accuracy and improving the model's robustness to noise data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217034A_ABST
    Figure CN120217034A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, a classification method and device, a medium and a program product, and relates to the field of artificial intelligence, and the method comprises the steps: inputting unlabeled data into a target model, so as to enable the unlabeled data to be input into a feature extraction module to obtain the features of the unlabeled data; the features of the unlabeled data are respectively input into the classification module and the uncertainty estimation module to obtain a prediction label of the unlabeled data and a first numerical value used for representing the uncertainty of the unlabeled data; obtaining a first loss of the classification module based on the pseudo label and the prediction label of the unlabeled data; obtaining a second loss of the uncertainty estimation module based on the first numerical value and the first loss; determining a weight corresponding to the first loss based on the first numerical value, and if the first numerical value is larger, the weight corresponding to the first loss is smaller; if the first numerical value is smaller, the weight corresponding to the first loss is larger; a target model is trained based on the first loss, the weight, and the second loss. Therefore, the model training precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of Artificial Intelligence (AI), and in particular, to a model training method, a classification method, a device, a medium, and a program product. Background Art

[0002] Semi-Supervised Learning (SSL) is a learning method that combines supervised learning and unsupervised learning. It can use labeled data and unlabeled data together to complete model training. Currently, for many fields where a large amount of labeled data cannot be obtained, such as the medical field, the semi-supervised learning method is often used.

[0003] Generally, the unlabeled data may contain noise, resulting in low model training accuracy. Currently, methods for eliminating noise from unlabeled data have been proposed. However, due to the large number of unlabeled data and the complex noise situation, using a relatively single denoising method cannot effectively eliminate noise, thus still resulting in low model training accuracy. Summary of the Invention

[0004] Embodiments of the present application provide a model training method, a classification method, a device, a medium, and a program product, which can improve the model training accuracy.

[0005] In a first aspect, embodiments of the present application provide a model training method. The target model includes a feature extraction module, a classification module, and an uncertainty estimation module. The method includes: inputting unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data; obtaining a first loss of the classification module based on the pseudo-labels and predicted labels of the unlabeled data; and obtaining a second loss of the uncertainty estimation module based on the first value and the first loss; determining the weight corresponding to the first loss based on the first value; training the target model based on the first loss, the weight corresponding to the first loss, and the second loss; where, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0006] In a second aspect, embodiments of the present application provide a classification method, including: obtaining data to be classified; inputting the data to be classified into the target model trained by the model training method provided in the first aspect or its various implementation manners to obtain the predicted labels of the data to be classified and a third value representing the uncertainty of the data to be classified.

[0007] In a third aspect, an embodiment of the present application provides a model training device. The target model includes a feature extraction module, a classification module, and an uncertainty estimation module. The model training device includes an input module, a calculation module, and a training module. Among them, the input module is used to input unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data. The calculation module is used to obtain a first loss of the classification module based on the pseudo-labels and predicted labels of the unlabeled data; and obtain a second loss of the uncertainty estimation module based on the first value and the first loss; determine the weight corresponding to the first loss based on the first value. The training module is used to train the target model based on the first loss, the weight corresponding to the first loss, and the second loss. Among them, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0008] In a fourth aspect, an embodiment of the present application provides a classification device, including an acquisition module and an input module. Among them, the acquisition module is used to acquire data to be classified. The input module is used to input the data to be classified into the target model trained by the model training method provided in the first aspect or its various implementation manners, to obtain the predicted label of the data to be classified and a third value representing the uncertainty of the data to be classified.

[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in any one of the first aspect to the second aspect or its various implementation manners.

[0010] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, and the computer program causes a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.

[0011] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer program instructions, and the computer program instructions cause a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.

[0012] In an eighth aspect, an embodiment of the present application provides a computer program, and the computer program causes a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.

[0013] Through the technical solution provided by this application, for unlabeled data with relatively high uncertainty, when calculating the loss of the classification module of the target model, the loss weight corresponding to this unlabeled data can be set relatively small. For unlabeled data with relatively low uncertainty, when calculating the loss of the classification module, the loss weight corresponding to this unlabeled data can be set relatively large, so that when training the target model, more reliance is placed on unlabeled data with relatively high certainty, thereby improving the model training accuracy. Brief Description of the Drawings

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] Figure 1 It is a schematic diagram of a system architecture related to an embodiment of this application;

[0016] Figure 2 It is a flowchart of a model training method provided by an embodiment of this application;

[0017] Figure 3 It is a schematic diagram of a target model provided by an embodiment of this application;

[0018] Figure 4 It is an input-output schematic diagram of a feature extraction module provided by an embodiment of this application;

[0019] Figure 5 It is an input-output schematic diagram of a classification module provided by an embodiment of this application;

[0020] Figure 6 It is an input-output schematic diagram of an uncertainty estimation module provided by an embodiment of this application;

[0021] Figure 7 It is a schematic diagram of a collaborative training process provided by an embodiment of this application;

[0022] Figure 8 It is a flowchart of another model training method provided by an embodiment of this application;

[0023] Figure 9 It is a flowchart of a classification method provided by an embodiment of this application;

[0024] Figure 10 It is a flowchart of a model training method provided by an embodiment of this application;

[0025] Figure 11 It is a flowchart of a classification method provided by an embodiment of this application;

[0026] Figure 12 Schematic diagram of a model training device 1200 provided by an embodiment of the present application;

[0027] Figure 13 Schematic diagram of a classification device 1300 provided by an embodiment of the present application;

[0028] Figure 14 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0032] The embodiments of the present application may relate to artificial intelligence technology, but are not limited thereto.

[0033] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0034] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the base model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0035] Among them, the embodiments of this application relate to pre-trained model technology.

[0036] The training device can pre-train the target model with training samples composed of labeled data and corresponding true labels to obtain the pre-trained target model. Subsequently, the pre-trained target model can be fine-tuned with training samples composed of unlabeled data and corresponding pseudo-labels. Therefore, in the embodiments of this application, the training process of the target model with training samples composed of unlabeled data and corresponding pseudo-labels is also referred to as the fine-tuning process of the target model.

[0037] The training device can also input unlabeled data into the pre-trained model to obtain the pseudo-labels of the unlabeled data. Among them, the pre-trained model can be the above-mentioned pre-trained target model or an open-source pre-trained model currently available. The embodiments of this application do not limit this.

[0038] Computer Vision Technology (CV) is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses cameras and computers to replace the human eye for target recognition, detection, and measurement, and further performs graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as swin-transformer, Vision Transformer (ViT), V-Model based Multi-Objective Evolutionary Algorithm (V-MOE), and Mean Absolute Error (MAE) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, Optical Character Recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3 Dimensions) technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0039] Among them, the embodiments of the present application may specifically relate to image processing. Specifically, when the data involved in the present application is an image, the classification module in the target model can classify the image. For example, when a medical image is input into the target model, a medical diagnosis result can be obtained through the classification module in the target model.

[0040] The relevant knowledge related to the present application will be elaborated below:

[0041] First, unlabeled data, also known as unmarked data or unlabeled data, refers to data without labels or tags, usually unlabeled images, texts, or other types of data. In machine learning, unlabeled data can be used to train semi-supervised learning models, etc. These models usually require a large amount of unlabeled data to improve their performance. The utilization of unlabeled data can greatly improve the generalization ability of the model, thereby improving the performance of the model.

[0042] Second, labeled data, also known as marked data or tagged data, refers to data that has been marked or labeled manually and is commonly used in fields such as machine learning, natural language processing, and image recognition. Labeled data is crucial for training models because the models need to learn the features and patterns extracted from the labeled data.

[0043] Third, pseudo-labels are labels used in semi-supervised learning. They are not obtained through manual annotation but are automatically generated by a certain method. For example, unlabeled data can be input into a pre-trained model to obtain the pseudo-labels of the unlabeled data.

[0044] Fourth, predicted labels refer to the prediction results of the model for the input data.

[0045] Fifth, true labels are data labels provided by manual annotation or a labeled dataset.

[0046] Sixth, semi-supervised learning is a learning method that combines supervised learning and unsupervised learning. It can use both labeled data and unlabeled data to jointly complete model training.

[0047] Seventh, a semi-supervised learning model is a model between supervised learning and unsupervised learning. It uses both unlabeled data and labeled data to jointly learn to improve the performance of the model.

[0048] Eighth, loss refers to the difference between the model's prediction result, i.e., the predicted label, and the true label. In machine learning, loss is used to measure the degree of the model's prediction error. Generally, the smaller the loss, the more accurate the model's prediction.

[0049] Next, the technical problems to be solved, the inventive concept, and the system architecture of the embodiments of the present application will be elaborated:

[0050] As mentioned above, generally, unlabeled data may contain noise, resulting in low accuracy in model training. Currently, methods for eliminating noise from unlabeled data have been proposed. However, due to the large number of unlabeled data and the complex noise situation, using a relatively single denoising method cannot effectively eliminate noise, thus still resulting in low accuracy in model training.

[0051] To solve the above technical problems, the embodiments of the present application propose that for unlabeled data with a high degree of uncertainty, when calculating the loss of the classification module of the target model, the loss weight corresponding to the unlabeled data can be set smaller. For unlabeled data with a low degree of uncertainty, when calculating the loss of the classification module, the loss weight corresponding to the unlabeled data can be set larger, so that when training the target model, more reliance is placed on the unlabeled data with a higher degree of certainty, thereby improving the accuracy of model training.

[0052] In some implementable manners, the system architecture of the embodiments of the present application is as Figure 1 shown.

[0053] Figure 1 FIG. is a schematic diagram of a system architecture related to the embodiments of the present application, including a user device 101, a data acquisition device 102, a training device 103, an execution device 104, a database 105, and a content library 106.

[0054] Among them, the data acquisition device 102 is used to read unlabeled data and the pseudo-labels corresponding to the unlabeled data from the content library 106. The data acquisition device 102 can store the read unlabeled data and the pseudo-labels corresponding to the unlabeled data into the database 105.

[0055] The training device 103 constructs training samples based on the unlabeled data and the pseudo-labels corresponding to the unlabeled data stored in the database 105. Further, the training device 103 can train a target model based on the training samples.

[0056] In addition, referring to Figure 1 , the execution device 104 is configured with an I / O interface 107 to interact with external devices. For example, the execution device 104 can receive the data to be classified sent by the user device 101 through the I / O interface, and the computing module 108 in the execution device 104 can input the data to be classified into the target model,

[0057] output the predicted label of the data to be classified, and send the predicted label of the data to be classified to the user device 101 through the I / O interface.

[0058] Among them, the user device 101 can include a mobile phone, a tablet computer, a notebook computer, a handheld computer, an MID, a desktop computer, or other terminal devices with a browser installation function.

[0059] The execution device 104 can be a server.

[0060] Exemplarily, the server can be a computing device such as a rack server, a blade server, a tower server, or a cabinet server. The server can be an independent test server or a test server cluster composed of multiple test servers.

[0061] In this embodiment, the execution device 104 is connected to the user device 101 through a network. The network can be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), the 4rd Generation (4G) network, the 5rd Generation (5G) network, Bluetooth, wireless fidelity (Wi-Fi), a call network, etc.

[0062] It should be understood that in the embodiments of the present application, if the target model is also pre-trained with labeled data and the true labels corresponding to the labeled data, then its training process can be similar to the training process of the target model with unlabeled data and the pseudo-labels corresponding to the unlabeled data as described above, and details will not be elaborated herein.

[0063] It should be noted that Figure 1 is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. In some embodiments, the above data acquisition device 102, user device 101, training device 103, and execution device 104 can be the same device. The above database 105 can be distributed on one server or multiple servers, and the above content library 106 can be distributed on one server or multiple servers.

[0064] The embodiments of the present application will be elaborated in detail below:

[0065] Figure 2 is a flowchart of a model training method provided by the embodiments of the present application. Among them, this method can be executed by a training device, for example, it can be Figure 1 the training device 103 in Figure 3 is a schematic diagram of the target model provided by the embodiments of the present application. As Figure 3 shown, the target model includes: a feature extraction module, a classification module, and an uncertainty estimation module. As Figure 2 shown, this model training method can include:

[0066] S210: Input the unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data.

[0067] It should be understood that the embodiments of this application mainly use the semi-supervised learning method to train the target model. Therefore, the target model can also be called a semi-supervised learning model.

[0068] In some implementable ways, the unlabeled data in the embodiments of this application can be images, texts, voices, etc., and the embodiments of this application do not limit this.

[0069] For example, the unlabeled data in the embodiments of this application can be medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI).

[0070] It should be understood that the feature extraction module can also be called a feature extraction network, a feature extraction unit, a feature extractor, etc., and it is mainly used to extract the features of data.

[0071] In some implementable ways, the feature extraction module can be a Deep Learning Network (DLN).

[0072] For example, the feature extraction module can be a Convolutional Neural Networks (CNN). Among them, the CNN can automatically learn the local features in the unlabeled data and form higher-level semantic features through multi-level feature combination and abstraction. These features can effectively capture the features of the unlabeled data. For example, they can capture information such as the structure and texture in the unlabeled image, which helps the subsequent classification module to use the features of the unlabeled data to obtain more accurate predicted labels, and helps the uncertainty estimation module to use the features of the unlabeled data to obtain a more accurate first value for ensuring the uncertainty of the unlabeled data.

[0073] In some implementable ways, transfer learning technology, etc., can be used to apply the feature extraction module trained in other fields or tasks to the target model provided by the embodiments of this application to improve the model training efficiency.

[0074] It should be understood that transfer learning is a machine learning technique commonly used to transfer knowledge between two different but related tasks. It improves the performance of the second task by applying the knowledge learned in one task to another task. In the embodiments of the present application, the feature extraction module trained in other fields or tasks is mainly transferred to the target model provided in the embodiments of the present application by using transfer learning technology.

[0075] In some implementable ways, the structure, parameters, etc. of the feature extraction module can be adjusted to adapt to different types of unlabeled data and application scenarios. For example, assuming that the original feature extraction module is adapted to CT, by adjusting the structure and / or parameters of the feature extraction module, the adjusted feature extraction module can be adapted to MRI.

[0076] In some implementable ways, before the feature extraction module extracts the features of the unlabeled data, the feature extraction module can first preprocess the unlabeled data, such as data augmentation, noise reduction, normalization, etc., to ensure the accuracy and stability of the features extracted subsequently.

[0077] It should be understood that the above feature extraction module may include a preprocessing unit for preprocessing the unlabeled data. Of course, the preprocessing unit can also be independent of the feature extraction module. For example, the preprocessing unit can be set before the feature extraction module, so that before the unlabeled data is input into the feature extraction module, the preprocessing unit first preprocesses the unlabeled data. In short, the embodiments of the present application do not limit the structure of the target model.

[0078] Figure 4 This is a schematic diagram of the input and output of the feature extraction module provided in the embodiments of the present application. As Figure 4 shown, the input of the feature extraction module can be unlabeled data or data preprocessed by the preprocessing unit, and the output of the feature extraction module is the features of the unlabeled data.

[0079] In some implementable ways, if the unlabeled data is an unlabeled image, the features of the unlabeled data include at least one of the following, but are not limited to: color features, texture features, shape features, and spatial relationship features of the unlabeled image.

[0080] Among them, the shape features may include contour features and region features.

[0081] It should be understood that the spatial relationship features refer to the mutual spatial positions or relative direction relationships between multiple objects segmented in the image.

[0082] In some implementable ways, if the unlabeled data is unlabeled text, the features of the unlabeled text include at least one of the following, but are not limited to: the language feature, content feature, structural feature, formal feature, and sentiment feature of the unlabeled text.

[0083] In some implementable ways, if the unlabeled data is unlabeled speech, the features of the unlabeled speech include at least one of the following, but are not limited to: the time-domain feature, frequency-domain feature, prosody feature, acoustic feature, and language feature of the unlabeled speech.

[0084] It should be understood that the classification module can also be referred to as a classification network, classification unit, classifier, etc. It is mainly used to classify data based on the features of the data, that is, to obtain the predicted label of the data.

[0085] In some implementable ways, the classification module can be a deep learning network, specifically a convolutional neural network, but is not limited to this.

[0086] In some implementable ways, the classification module can include: a convolutional layer, a pooling layer, a fully connected layer, and an output layer. Among them, the convolutional layer, pooling layer, and fully connected layer are used to learn feature representations at different levels.

[0087] In some implementable ways, the output layer can adopt the softmax function, which can compress the K-dimensional vector containing arbitrary real numbers output from the fully connected layer into another K-dimensional real vector, so that the range of each element in the compressed K-dimension is between 0 and 1, and the sum of all elements is 1, where K is the number of categories.

[0088] Figure 5 This is the input-output schematic diagram of the classification module provided by the embodiments of this application. As Figure 5 shown, the input of the classification module can be the features of the unlabeled data, and the output is the predicted label of the unlabeled data.

[0089] In some implementable ways, the uncertainty estimation module can also be referred to as an uncertainty estimation network, uncertainty estimation unit, uncertainty estimator, etc. It mainly obtains a first value for representing the uncertainty of the unlabeled data based on the features of the data.

[0090] It should be understood that this first value can also be understood as the cross-entropy loss of the classification module fitted by the uncertainty estimation module for this unlabeled data.

[0091] In some implementable ways, the uncertainty estimation module can be a deep learning network, specifically a convolutional neural network, but is not limited to this.

[0092] In some implementable ways, the uncertainty estimation module can include: a structure similar to a regression model and an output layer.

[0093] In some implementable ways, a structure similar to a regression model can adopt a Multilayer Perceptron (MLP), etc., but is not limited thereto.

[0094] In some implementable ways, the output layer can adopt a softplus activation function.

[0095] In some implementable ways, softplus activation is a continuous non-linear activation function, and scalar uncertainty is output through the softplus activation function. The softplus activation function can ensure that the first value of the output is a non-negative value, and its range is [0, +∞), thus meeting the requirements of practical problems.

[0096] In some implementable ways, the uncertainty estimation module and the classification module can be two independent modules, or the uncertainty estimation module and the classification module can share some structures, and the embodiments of the present application do not limit this.

[0097] Figure 6 For the input-output schematic diagram of the uncertainty estimation module provided by the embodiments of the present application, as Figure 6 shown, the input of this uncertainty estimation module can be the features of unlabeled data, and the output is the first value used to represent the uncertainty of unlabeled data.

[0098] S220: Based on the pseudo-label and the predicted label of the unlabeled data, obtain the first loss of the classification module; and based on the first value and the first loss, obtain the second loss of the uncertainty estimation module; determine the weight corresponding to the first loss based on the first value;

[0099] Among them, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0100] In some implementable ways, the unlabeled data can be input into a pre-trained model to obtain the pseudo-label of the unlabeled data.

[0101] In some implementable ways, this pre-trained model can be a pre-trained target model obtained by training a target model with a training sample composed of labeled data and the true label corresponding to the labeled data.

[0102] In some other implementable ways, this pre-training module can be an open-source pre-trained model, and this open-source pre-trained model can be a pre-trained classification model.

[0103] It should be understood that the first loss reflects the difference between the pseudo-label and the predicted label of the unlabeled data, which can measure the uncertainty of the classification module's prediction for the unlabeled data. Among them, the greater the difference between the pseudo-label and the predicted label of the unlabeled data, the greater the uncertainty of the classification module's prediction for the unlabeled data; the smaller the difference between the pseudo-label and the predicted label of the unlabeled data, the smaller the uncertainty of the classification module's prediction for the unlabeled data. Based on this, the embodiments of the present application consider training the uncertainty estimation module by using the second loss calculated from the first value and the first loss. In this case, the first loss can be regarded as the loss supervision information of the uncertainty estimation module, enabling the uncertainty estimation module to better predict the uncertainty of the data.

[0104] In some implementable ways, the first loss can be any one of the following, but not limited to: Cross Entropy Loss, Log Loss.

[0105] Taking the first loss being the cross entropy loss as an example, it can be calculated by the following formula (1):

[0106] u(x) = -y × log(p(y)) (1)

[0107] Where x represents the unlabeled data, u(x) represents the first loss, y represents the pseudo-label of the unlabeled data, p(y) represents the predicted label of the unlabeled data, the value of y is 0 or 1, and the value range of p(y) is [0, 1].

[0108] In some implementable ways, the second loss can be any one of the following, but not limited to: Mean Squared Error (MSE), Mean Absolute Error (MAE).

[0109] It should be understood that if the second loss is MSE, then the training device can calculate the second loss in the following way: calculate the square of the difference between the first value and the first loss to obtain the second loss.

[0110] Taking the second loss being MSE as an example, it can be calculated by the following formula (2):

[0111]

[0112] Where L2 represents the second loss, represents the first value, and u(x) represents the first loss.

[0113] In some implementable ways, the training device can convert the first value into the weight corresponding to the first loss through an exponential function.

[0114] In some implementable ways, the base of the exponential function can be greater than 1. For example, it can take the value of e.

[0115] In other implementable ways, the base of the exponential function can be a number greater than 0 and less than 1. For example, it can take the value of 0.5.

[0116] In some implementable ways, when the base of the indicator function is greater than 1, the training device can determine the opposite of the first value; use the opposite as the exponent of the base to obtain the weight corresponding to the first loss.

[0117] Taking the base as e as an example, the weight corresponding to the first loss can be calculated by formula (3):

[0118]

[0119] As can be seen from formula (3), if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0120] In some implementable ways, when the base of the indicator function is a number greater than 0 and less than 1, the training device can use the first value as the exponent of the base to obtain the weight corresponding to the first loss.

[0121] Taking the base as 0.5 as an example, the weight corresponding to the first loss can be calculated by formula (4):

[0122]

[0123] As can be seen from formula (4), if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0124] S230: Train the target model based on the first loss, the weight corresponding to the first loss, and the second loss.

[0125] In some implementable ways, the training device can obtain the third loss of the classification module based on the first loss and the weight corresponding to the first loss; train the target model based on the second loss and the third loss.

[0126] In some implementable ways, the training device can calculate the product of the first loss and the weight corresponding to the first loss to obtain the third loss of the classification module.

[0127] For example, the third loss can be calculated by formula (5):

[0128]

[0129] Among them, L3 represents the third loss, Let \(w_1\) denote the weight of the first loss, and \(-y\times\log(p(y))\) denote the first loss. For the explanations of other parameters in formula (5), please refer to the above text, and they will not be elaborated in the embodiments of this application.

[0130] For another example, the third loss can be calculated by formula (6):

[0131]

[0132] where \(L_3\) represents the third loss, Let \(w_1\) denote the weight of the first loss, and \(-y\times\log(p(y))\) denote the first loss. For the explanations of other parameters in formula (6), please refer to the above text, and they will not be elaborated in the embodiments of this application.

[0133] In some implementable ways, the training device can calculate the sum of the second loss and the third loss to obtain the fourth loss; and train the target model based on the fourth loss.

[0134] For example, the fourth loss can be calculated by formula (7):

[0135]

[0136] where \(L_4\) represents the fourth loss, \(w_2\) represents the second loss, \(w_3\) represents the third loss. For the explanations of other parameters in formula (7), please refer to the above text, and they will not be elaborated in the embodiments of this application.

[0137] For example, the fourth loss can be calculated by formula (8):

[0138]

[0139] where \(L_4\) represents the fourth loss, \(w_2\) represents the second loss, \(w_3\) represents the third loss. For the explanations of other parameters in formula (8), please refer to the above text, and they will not be elaborated in the embodiments of this application.

[0140] In some other implementable ways, the training device can calculate the product of the second loss and the first weight to obtain the first product, and calculate the product of the third loss and the second weight to obtain the second product. Further, calculate the sum of the first product and the second product to obtain the fourth loss; and train the target model based on the fourth loss.

[0141] In some implementable ways, the value ranges of the first weight and the second weight are both \([0, 1]\), and the sum of the first weight and the second weight is 1.

[0142] It should be understood that the first weight and the second weight are respectively used to represent the importance of the second loss and the third loss. During the actual training process, if the first weight is greater than the second weight, it means that the training device focuses on training the uncertainty estimation module; if the first weight is less than the second weight, it means that the training device focuses on training the classification module; if the first weight is equal to the second weight, it means that when the training device trains the target model, it considers the classification module and the uncertainty estimation module to be equally important.

[0143] By setting the first weight and the second weight, the model training can be made more targeted and flexible.

[0144] For example, the fourth loss can be calculated by formula (9):

[0145]

[0146] where L4 represents the fourth loss, ω1 represents the first weight, ω2 represents the second weight, represents the second loss, represents the third loss. For the explanations of other parameters in formula (9), reference can be made to the above text, and the embodiments of this application will not elaborate on this.

[0147] For example, the fourth loss can be calculated by formula (10):

[0148]

[0149] where L4 represents the fourth loss, ω1 represents the first weight, ω2 represents the second weight, represents the second loss, represents the third loss. For the explanations of other parameters in formula (10), reference can be made to the above text, and the embodiments of this application will not elaborate on this.

[0150] It should be understood that during the model training process, the number of training samples composed of unlabeled data and their corresponding pseudo-labels is huge. N training samples can form a training set, where N is an integer greater than 1. Based on this, the training device can extract a part of the training data from the training set. After obtaining the fourth loss corresponding to each training data, the fourth losses of these training data can be summed to obtain the total loss.

[0151] Furthermore, the training device can use the backpropagation algorithm to propagate the total loss layer by layer from the output layer of the target model to the input layer, and calculate the gradient of each layer. According to the calculated gradient, an optimization algorithm such as the gradient descent method is used to update the parameters of the model.

[0152] Among them, the training device can perform multiple iterative training processes. In each iteration, a new batch of training data (usually a part randomly selected from the entire training set) is used to obtain a more comprehensive training result. When the model evaluation parameters reach the standard or the number of training times reaches the preset number, the training device can stop training.

[0153] It should be understood that the backpropagation algorithm is a learning algorithm suitable for multi-layer neural networks and is based on the gradient descent method. It inputs the training set data into the input layer of the network, passes through the hidden layer, and finally reaches the output layer and outputs the result, and then adjusts the parameters backward according to the output error. Specifically, it calculates the error between the result of the output layer and the actual result, and propagates this error backward from the output layer to the hidden layer until it reaches the input layer. During the backpropagation process, various parameter values are adjusted according to the error, and the above process is continuously iterated until convergence. The basis for the application of the backpropagation algorithm is that its information processing ability comes from the multiple compositions of simple non-linear functions, so it has a strong function reproduction ability.

[0154] It should be understood that the gradient descent method is an optimization algorithm used to minimize the loss function. It iteratively adjusts the parameters to make the value of the function closer and closer to the minimum value.

[0155] Among them, the gradient descent method has two main forms: batch gradient descent and stochastic gradient descent. Batch gradient descent uses the entire training set to calculate the gradient each time, while stochastic gradient descent uses only one training sample (or a small part of the samples) to calculate the gradient each time.

[0156] In some implementable ways, in the model training method provided in the embodiments of the present application, the training device can adjust the parameters of the classification module and the uncertainty estimation module, without adjusting the parameters of the feature extraction module, or the training device can adjust the parameters of the classification module, the uncertainty estimation module, and the feature extraction module.

[0157] It should be understood that the model training method provided in the embodiments of the present application can also be understood as the co-training of the classification module and the uncertainty estimation module. The following elaborates on this co-training process through an example:

[0158] Figure 7 For the schematic diagram of the co-training process provided in the embodiments of the present application, as Figure 7As shown, the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data; based on the pseudo-labels and predicted labels of the unlabeled data, a first loss of the classification module is obtained; and based on the first value and the first loss, a second loss of the uncertainty estimation module is obtained; based on the first value, the weight corresponding to the first loss is determined; based on the first loss, the weight corresponding to the first loss, and the second loss, the target model is trained.

[0159] An embodiment of the present application provides a model training method, including: a training device inputs unlabeled data into a target model, so that the unlabeled data is input into a feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into a classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into an uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data; based on the pseudo-labels and predicted labels of the unlabeled data, a first loss of the classification module is obtained; and based on the first value and the first loss, a second loss of the uncertainty estimation module is obtained; based on the first value, the weight corresponding to the first loss is determined; based on the first loss, the weight corresponding to the first loss, and the second loss, the target model is trained. Among them, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger. In other words, for unlabeled data with higher uncertainty, the loss weight corresponding to this unlabeled data can be set smaller, and for unlabeled data with lower uncertainty, the loss weight corresponding to this unlabeled data can be set larger, so that when training the target model, more reliance is placed on the unlabeled data with higher certainty, thereby improving the model training accuracy.

[0160] It should be understood that the unlabeled data in the real world is often disturbed by various noises. In the embodiment of the present application, there is no need to distinguish the noises or perform denoising processing. As long as the classification module and the uncertainty estimation module are co-trained, the uncertainty estimation module can accurately predict the uncertainty of the unlabeled data. For unlabeled data with different uncertainties, different weights can be set for this unlabeled data when calculating the loss of the classification module. Specifically, for unlabeled data with higher uncertainty, the corresponding weight is smaller, and for unlabeled data with lower uncertainty, the corresponding weight is larger, so that when training the target model, more reliance is placed on the unlabeled data with higher certainty. Therefore, from this perspective, the model training method provided by the embodiment of the present application can be applied to unlabeled data with any type of noise.

[0161] In addition, through the collaborative training of the classification module and the uncertainty estimation module, the classification module can better utilize the unlabeled data according to the uncertainty provided by the uncertainty estimation module, that is, it focuses on using the unlabeled data with less uncertainty, while the uncertainty estimation module can adjust itself according to the first loss of the classification module, realizing the mutual promotion between the classification module and the uncertainty estimation module, so that the two can jointly improve the prediction performance.

[0162] Figure 8 FIG. 4 is a flowchart of another model training method provided by an embodiment of the present application. The method can be executed by a training device, for example, it can be Figure 1 the training device 103 in Figure 8 As shown in FIG. 4, the model training method may include:

[0163] S810: Input the labeled data into the target model, so that the labeled data is input into the feature extraction module to obtain the features of the labeled data; and the features of the labeled data are input into the classification module to obtain the predicted labels of the labeled data; the features of the labeled data are input into the uncertainty estimation module to obtain a second value for representing the uncertainty of the labeled data;

[0164] In some implementable manners, the labeled data in the embodiments of the present application may be images, texts, voices, etc., and the embodiments of the present application do not limit this.

[0165] For example, the labeled data in the embodiments of the present application may be medical images, such as CT and MRI.

[0166] It should be understood that the explanation of the target model can be referred to Figure 2 the corresponding embodiment, and the embodiments of the present application will not elaborate on this.

[0167] It should be understood that the second value can also be understood as the cross-entropy loss of the classification module fitted by the uncertainty estimation module for the labeled data.

[0168] S820: Based on the true label and the predicted label of the labeled data, obtain the fifth loss of the classification module; and based on the second value and the fifth loss, obtain the sixth loss of the uncertainty estimation module; determine the weight corresponding to the fifth loss based on the second value;

[0169] Wherein, if the second value is larger, the weight corresponding to the fifth loss is smaller; if the second value is smaller, the weight corresponding to the fifth loss is larger.

[0170] It should be understood that the fifth loss reflects the difference between the true label and the predicted label of the labeled data, which can measure the uncertainty of the prediction of the classification module for the labeled data. Among them, the greater the difference between the true label and the predicted label of the labeled data, the greater the uncertainty of the prediction of the classification module for the labeled data, and the smaller the difference between the pseudo-label and the predicted label of the labeled data, the smaller the uncertainty of the prediction of the classification module for the labeled data. Based on this, the embodiments of the present application consider training the uncertainty estimation module through the sixth loss calculated from the second value and the fifth loss. In this case, the fifth loss can be regarded as the loss supervision information of the uncertainty estimation module, enabling the uncertainty estimation module to better predict the uncertainty of data (including unlabeled data and labeled data).

[0171] In some implementable ways, the fifth loss can be any of the following, but not limited to: Cross Entropy Loss, Log Loss.

[0172] It should be understood that for the calculation formula of the fifth loss, reference can be made to the calculation formula of the first loss, and the embodiments of the present application will not elaborate on this.

[0173] In some implementable ways, the sixth loss can be any of the following, but not limited to: MSE, MAE.

[0174] It should be understood that for the calculation formula of the sixth loss, reference can be made to the calculation formula of the second loss, and the embodiments of the present application will not elaborate on this.

[0175] In some implementable ways, the training device can convert the second value into the weight corresponding to the fifth loss through an exponential function.

[0176] In some implementable ways, the base of the exponential function can be greater than 1. For example, it can take the value of e.

[0177] In some other implementable ways, the base of the exponential function can be a number greater than 0 and less than 1. For example, it can take the value of 0.5.

[0178] In some implementable ways, when the base of the exponential function is greater than 1, the training device can determine the opposite number of the second value; use the opposite number as the exponent of the base to obtain the weight corresponding to the fifth loss.

[0179] In some implementable ways, when the base of the exponential function is a number greater than 0 and less than 1, the training device can use the second value as the exponent of the base to obtain the weight corresponding to the fifth loss.

[0180] It should be understood that the calculation formula for the weight corresponding to the fifth loss can refer to the calculation formula for the weight corresponding to the first loss, and the embodiments of the present application will not elaborate on this.

[0181] S830: Train the target model based on the fifth loss, the weight corresponding to the fifth loss, and the sixth loss;

[0182] In some implementable ways, the training device can obtain the seventh loss of the classification module based on the fifth loss and the weight corresponding to the fifth loss; and train the target model based on the sixth loss and the seventh loss.

[0183] In some implementable ways, the training device can calculate the product of the fifth loss and the weight corresponding to the fifth loss to obtain the seventh loss of the classification module.

[0184] It should be understood that the calculation formula for the seventh loss can refer to the calculation formula for the third loss, and the embodiments of the present application will not elaborate on this.

[0185] In some implementable ways, the training device can calculate the sum of the sixth loss and the seventh loss to obtain the eighth loss; and train the target model based on the eighth loss.

[0186] In some other implementable ways, the training device can calculate the product of the sixth loss and the third weight to obtain the third product, and calculate the product of the seventh loss and the fourth weight to obtain the fourth product. Further, calculate the sum of the third product and the fourth product to obtain the eighth loss; and train the target model based on the eighth loss.

[0187] In some implementable ways, the value ranges of the third weight and the fourth weight are both [0, 1], and the sum of the third weight and the fourth weight is 1.

[0188] It should be understood that the third weight and the fourth weight are respectively used to represent the importance of the sixth loss and the seventh loss. In the actual training process, if the third weight is greater than the fourth weight, it means that the training device focuses on training the uncertainty estimation module; if the third weight is less than the fourth weight, it means that the training device focuses on training the classification module; if the third weight is equal to the fourth weight, it means that the training device considers the classification module and the uncertainty estimation module to be equally important when training the target model.

[0189] By setting the third weight and the fourth weight, the model training can be made more targeted and flexible.

[0190] It should be understood that the calculation formula for the eighth loss can refer to the calculation formula for the fourth loss, and the embodiments of the present application will not elaborate on this.

[0191] It should be understood that during the model training process, the number of training samples composed of labeled data and their corresponding true labels is huge. M training samples can form a training set, where M is an integer greater than 1. Based on this, the training device can extract a part of the training data from the training set. After obtaining the eighth loss corresponding to each training data, the training device can sum up the eighth losses of these training data to obtain the total loss.

[0192] Furthermore, the training device can use the backpropagation algorithm to propagate the total loss layer by layer from the output layer to the input layer of the target model, and calculate the gradient of each layer. According to the calculated gradient, the training device can use an optimization algorithm, such as the gradient descent method, to update the parameters of the model.

[0193] Among them, the training device can perform multiple iterative training processes. Each iteration will use a new batch of training data (usually a part randomly selected from the entire training set) to obtain more comprehensive training results. When the model evaluation parameters meet the standards or the number of training times reaches the preset number, the training device can stop training.

[0194] In some implementable ways, in the model training method provided in the embodiments of the present application, the training device can adjust the parameters of the classification module and the uncertainty estimation module, without adjusting the parameters of the feature extraction module, or the training device can adjust the parameters of the classification module, the uncertainty estimation module, and the feature extraction module.

[0195] It should be understood that S810 - S830 can be understood as the pre-training process of the target model. After the target model is pre-trained, a pre-trained target model can be obtained. Therefore, S840 - S860 can be understood as the fine-tuning process of the pre-trained target model, and the target model in S840 - S860 can also be considered as the pre-trained target model.

[0196] S840: Input the unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain the first value representing the uncertainty of the unlabeled data;

[0197] S850: Based on the pseudo-labels and predicted labels of the unlabeled data, obtain the first loss of the classification module; and based on the first value and the first loss, obtain the second loss of the uncertainty estimation module; determine the weight corresponding to the first loss based on the first value;

[0198] S860: Based on the first loss, the weight corresponding to the first loss, and the second loss, train the target model.

[0199] It should be understood that for the explanations of S840 - S860, reference can be made to the explanations of S210 - S230, and the embodiments of the present application will not elaborate on this again.

[0200] In the embodiments of the present application, the training device can first pre - train the target model based on the training samples composed of labeled data and their corresponding true labels, and then fine - tune the target model through the training samples composed of labeled data and their corresponding pseudo - labels. Since the labeled data usually has relatively higher certainty or accuracy than the unlabeled data, therefore, pre - training the target model through the training samples composed of labeled data and their corresponding true labels can ensure that the target model has a certain accuracy. On this basis, further model fine - tuning can improve the model training efficiency.

[0201] Figure 9 It is a flowchart of a classification method provided by the embodiments of the present application. As Figure 9 shown, this method can be executed by an execution device, for example, it can be Figure 1 the execution device 104 in Figure 9 shown. This classification method can include:

[0202] S910: Obtain the data to be classified;

[0203] In some implementable ways, the data to be classified in the embodiments of the present application can be images, texts, voices, etc., and the embodiments of the present application do not limit this.

[0204] For example, the data to be classified in the embodiments of the present application can be medical images, such as CT and MRI.

[0205] S920: Input the data to be classified into the target model trained by the above - mentioned model training method, and obtain the predicted label of the data to be classified and a third value representing the uncertainty of the data to be classified.

[0206] It should be understood that the predicted label of the data to be classified represents the probability that the data to be classified belongs to any category. Among them, the predicted label of the data to be classified is the output result of the classification module in the target model, and the third value is the output result of the uncertainty estimation module in the target model.

[0207] It should be understood that for the explanation of the target model, reference can be made to Figure 2 the explanation of the corresponding embodiment, and the embodiments of the present application will not elaborate on this again.

[0208] In some implementable ways, in this classification method, the execution device can also delete or mask the uncertainty estimation module of the target model, so that the target model includes: a feature extraction module and a classification module. In this case, the execution device inputs the data to be classified into the target model trained by the above model training method, and obtains the predicted label of the data to be classified, without outputting the third value.

[0209] An embodiment of the present application provides a classification method, including: obtaining data to be classified, inputting the data to be classified into a target model trained by the above model training method, and obtaining a predicted label of the data to be classified and a third value used to represent the uncertainty of the data to be classified. Since a relatively accurate target model can be obtained through the above model training method, the classification accuracy can be improved.

[0210] The embodiments of the present application will be described in detail below in combination with the medical field scenario:

[0211] Figure 10 It is a flowchart of a model training method provided by an embodiment of the present application. Among them, this method can be executed by a training device, for example, it can be Figure 1 the training device 103 therein. This model training method is used to train a target model, as Figure 3 shown. This target model includes: a feature extraction module, a classification module, and an uncertainty estimation module, as Figure 10 shown. This model training method can include:

[0212] S1010: Input an unlabeled medical image into the target model, so that the unlabeled medical image is input into the feature extraction module to obtain the features of the unlabeled medical image; and the features of the unlabeled medical image are input into the classification module to obtain the predicted label of the unlabeled medical image; the features of the unlabeled medical image are input into the uncertainty estimation module to obtain a first value used to represent the uncertainty of the unlabeled medical image;

[0213] It should be understood that the explanation of the target model can refer to Figure 2 the corresponding method embodiment, and this embodiment will not be elaborated here.

[0214] In some implementable ways, the unlabeled medical image can be CT, MRI, etc., but not limited thereto.

[0215] In some implementable ways, before the feature extraction module extracts the features of the unlabeled medical image, the feature extraction module can first preprocess the unlabeled medical image, such as image enhancement, noise reduction, normalization, etc., to ensure the accuracy and stability of the subsequent extracted features.

[0216] In some implementable ways, the features of the unlabeled medical image include at least one of the following, but are not limited to: the color feature, texture feature, shape feature, and spatial relationship feature of the unlabeled image.

[0217] Among them, the shape feature may include a contour feature and a region feature.

[0218] It should be understood that this first value can also be understood as the cross-entropy loss of the classification module fitted by the uncertainty estimation module with respect to the unlabeled medical image.

[0219] In some implementable ways, the uncertainty estimation module and the classification module can be two independent modules, or the uncertainty estimation module and the classification module can share some structures, and the embodiments of the present application do not limit this.

[0220] S1020: Obtain the first loss of the classification module based on the pseudo-label and the prediction label of the unlabeled medical image; and obtain the second loss of the uncertainty estimation module based on the first value and the first loss; determine the weight corresponding to the first loss based on the first value;

[0221] Among them, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0222] In some implementable ways, the unlabeled medical image can be input into a pre-trained model to obtain the pseudo-label of the unlabeled medical image.

[0223] In some implementable ways, the pre-trained model can be a pre-trained target model obtained by training a target model with a training sample composed of a labeled medical image and the true label corresponding to the labeled medical image.

[0224] In some other implementable ways, the pre-training module can be a currently open-source pre-trained model, and the open-source pre-trained model can be a pre-trained classification model.

[0225] It should be understood that the first loss reflects the difference between the pseudo-label and the prediction label of the unlabeled medical image, and it can measure the uncertainty of the prediction of the classification module for the unlabeled medical image. Among them, the greater the difference between the pseudo-label and the prediction label of the unlabeled medical image, the greater the uncertainty of the prediction of the classification module for the unlabeled medical image, and the smaller the difference between the pseudo-label and the prediction label of the unlabeled medical image, the smaller the uncertainty of the prediction of the classification module for the unlabeled medical image. Based on this, the embodiments of the present application consider training the uncertainty estimation module through the second loss calculated from the first value and the first loss. In this case, the first loss can be regarded as the loss supervision information of the uncertainty estimation module, enabling the uncertainty estimation module to better predict the uncertainty of the data.

[0226] It should be understood that the calculation methods for the first loss, the weight corresponding to the first loss, and the second loss can be referred to Figure 2 the corresponding method embodiments, and will not be elaborated herein.

[0227] S1030: Train the target model based on the first loss, the weight corresponding to the first loss, and the second loss.

[0228] In some implementable ways, the training device can obtain the third loss of the classification module based on the first loss and the weight corresponding to the first loss; and train the target model based on the second loss and the third loss.

[0229] In some implementable ways, the training device can calculate the product of the first loss and the weight corresponding to the first loss to obtain the third loss of the classification module.

[0230] It should be understood that the calculation method for the third loss can be referred to Figure 2 the corresponding method embodiments, and will not be elaborated herein.

[0231] In some implementable ways, the training device can calculate the sum of the second loss and the third loss to obtain the fourth loss; and train the target model based on the fourth loss.

[0232] It should be understood that the calculation method for the fourth loss and the training method for the target model based on the fourth loss can be referred to Figure 2 the corresponding method embodiments, and will not be elaborated herein.

[0233] It should be understood that the model training method provided by the embodiments of the present application can also be understood as co-training of the classification module and the uncertainty estimation module.

[0234] The embodiments of the present application provide a model training method. Among them, for unlabeled medical images with high uncertainty, the loss weight corresponding to the unlabeled medical image can be set smaller; for unlabeled medical images with low uncertainty, the loss weight corresponding to the unlabeled medical image can be set larger, so that when training the target model, more reliance is placed on unlabeled medical images with higher certainty, thereby improving the model training accuracy.

[0235] Figure 11 is a flowchart of a classification method provided by the embodiments of the present application. As Figure 11 shown, this method can be executed by an execution device, for example, it can be Figure 1 the execution device 104 in Figure 11 shown. This classification method can include:

[0236] S1110: Obtain the medical image to be classified;

[0237] In some implementable ways, the medical image to be classified can be a CT or MRI to be classified, etc., but not limited thereto.

[0238] S1120: Input the medical image to be classified into the target model trained by the above model training method, and obtain the predicted label of the medical image to be classified and a third value used to represent the uncertainty of the medical image to be classified.

[0239] It should be understood that the predicted label of the medical image to be classified represents the probability that the medical image to be classified belongs to any category. Among them, the predicted label of the medical image to be classified is the output result of the classification module in the target model, and the third value is the output result of the uncertainty estimation module in the target model.

[0240] In some implementable ways, the classification method provided by the embodiments of the present application can be a binary classification method. For example, based on the medical image to be classified, it is determined whether the patient has a tumor. One classification result is: having a tumor, and the other classification result is: not having a tumor.

[0241] In other implementable ways, the classification method provided by the embodiments of the present application can be an N-classification method, where N is an integer greater than 2. For example, based on the medical image to be classified, it is determined whether the patient has a tumor, and if the patient has a tumor, the specific tumor classification results are, for example, liver cysts, lipomas, breast fibroids, etc.

[0242] The embodiments of the present application provide a classification method, including: obtaining the medical image to be classified, inputting the medical image to be classified into the target model trained by the above model training method, and obtaining the predicted label of the medical image to be classified and a third value used to represent the uncertainty of the medical image to be classified. Since a relatively accurate target model can be obtained through the above model training method, the accuracy and reliability of medical image analysis can be improved through this classification method, and it is convenient to assist doctors in giving relatively accurate diagnostic results.

[0243] The preferred embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, in the various specific technical features described in the above specific embodiments, they can be combined in any suitable way without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. Also, for example, any combination can be made between various different embodiments of the present application as long as it does not violate the idea of the present application, and it should also be regarded as the content disclosed by the present application.

[0244] It should also be understood that in various method embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0245] The method provided by the embodiments of the present application has been described above. Next, the model training device provided by the embodiments of the present application will be described.

[0246] Figure 12 It is a schematic diagram of a model training device 1200 provided by an embodiment of the present application. Among them, the target model includes: a feature extraction module, a classification module, and an uncertainty estimation module. As Figure 12 shown, the device 1200 includes:

[0247] An input module 1210, configured to input unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data;

[0248] A calculation module 1220, configured to obtain a first loss of the classification module based on the pseudo-label and the predicted label of the unlabeled data; and obtain a second loss of the uncertainty estimation module based on the first value and the first loss; determine a weight corresponding to the first loss based on the first value;

[0249] A training module 1230, configured to train the target model based on the first loss, the weight corresponding to the first loss, and the second loss;

[0250] Among them, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

[0251] In some implementable manners, the calculation module 1220 is specifically configured to: convert the first value into a weight corresponding to the first loss through an exponential function.

[0252] In some implementable manners, the calculation module 1220 is specifically configured to: determine the opposite number of the first value; use the opposite number as the exponent of e to obtain a weight corresponding to the first loss.

[0253] In some implementable manners, the calculation module 1220 is specifically configured to: calculate the square of the difference between the first value and the first loss to obtain the second loss.

[0254] In some implementable manners, the training module 1230 is specifically configured to: obtain a third loss of the classification module based on the first loss and the weight corresponding to the first loss; and train the target model based on the second loss and the third loss.

[0255] In some implementable manners, the training module 1230 is specifically configured to: calculate the product of the first loss and the weight corresponding to the first loss to obtain a third loss of the classification module.

[0256] In some implementable manners, the training module 1230 is specifically configured to: calculate the sum of the second loss and the third loss to obtain a fourth loss; and train the target model based on the fourth loss.

[0257] In some implementable manners, the input module 1210 is further configured to: before inputting the unlabeled data into the target model, input the unlabeled data into a pre-trained model to obtain pseudo-labels of the unlabeled data.

[0258] In some implementable manners, before the input module 1210 inputs the unlabeled data into the target model, the input module 1210 is further configured to: input the labeled data into the target model so that the labeled data is input into the feature extraction module to obtain features of the labeled data; and the features of the labeled data are input into the classification module to obtain predicted labels of the labeled data; the features of the labeled data are input into the uncertainty estimation module to obtain a second value representing the uncertainty of the labeled data; the calculation module 1220 is further configured to: obtain a fifth loss of the classification module based on the true label and the predicted label of the labeled data; and obtain a sixth loss of the uncertainty estimation module based on the second value and the fifth loss; determine the weight corresponding to the fifth loss based on the second value; the training module 1230 is further configured to: train the target model based on the fifth loss, the weight corresponding to the fifth loss, and the sixth loss; wherein, if the second value is larger, the weight corresponding to the fifth loss is smaller; if the second value is smaller, the weight corresponding to the fifth loss is larger.

[0259] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 12 the illustrated apparatus 1200 can execute Figure 2 、 Figure 8 and Figure 10 the corresponding method embodiments, and the foregoing and other operations and / or functions of each module in the apparatus 1200 are respectively for implementing Figure 2 、 Figure 8 and Figure 10 the corresponding processes in each method in, for the sake of brevity, they are not elaborated here.

[0260] In the above, the apparatus 1200 of the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0261] Figure 13 FIG. is a schematic diagram of a classification apparatus 1300 provided by an embodiment of the present application. As Figure 13 shown, the apparatus 1300 includes:

[0262] An acquisition module 1310, configured to acquire data to be classified;

[0263] An input module 1320, configured to input the data to be classified into a target model trained by the above model training method, and obtain a predicted label of the data to be classified and a third value indicating the uncertainty of the data to be classified.

[0264] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 13 the apparatus 1300 shown can execute Figure 9 and Figure 11 the corresponding method embodiments, and the foregoing and other operations and / or functions of each module in the apparatus 1300 are respectively for implementing Figure 9 and Figure 11 the corresponding processes in each method in, and for the sake of brevity, it will not be elaborated here.

[0265] Above, the device 1300 of the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that these functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0266] Figure 14 is a schematic block diagram of an electronic device provided by an embodiment of the present application.

[0267] As Figure 14 shown, the electronic device may include:

[0268] A memory 1410 and a processor 1420. The memory 1410 is used to store a computer program and transmit the program code to the processor 1420. In other words, the processor 1420 can call and run the computer program from the memory 1410 to implement the method in the embodiments of the present application.

[0269] For example, the processor 1420 can be used to execute the above method embodiments according to the instructions in the computer program.

[0270] In some embodiments of the present application, the processor 1420 may include, but is not limited to:

[0271] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0272] In some embodiments of the present application, the memory 1410 includes, but is not limited to:

[0273] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0274] In some embodiments of the present application, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 1410 and executed by the processor 1420 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0275] As Figure 14 shown, the electronic device may further include:

[0276] A transceiver 1430, which can be connected to the processor 1420 or the memory 1410.

[0277] Among them, the processor 1420 can control the transceiver 1430 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 1430 can include a transmitter and a receiver. The transceiver 1430 may further include an antenna, and the number of antennas can be one or more.

[0278] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0279] This application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, the embodiments of this application also provide a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0280] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0281] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0282] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0283] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0284] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, The target model includes: a feature extraction module, a classification module, and an uncertainty estimation module, and the method includes: Inputting the unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted labels of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data; Based on the pseudo-labels and predicted labels of the unlabeled data, obtaining a first loss of the classification module; and based on the first value and the first loss, obtaining a second loss of the uncertainty estimation module; determining the weight corresponding to the first loss based on the first value; Training the target model based on the first loss, the weight corresponding to the first loss, and the second loss; Wherein, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

2. The method according to claim 1, characterized in that The determining the weight corresponding to the first loss based on the first value includes: Converting the first value into the weight corresponding to the first loss through an exponential function.

3. The method according to claim 2, wherein The converting the first value into the weight corresponding to the first loss through an exponential function includes: Determining the opposite number of the first value; Taking the opposite number as the exponent of e to obtain the weight corresponding to the first loss.

4. The method according to any one of claims 1 to 3, characterized in that, The obtaining the second loss of the classification module based on the first value and the first loss includes: Calculating the square of the difference between the first value and the first loss to obtain the second loss.

5. The method according to any one of claims 1-3, characterized in that, The training the target model based on the first loss, the weight corresponding to the first loss, and the second loss includes: Based on the first loss and the weight corresponding to the first loss, obtaining a third loss of the classification module; Training the target model based on the second loss and the third loss.

6. The method according to claim 5, wherein The obtaining the third loss of the classification module based on the first loss and the weight corresponding to the first loss includes: Calculating the product of the first loss and the weight corresponding to the first loss to obtain the third loss of the classification module.

7. The method according to claim 5, characterized in that The training the target model based on the second loss and the third loss includes: Calculating the sum of the second loss and the third loss to obtain a fourth loss; Training the target model based on the fourth loss.

8. The method according to any one of claims 1-3, characterized in that, Before inputting the unlabeled data into the target model, it further includes: Inputting the unlabeled data into a pre-trained model to obtain the pseudo-labels of the unlabeled data.

9. The method according to any one of claims 1 to 3, characterized in that, Before inputting the unlabeled data into the target model, it further includes: Inputting the labeled data into the target model, so that the labeled data is input into the feature extraction module to obtain the features of the labeled data; and the features of the labeled data are input into the classification module to obtain the predicted labels of the labeled data; the features of the labeled data are input into the uncertainty estimation module to obtain a second value representing the uncertainty of the labeled data; Based on the true label and the predicted label of the labeled data, obtain the fifth loss of the classification module; and based on the second value and the fifth loss, obtain the sixth loss of the uncertainty estimation module; determine the weight corresponding to the fifth loss based on the second value; Train the target model based on the fifth loss, the weight corresponding to the fifth loss, and the sixth loss; wherein, if the second value is larger, the weight corresponding to the fifth loss is smaller; if the second value is smaller, the weight corresponding to the fifth loss is larger.

10. A classification method, characterized in that, It includes: Obtain the data to be classified; Input the data to be classified into the target model trained by the model training method according to any one of claims 1-9, and obtain the predicted label of the data to be classified and a third value representing the uncertainty of the data to be classified.

11. A model training device, characterized in that, The target model includes: a feature extraction module, a classification module, and an uncertainty estimation module. The model training device includes: An input module, configured to input unlabeled data into the target model, so that the unlabeled data is input into the feature extraction module to obtain the features of the unlabeled data; and the features of the unlabeled data are input into the classification module to obtain the predicted label of the unlabeled data; the features of the unlabeled data are input into the uncertainty estimation module to obtain a first value representing the uncertainty of the unlabeled data; A calculation module, configured to obtain the first loss of the classification module based on the pseudo-label and the predicted label of the unlabeled data; and based on the first value and the first loss, obtain the second loss of the uncertainty estimation module; determine the weight corresponding to the first loss based on the first value; A training module, configured to train the target model based on the first loss, the weight corresponding to the first loss, and the second loss; wherein, if the first value is larger, the weight corresponding to the first loss is smaller; if the first value is smaller, the weight corresponding to the first loss is larger.

12. A classification device, characterized in that, It includes: An acquisition module, configured to acquire the data to be classified; An input module, configured to input the data to be classified into the target model trained by the model training method according to any one of claims 1-9, and obtain the predicted label of the data to be classified and a third value representing the uncertainty of the data to be classified.

13. An electronic device, characterized in that, It includes: A processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 1 to 10.

15. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Solid state disk state information prediction model training method, prediction method and device

    CN120821647A