Prediction model training method, prediction method and device
By calculating regularization losses in multimodal learning and updating model parameters, the problem of unreliable confidence in multimodal learning is solved, and the reliability and robustness of the prediction model is improved, and applied to medical diagnosis and advertising delivery models.
Patent Information
- Application Number
- CN202410111214.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-29
AI Technical Summary
The existing multimodal learning methods have problems with unreliability and lack of robustness in confidence estimation, resulting in insufficient credibility of the prediction results.
By acquiring at least two modal data of the sample, the first and second confidence levels are calculated, and the regularization loss is determined when the second confidence level is less than the first confidence level, the loss constraint model parameters are updated to ensure that the predicted confidence does not decrease when the modal data is added.
It improves the confidence reliability and robustness of the prediction model, reduces the probability of misdiagnosis and missed diagnosis, and improves the accuracy of medical diagnosis and advertising delivery models.
Smart Images

Figure CN120387059A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method, a prediction method, and a device for training a prediction model. Background Art
[0002] With the development of technology, people can obtain data from a variety of sensors and data sources. For example, in medical diagnosis, computer tomography (CT) images, pathological images, and clinical indicator data can be utilized simultaneously; in Internet advertising, object click data, web browsing history data, and profile data, etc. can be utilized simultaneously. In order to make more full use of the complementary information of heterogeneous multi-source data, multi-modal learning has emerged. Multi-modal learning attempts to integrate all available multi-modal information by training a model to achieve downstream analysis tasks.
[0003] At present, great progress has been made in multi-modal learning, but the reliability of current multi-modal learning methods has mostly not been explored. In related technologies, for a classification task, a high-quality confidence estimator can be constructed to quantitatively describe the probability of a correct prediction. However, there are still problems in the confidence estimation of related solutions, such as the estimated prediction confidence being unreliable and lacking robustness. Summary of the Invention
[0004] The present application provides a method, a prediction method, and a device for training a prediction model, which are beneficial to improving the reliability and robustness of the prediction confidence of the prediction model.
[0005] In a first aspect, an embodiment of the present application provides a method for training a prediction model, including:
[0006] Obtain training data, where the training data includes at least two types of modal data of a sample and the true label corresponding to the sample;
[0007] Input a first set of modal data in the at least two types of modal data into the prediction model to obtain a first prediction result and a first confidence of the sample; and input a second set of modal data in the at least two types of modal data into the prediction model to obtain a second prediction result and a second confidence of the sample; where the first set of modal data is a non-empty proper subset of the second set of modal data;
[0008] When the second confidence is less than the first confidence, determine a regularization loss according to the first confidence and the second confidence;
[0009] Determine a target loss according to the regularization loss, the first prediction result, the second prediction result, and the true label;
[0010] Update the parameters of the prediction model according to the target loss to obtain the trained prediction model.
[0011] In a second aspect, an embodiment of the present application provides a prediction method, including:
[0012] Obtain at least one modality data of a sample to be predicted;
[0013] Input the at least one modality data into a prediction model to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the method described in the first aspect.
[0014] In a third aspect, a prediction model training apparatus is provided, including:
[0015] An acquisition unit for acquiring training data, where the training data includes at least two modality data of a sample and a true label corresponding to the sample;
[0016] A prediction unit for inputting a first modality data set among the at least two modality data into a prediction model to obtain a first prediction result and a first confidence level of the sample; and inputting a second modality data set among the at least two modality data into the prediction model to obtain a second prediction result and a second confidence level of the sample; wherein, the first modality data set is a non-empty proper subset of the second modality data set;
[0017] A determination unit for determining a regularization loss according to the first confidence level and the second confidence level when the second confidence level is less than the first confidence level;
[0018] The determination unit is further configured to determine a target loss according to the regularization loss, the first prediction result, the second prediction result, and the true label;
[0019] A parameter update unit for updating the parameters of the prediction model according to the target loss to obtain the trained prediction model.
[0020] In a fourth aspect, a prediction apparatus is provided, including:
[0021] An acquisition unit for acquiring at least one modality data of a sample to be predicted;
[0022] A prediction model for inputting the at least one modality data to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the method described in the first aspect.
[0023] Fifth aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or the method in the second aspect.
[0024] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on a computer cause the computer to execute the method in the first aspect or the method in the second aspect.
[0025] Fifth aspect, an embodiment of the present application provides a computer program product, including computer program instructions, which cause the computer to execute the method in the first aspect or the method in the second aspect.
[0026] Sixth aspect, an embodiment of the present application provides a computer program, which causes the computer to execute the method in the first aspect or the method in the second aspect.
[0027] In the above technical solution, by inputting the first modal data set of the sample into the prediction model to obtain the first confidence level, and inputting the second modal data set of the sample into the prediction model to obtain the second confidence level, where the second modal data set adds modal data relative to the first modal data set, and further determining the regularization loss when the second confidence level is less than the first confidence level, and updating the parameters of the model according to the regularization loss, to achieve the penalty for samples whose prediction confidence level does not increase but decreases when adding modal data, thereby improving the reliability and robustness of the prediction confidence level of the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of a system architecture according to an embodiment of the present application;
[0029] Figure 2 It is an optional schematic diagram of an application scenario according to an embodiment of the present application;
[0030] Figure 3 It is an optional schematic diagram of an application scenario according to an embodiment of the present application;
[0031] Figure 4 It is a schematic flowchart of a prediction model training method according to an embodiment of the present application;
[0032] Figure 5 It is a schematic diagram of a network architecture of a training model according to an embodiment of the present application;
[0033] Figure 6 It is a schematic diagram of a sample corresponding to three modal data according to an embodiment of the present application;
[0034] Figure 7 It isFigure 6 An optional schematic diagram of adding a penalty term to the sample in
[0035] Figure 8 A schematic flowchart of a prediction method according to an embodiment of the present application;
[0036] Figure 9 A schematic flowchart of a prediction model training device according to an embodiment of the present application;
[0037] Figure 10 A schematic block diagram of a prediction device according to an embodiment of the present application;
[0038] Figure 11 A schematic block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0040] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0041] In the description of the present application, unless otherwise specified, "at least one" means one or more, and "a plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0042] It should also be understood that the first, second, etc. descriptions that appear in the embodiments of the present application are only for schematic and distinguishing the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0043] It should also be understood that the specific features, structures, or characteristics related to the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.
[0044] In addition, the terms "comprising", "having", and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that comprises a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to those processes, methods, products, or devices.
[0045] The embodiments of this application are applied to the field of artificial intelligence technology.
[0046] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0048] The embodiments of the present application may relate to natural language processing (NLP) in the field of artificial intelligence technology. NLP is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0049] The embodiments of the present application may also relate to machine learning (ML) in the field of artificial intelligence technology. ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0050] In related technologies, multi-modal learning realizes downstream analysis tasks by training a model to integrate all available multi-modal information. At present, great progress has been made in multi-modal learning, but the reliability of current multi-modal learning methods has mostly not been explored. For example, for classification tasks, a high-quality confidence estimator can be constructed to quantitatively describe the probability of correct prediction. However, the related solutions still have problems such as unreliable estimated prediction confidence and lack of robustness in confidence estimation.
[0051] In view of this, the embodiments of the present application provide a prediction model training method, a prediction method, and a device, which can help improve the reliability and robustness of prediction confidence.
[0052] Specifically, training data can be obtained, which includes at least two modalities of data of the samples and the true labels corresponding to the samples; input the first modality data set among the at least two modalities of data into the prediction model to obtain the first prediction result and the first confidence level of the samples; input the second modality data set among the at least two modalities of data into the prediction model to obtain the second prediction result and the second confidence level of the samples; wherein, the first modality data set is a non-empty proper subset of the second modality data set; when the second confidence level is less than the first confidence level, determine the regularization loss according to the first confidence level and the second confidence level; determine the target loss according to the regularization loss, the first prediction result, the second prediction result and the true label; update the parameters of the prediction model according to the target loss to obtain the trained prediction model.
[0053] Therefore, in the embodiment of the present application, by inputting the first modality data set of the sample into the prediction model to obtain the first confidence level, and inputting the second modality data set of the sample into the prediction model to obtain the second confidence level, where the second modality data set has more modality data than the first modality data set, and further determining the regularization loss when the second confidence level is less than the first confidence level, and constraining the model to update the parameters according to the regularization loss, the penalty for the samples whose prediction confidence level does not increase but decreases when adding modality data is realized, thereby improving the reliability and robustness of the prediction confidence level of the prediction model.
[0054] The embodiment of the present application can be used in any prediction model based on multi-modal data. For example, in the medical field, with the development of medical imaging technology, doctors can obtain patient information from multiple modalities of data, such as CT images, Magnetic Resonance Imaging (MRI) images, pathological images, etc. Medical diagnosis models attempt to comprehensively use these different modalities of medical data for disease diagnosis and prognosis judgment. However, due to the over-reliance on certain modalities of data in current medical diagnosis models, the confidence level of the predicted diagnosis results is unreliable. The method for training the prediction model provided by the embodiment of the present application can be used for training a medical diagnosis model based on multi-modal medical data, improving the reliability of the prediction results of the medical diagnosis model. Exemplarily, by integrating an additional regularization loss term into an existing medical diagnosis model based on multi-modal data, since the regularization loss penalizes the samples whose prediction confidence level does not increase but decreases when adding modality data to train the model parameters, it is beneficial to improve the calibration degree of the model prediction results, making them more reliable, thereby reducing the probability of misdiagnosis and missed diagnosis. Further, by improving the reliability of the prediction results of the medical diagnosis model based on multi-modal medical data, the embodiment of the present application can help improve the diagnostic level of clinicians, help doctors customize more reasonable diagnosis and treatment plans according to the model prediction results, so that the medical diagnosis model can better assist in the diagnosis of various diseases.
[0055] For another example, in the field of Internet advertising, object click data, browsing history data, profile data, etc. can be utilized simultaneously. An advertising placement model can perform targeted advertising placement by integrating these object data of different modalities. The method for training a prediction model provided by the embodiments of the present application can be used for training an advertising placement model based on multi-modal data, improving the reliability of targeted advertising placement of the advertising placement model. Exemplarily, the embodiments of the present application integrate an additional regularization loss term into an existing advertising placement model based on multi-modal data. Since the regularization loss penalizes samples whose prediction confidence does not increase but decreases when adding modal data to train model parameters, it is beneficial to improve the calibration degree of the model prediction results, making them more reliable, thereby reducing the probability of incorrect advertising placement.
[0056] Figure 1 It is a schematic diagram of a system architecture related to the embodiments of the present application. As Figure 1 shown, the system architecture may include a user device 101, a data collection device 102, a training device 103, an execution device 104, a database 105, and a content library 106.
[0057] Among them, the data collection device 102 is used to read training data from the content library 106 and store the read training data in the database 105. The training data related to the embodiments of the present application includes at least two modal data of samples and the true labels of the samples, such as multi-modal medical data and corresponding disease types.
[0058] The training device 103 trains a machine learning model based on the training data maintained in the database 105. The model obtained by the training device 103 can output a prediction result. Optionally, the model can further be connected to other downstream models. The model obtained by the training device 103 can be applied to different systems or devices.
[0059] In addition, referring to Figure 1 , the execution device 104 is configured with an I / O interface 107 to interact with external devices. For example, it receives multi-modal data sent by the user device 101 through the I / O interface. The computing module 109 in the execution device 104 processes the input data using the trained machine learning model, outputs a prediction result, and sends the corresponding result to the user device 101 through the I / O interface.
[0060] Among them, the user device 101 may include a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted terminal, a mobile internet device (MID), or other terminal devices.
[0061] The execution device 104 may be a server. Exemplarily, the server may be a computing device such as a rack server, a blade server, a tower server, or a cabinet server. The server may be an independent server or a server cluster composed of multiple servers.
[0062] In this embodiment, the execution device 104 is connected to the user device 101 through a network. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, or a call network.
[0063] It should be noted that Figure 1 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. In some embodiments, the above data acquisition device 102, user device 101, training device 103, and execution device 104 may be the same device. The above database 105 may be distributed on one server or on multiple servers, and the above content library 106 may be distributed on one server or on multiple servers.
[0064] Figure 2 is an optional schematic diagram of the application scenario of the prediction model. As Figure 2 shown, it includes an input sample 210, a classifier 220, and an output category 230. Among them, after the sample 210 is input into the classifier 220, the classifier 220 can output the category 230 to which the sample belongs. This prediction model may be called a classification model.
[0065] See Figure 2 , the goal of the classification model is to learn a parameterized function: f(x,θ)→z, where θ is the parameter of the network and z is the output of the network. z may be called logits, and the logits vector, that is, the normalized probability distribution, is obtained after the logits are transformed by the softmax layer. Among them, the category label corresponding to the maximum probability in the probability distribution of the sample x is the predicted category label of the sample x This maximum probability is the confidence level of the category label .
[0066] Figure 2The prediction process corresponding to the shown prediction model is the class prediction process of a traditional classification network, and the process of determining the confidence of the class label therein is a traditional confidence estimation process. However, in a medical diagnosis model based on multimodal data, this confidence estimation process cannot be guaranteed to be applicable in all cases. For example, in non-class prediction tasks such as predicting the degree of a patient's lesion and survival time, the classification model will no longer be used.
[0067] Based on this, the embodiments of the present application provide a confidence estimation scheme, which uses a neural network model to fit the confidence estimation, so that the prediction method provided by the embodiments of the present application can be extended to other models other than the classification model, such as the prediction model corresponding to a generative task.
[0068] Figure 3 Fig. shows an optional schematic diagram of the application scenario of the prediction model provided by the embodiments of the present application. As Figure 3 shown, it includes an input sample 310, a confidence estimation network 320, a classifier 330, a confidence 340, and an output class 350. Among them, after the sample 310 is input into the classifier 330, the classifier 330 can output the class 350 to which the sample belongs. After the sample 310 is input into the confidence estimation network 320, the confidence estimation network 320 can output the confidence 340 corresponding to the class to which it belongs.
[0069] Relative to Figure 2 the network architecture in Figure 3 a confidence estimation network 320 is added in
[0070] to fit the confidence estimation of the output class 350, making the output class 350 more accurate. Exemplarily, the confidence estimation network 320 can be a neural network model, and the present application does not limit this. Figure 2 Figure 2 Figure 2 For details, the sample 310, the classifier 330, and the output class 350 are similar to the corresponding modules in
[0071] It can be understood that Figure 3 the classifier 330 in
[0072] can also be replaced by other types of prediction models, such as a generative model. Thus, the confidence estimation 320 can estimate the confidence of the prediction result output by the generative model, enabling models other than the classification model to also estimate the confidence of their output results, effectively expanding the usage scenario of the solution. Exemplarily, a disease diagnosis model based on a generative model can generate predictions of the degree of a disease lesion or the survival time of a patient, etc., and the present application does not limit this.The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0073] Figure 4 FIG. 400 is a schematic flowchart of a prediction model training method 400 according to an embodiment of the present application. This method 400 can be executed by any electronic device with data processing capabilities. For example, this electronic device can be implemented as a server or a terminal device. Another example is that this electronic device can be implemented as Figure 1 the training device 103 in
[0074] In some embodiments, the electronic device may include (such as deploy) a network architecture for training a model to execute the model training method 400. Figure 5 FIG. shows a schematic diagram of a network architecture for training a model, which can be used to execute method 400. As Figure 5 shown, this network architecture includes a feature extraction network 510, a prediction network 520, and a confidence estimation network 530. Among them, the feature extraction network 510 is used to extract feature data of the input modal data, the prediction network 520 is used to obtain the prediction result of the sample according to the input feature data; the confidence estimation network 530 is used to obtain the confidence of the prediction result according to the input feature data.
[0075] In some embodiments, Figure 5 the network architecture in
[0076] Optionally, the prediction network 520 may include a classification network or a generative network, and the embodiments of the present application do not make any limitations in this regard. Exemplarily, the classification network can be a classifier, and the generative network can be a regression network.
[0077] It should be understood that Figure 5 FIG. shows an example of a network architecture for model training. This example is only to help those skilled in the art understand and implement the embodiments of the present application, rather than limiting the scope of the embodiments of the present application. Those skilled in the art can make equivalent transformations or modifications according to the examples given here, and such transformations or modifications should still fall within the protection scope of the embodiments of the present application.
[0078] As Figure 4 shown, the prediction model training method 400 includes steps 410 to 450.
[0079] 410. Obtain training data, where the training data includes at least two modalities of data of the samples and the true labels corresponding to the samples.
[0080] Exemplarily, the training data set can be defined as where represents the m-th modality of the i-th sample, and y i ∈{1,…,K} is the class label of the i-th sample. Taking ophthalmic examination data as an example, represents the CFP modality of the first patient, represents the OCT modality of the first patient.
[0081] In some embodiments, to distinguish one modality or a group of modalities, x m and x (S) are respectively used to represent the m-th modality and multiple modalities, where S is the set of modality indices. For example, if S = {1, 2}, then x (S) represents the feature set composed of x 1 and x 2 . Exemplarily, x (M) = {x 1 ,…,x M} represents the complete M modalities of data.
[0082] Exemplarily, when the prediction model is a medical diagnosis model, the at least two modalities of data of the above samples include at least two multi-modal medical examination data of the patient, and the corresponding true label can be the diagnosis result corresponding to the sample. Exemplarily, the diagnosis result can be a diagnosis class label, such as the type of disease; or the diagnosis result can be a generative prediction result, such as the prediction of the degree of lesion, the survival time of the patient, etc., and the present application does not make any limitations in this regard.
[0083] 420. Input the first modality data set in the at least two modalities of data into the prediction model to obtain the first prediction result and the first confidence level of the sample; and input the second modality data set in the at least two modalities of data into the prediction model to obtain the second prediction result and the second confidence level of the sample; wherein, the first modality data set is a non-empty proper subset of the second modality data set.
[0084] Exemplarily, taking the training data set as as an example, the first modality data set can be represented as x (S) , and the second modality data set can be represented as x (T) . x (S) can also be the source modality data set, and x (T) is the modality data set obtained by adding one or more modalities of data to x (S) , and can also be called the target modality data set, and the embodiments of the present application do not make any limitations in this regard.
[0085] Optionally, a first modality data can be added to the first modality data set to obtain a second modality data set.
[0086] Optionally, the second modality data set after adding the first modality data can be determined as the first modality data set. Subsequently, a modality data can be continuously added to the first modality data set to obtain a new second modality data set. In this way, it is possible to add modality data one by one to the first modality data set to obtain different numbers of modality data sets as the second modality data set.
[0087] As an implementable manner, for each sample, the first modality data set x (S) can be initialized to a single modality data. Then, the second modality data set x (T) can be obtained by randomly adding modality data one by one. For the second modality data set x (T) after adding the modality, it can be regarded as another first modality data set x (S) , and the process of adding modality data is repeated until the second modality data set x (T) contains complete modality data.
[0088] Exemplarily, referring to Figure 6 , assume that each sample corresponds to three modality data. First, the first modality data M1 can be used as a modality data set and input into the prediction model to obtain the corresponding prediction result and confidence level. Then, the first modality data M1 and the second modality data M2 can be used as a modality data set and input into the prediction model to obtain the corresponding prediction result and confidence level. Subsequently, the first modality data M1, the second modality data M2, and the third modality data M3 can be used as a modality data set and input into the prediction model to obtain the corresponding prediction result to be predicted and confidence level.
[0089] As another implementable manner, for each sample, first, each sample can initialize the second modality data set x (T) to contain complete modality data. Then, the first modality data set x (S) can be obtained by randomly removing modality data one by one. For the first modality data set x (s) after removing the modality, it can be regarded as another second modality data set x (T) , and the process of removing modality data is repeated until the first modality data set x (S) contains a single modality data.
[0090] Exemplarily, continue to refer to Figure 6, assume that each sample corresponds to three types of modal data. First, the first type of modal data M1, the second type of modal data M2, and the third type of modal data M3 can be used as a modal data set, and input into the prediction model to obtain the corresponding prediction result and confidence level. Then, the first type of modal data M1 and the second type of modal data M2 can be used as a modal data set, and input into the prediction model to obtain the corresponding prediction result and confidence level. After that, the first type of modal data M1 can be used as a modal data set, and input into the prediction model to obtain the corresponding prediction result to be predicted and confidence level.
[0091] Optionally, for any two modal data sets, the one with fewer corresponding types of modal data can be called the first modal data set x (S) , and the one with more corresponding types of modal data can be called the second modal data set x (T) , that is where x (M) represents the complete M types of modal data in the training data set.
[0092] In some embodiments, the prediction result of the prediction model depends more on the first modal data than on any modal data in the first modal data set.
[0093] Among them, the first modal data is the modal data added to the first modal data set. After adding this first modal data to the first modal data set, the second modal data can be obtained. That is to say, the first modal data is not included in the first modal data set, while the second modal data set includes this first modal data.
[0094] Generally, the greater the dependence of the prediction result on a modal data, the more important this modal data is. That is to say, a relatively important modal data can be added to the first modal data to obtain the second modal data set. In other words, the second modal data set includes modal data with higher importance, while the first modal data set does not include this modal data with higher importance. Since the second modal data set contains modal data with higher importance, the second confidence level of the second prediction result corresponding to the second modal data should be higher.
[0095] As an example, when adding modal data one by one to the first modal data set x (S) to obtain the second modal data set x (T) , modal data with increasingly higher importance levels can be added in sequence according to the importance degree of the modal data. For example, in the Figure 6 example shown, the importance levels of the modal data M1, M2, and M3 increase in sequence. Therefore, for each input modal data set, the modal data M2 and M3 can be added in sequence.
[0096] As another example, when removing modal data one by one from the second-modal data set x (T) to obtain the first-modal data set x (S) , modal data with decreasing importance levels can be removed in sequence according to the importance degree of the modal data. For example, in the example shown in Figure 6 , the importance degrees of the modal data M1, M2, and M3 increase in sequence. Therefore, for each input modal data set, the modal data M3 and M2 can be removed in sequence.
[0097] Exemplarily, referring to Figure 5 , for the input first-modal data set x (S) , the feature extraction network 510 can perform feature extraction on it to obtain the first feature data corresponding to the first-modal data set x (S) . Then, by inputting the first feature data into the prediction network 520, the corresponding first prediction result can be obtained, and by inputting the first feature data into the confidence estimation network 530, the corresponding first confidence level can be obtained, where the first confidence level is used to evaluate the reliability of the first prediction result.
[0098] Similarly, for the input second-modal data set x (T) , the feature extraction network 510 can perform feature extraction on it to obtain the second feature data corresponding to the second-modal data set x (T) . Then, by inputting the second feature data into the prediction network 520, the corresponding second prediction result can be obtained, and by inputting the second feature data into the confidence estimation network 530, the corresponding second confidence level can be obtained, where the second confidence level is used to evaluate the reliability of the second prediction result.
[0099] Exemplarily, the first prediction result can be expressed as The first confidence level can be expressed as Conf(x (S) ); the second prediction result can be expressed as The second confidence level can be expressed as Conf(x (T) ).
[0100] Exemplarily, when the prediction network 520 is a classifier, the first prediction result and the second prediction result respectively correspond to the predicted class labels of the input sample x The first confidence level and the second confidence level respectively correspond to the confidence levels of the class labels of the sample x .
[0101] When the prediction network 520 is a generative network, the first prediction result and the second prediction result respectively correspond to two prediction results of the input sample x The first confidence level and the second confidence level respectively correspond to the confidence levels of two prediction results of the sample x .
[0102] In some embodiments, when the prediction network 520 includes a classification network, the output of the confidence network 530 is determined according to the probability of the predicted class of the sample. At this time, the output of the prediction network 520 is the predicted class of the sample and the probability corresponding to the predicted class.
[0103] As a specific example, the goal of the classifier is to learn a parameterized function: f(x,θ)→z, where θ is the parameter of the network and z is the output of the network. z can be called logits, and the logits are converted through a softmax layer to obtain a logits vector: That is, the normalized probability. Among them, the probability distribution of the sample x is defined as: The predicted class label is: That is, the class label corresponding to the maximum probability in the probability distribution of the sample x; the class label The probability of is: max y P(y∣∣θ,x (M) ), that is, the maximum probability value in the probability distribution of the sample x.
[0104] Continuing to refer to Figure 5 , the confidence estimation network 530 can be constrained according to the probability of the class label output by the prediction network 520. For example, according to the output of the confidence estimation network 530 and the class label The probability C * between the loss L conf constrains the confidence estimation network 530 so that the confidence output by the confidence estimation network 530 is the probability of the predicted class of the sample. Exemplarily, the confidence can be expressed as Conf(x (M) ) = max y P(y∣∣θ,x (M) ). Further, the prediction network 530 can also be constrained according to the class loss L between the class label ce output by the prediction network 520 and the true label y, so that it predicts the correct class label.
[0105] In some embodiments, when the prediction network 520 includes a generative network, the output of the confidence network 530 is determined according to the logits output by the generative network. At this time, the output of the prediction network 520 is the prediction result of a certain index of the sample.
[0106] As a specific example, the function of the generative network is: f(x,θ)→z, where θ is the parameter of the network and z is the output of the network. z can be referred to as logits. In this case, the logits may not need to be transformed by the softmax layer, and it is the prediction result of the model for specific metrics of the sample. Then, based on the prediction result and the true label y of the sample, the confidence of the prediction result is determined and expressed as:
[0107] Exemplarily, the prediction result output by the prediction network 520 can be used to constrain the confidence estimation network 530. For example, based on the error between the output of the confidence estimation network 530 and the prediction result the loss L is calculated conf for constraint, such that the confidence output by the confidence estimation network 530 is inversely proportional to the error from the prediction result, e.g., as Furthermore, based on the prediction result output by the prediction network 520 and the loss L between the prediction result and the true label y ce the prediction network 530 can be constrained to predict the correct prediction result.
[0108] 430, when the second confidence is less than the first confidence, the regularization loss is determined based on the first confidence and the second confidence.
[0109] Specifically, for a reliable prediction network, when one or more modalities of data are added for prediction, the confidence of the prediction result should not decrease. Since the set of second-modal data corresponding to the second confidence has more modalities than the set of first-modal data corresponding to the first confidence, when the second confidence is less than the first confidence, it means that the confidence of the prediction network decreases instead of increasing when the input modalities of data increase. At this time, based on the first confidence and the second confidence, a regularized learning (Calibrating Multimodal Learning, CML) strategy can be formulated to ensure that the confidence should increase when more modal data is input.
[0110] Optionally, the regularization loss can be determined based on the difference between the second confidence and the first confidence. Exemplarily, when the second confidence is less than the first confidence, the regularization loss is the difference between the second confidence and the first confidence, and this difference is a negative value; when the second confidence is equal to or greater than the first confidence, the regularization loss is zero.
[0111] Optionally, a confidence growth value (i.e., the difference between the second confidence level and the first confidence level) can be determined based on the first confidence level and the second confidence level. Exemplarily, the confidence growth value CI can be as shown in the following formula (1):
[0112] CI(x (T) ,x (S) )=Conf(x (T) )-Conf(x (S) ) (1)
[0113] where CI(x (T) ,x (S) ) being a negative value indicates that the confidence level of the prediction result decreases when one or more modal data are added to the input; CI(x (T) ,x (S) ) being a positive value indicates that the confidence level of the prediction result increases when one or more modal data are added to the input; CI(x (T) ,x (S) ) being 0 indicates that the confidence level of the prediction result remains unchanged when one or more modal data are added to the input.
[0114] Exemplarily, for any pair of inputs satisfying , it can be denoted as (S,T), and the corresponding regularization loss can be determined based on the above confidence growth value CI. For example, the regularization loss L (T,S) can be as shown in the following formula (2):
[0115] L (T,S) =min(0,CI(x (T) ,x (S) ))=min(0,Conf(x (T) )-Conf(x (s) )) (2)
[0116] That is to say, when the second confidence level Conf(x (T) ) corresponding to the second modal data set x (T) is less than the first confidence level Conf(x (S) ) corresponding to the first modal data set x (S) , the difference between the second confidence level Conf(x (T) ) and the first confidence level Conf(x (S) ) is used as the regularization loss to penalize the samples with decreased confidence levels, so as to ensure that the confidence level remains unchanged or increases when modal data are added.
[0117] Optionally, during the process of model training, for each sample, its total regularization loss can be jointly determined based on the regularization losses of input pairs with different numbers of modal data. Exemplarily, the total regularization loss L of each sampleCML Can satisfy all The input pair is integrated to obtain the following formula (3):
[0118]
[0119] Among them, for any input pair (T, S), it satisfies
[0120] Figure 7 Shows the Figure 6 Each satisfied A diagram showing whether to add regularization loss as a penalty for the input pair (T, S). Figure 7 As shown, for sample 1, when the two modal data sets are {M1} and {M1, M2}, if the confidence of the prediction result corresponding to the modal data set {M1, M2} is higher than the confidence of the prediction result corresponding to the modal data set {M1}, the corresponding regularization loss is 0, that is, there is no penalty; when the two modal data sets are {M1, M2} and {M1, M2, M3}, if the confidence of the prediction result corresponding to the modal data set {M1, M2, M3} is higher than the confidence of the prediction result corresponding to the modal data set {M1, M2}, the corresponding regularization loss is determined according to the difference between the two confidences, and the sample is penalized at this time.
[0121] For sample 2, when the two modal data sets are {M1} and {M1, M2} respectively, if the confidence of the prediction result corresponding to the modal data set {M1, M2} is lower than the confidence of the prediction result corresponding to the modal data set {M1}, the corresponding regularization loss is determined according to the difference between the two confidences, and the sample is penalized at this time; when the two modal data sets are {M1, M2} and {M1, M2, M3} respectively, if the confidence of the prediction result corresponding to the modal data set {M1, M2, M3} is equal to the confidence of the prediction result corresponding to the modal data set {M1, M2}, the corresponding regularization loss is 0, and there is no penalty at this time.
[0122] It should be noted that the exact calculation of the loss in formula (3) requires enumerating the regularization loss of all input pairs (T, S), which usually requires expensive computational costs. For example, for a sample containing three modal data M1, M2, and M3, when only one modal data is added each time, it satisfies The number of all input pairs is 9, which are {M1}{M1, M2}, {M1}{M1, M3}, {M2}{M1, M2}, {M2}{M2, M3}, {M3}{M1, M3}, {M3}{M2, M3}, {M1, M2}{M1, M2, M3}, {M2, M3}{M2, M3, M1}, {M1, M3}{M2, M3, M1}. And when successively adding modal data with increasing importance, it satisfies The number of all input pairs is 4. Assuming that the weights of M1, M2, and M3 increase in sequence, then the 4 input pairs are {M1}{M1, M2}, {M1, M2}{M1, M2, M3}, {M1}{M1, M3}, {M2}{M2, M3}.
[0123] Therefore, by adding modal data with higher dependence to the first modal data set to obtain the second modal data set, sampling of important input pairs among all input pairs of the sample can be realized, effectively reducing the number of input pairs that meet the conditions, thereby reducing the computational cost of the regularization loss.
[0124] In addition, in the medical field, usually the degree of dependence of the disease prediction result on the modal data corresponding to a certain medical examination is related to the cost of the medical examination. The higher the degree of dependence on a certain medical examination, the more expensive the cost corresponding to the examination. In related technologies, medical diagnosis models often have a relatively high degree of dependence on the modal data corresponding to these expensive examinations. However, in the embodiments of the present application, by successively adding modal data with higher dependence to the first modal data set to train the model, on the premise of ensuring the confidence of the prediction result, the dependence of the examination result on the successively added modal data can be reduced, which is beneficial to improving the utilization of all modal data (especially medical modal data with relatively low importance) by the model, thereby reducing the dependence of the model on expensive medical examinations and reducing the examination cost of patients.
[0125] 440. Determine the target loss according to the regularization loss, the first prediction result, the second prediction result, and the true label.
[0126] Specifically, the regularization loss determined in step 430 can be combined with the prediction losses between the first prediction result, the second prediction result, and the true label to jointly determine the target loss for constraining the prediction model.
[0127] As an implementable manner, the first joint loss can be determined according to the first prediction result, the second prediction result, and the true label; and the target loss can be determined according to the regularization loss and the first joint loss.
[0128] Specifically, the first combined loss can be determined according to the differences between the prediction results corresponding to different numbers of modal data sets and the true labels. Exemplarily, for Figure 7 the sampled input pairs shown in
[0129] L CE = CE(M1, y)+CE(M1 + M2, y)+CE(M1 + M2 + M3, y) (4)
[0130] where y is the target value, i.e., the prediction result. CE(M1, y) is the loss between the prediction result and the true label when the input is the modal data set {M1}, CE(M1 + M2, y) is the loss between the prediction result and the true label when the input is the modal data set {M1, M2}; CE(M1 + M2 + M3, y) is the loss between the prediction result and the true label when the input is the modal data set {M1, M2, M3}. The sum of the losses corresponding to different numbers of modal data sets is the combined loss corresponding to the prediction network, i.e., an example of the above first combined loss.
[0131] In some embodiments, the sum of the first combined loss and the regularization loss can be determined as the target loss. As shown in the following formula (5):
[0132] L = L CE + L CML (5)
[0133] As an implementable manner, the second combined loss can also be determined according to the first confidence, the second confidence, and the true label; and the target loss can be determined according to the above regularization loss, the first combined loss, and the second combined loss.
[0134] Specifically, the second combined loss can be determined according to the differences between the confidences corresponding to different numbers of modal data sets and the true label. Exemplarily, for Figure 7 the sampled input pairs shown in
[0135] L MSE = MSE(M1, u)+MSE(M1 + M2, u)+MSE(M1 + M2 + M3, u) (4)
[0137] Among them, u is the target value, i.e., the confidence level. MSE(M1, u) is the loss between the confidence level of the prediction result and the true label when the input is the modal data set {M1}, and MSE(M1 + M2, u) is the loss between the confidence level of the prediction result and the true label when the input is the modal data set {M1, M2}; MSE(M1 + M2 + M3, u) is the loss between the confidence level of the prediction result and the true label when the input is the modal data set {M1, M2, M3}. The sum of the losses corresponding to different numbers of modal data sets is the joint loss corresponding to the confidence level estimation network, which is an example of the above-mentioned second joint loss.
[0138] In some embodiments, the sum of the first joint loss, the second joint loss, and the regularization loss can be determined as the target loss, as shown in the following formula (5):
[0139] L = L CE + L MSE + L CML (5)
[0140] Among them, for the classification network, u is the maximum probability value in the probability distribution; for the generative network (such as the regression network), u is
[0141] 450. According to the target loss, the parameters of the prediction model are updated to obtain the trained prediction model.
[0142] Specifically, since the objective function includes the regularization loss for penalizing the samples whose confidence level decreases when adding modal data, the sample penalty can ensure that the confidence level increases or remains unchanged when adding modal data, which is beneficial to improving the reliability and robustness of the prediction confidence level of the prediction model.
[0143] Exemplarily, according to the target loss, the objective loss function can be optimized using the stochastic gradient descent algorithm, and then backpropagation is used to update the parameters of each network module in the prediction network until the training stop condition is met. The prediction network determined to meet the training stop condition is output as the trained hash model. For example, the parameters of Figure 5 the feature extraction network 510, the prediction network 520, and the confidence level estimation network 530 can be updated to obtain the trained prediction model. Exemplarily, for Figure 5 the network architecture in, after the model training is completed, the feature extraction network 510 and the confidence level estimation network 530 can be output as the trained confidence level estimation model; and the feature extraction network 510 and the prediction network 520 are output as the trained prediction model.
[0144] Therefore, in this application, the first confidence level is obtained by inputting the first modal data set of the sample into the prediction model, and the second confidence level is obtained by inputting the second modal data set of the sample into the prediction model, where the second modal data set has additional modal data compared to the first modal data set. Further, when the second confidence level is less than the first confidence level, the regularization loss is determined, and the model is constrained to update its parameters according to this regularization loss, so as to impose a penalty on the samples whose prediction confidence level decreases rather than increases when additional modal data is added, thereby improving the reliability and robustness of the prediction confidence level of the prediction model.
[0145] As a specific embodiment, training data can be obtained. The training data includes at least two types of modal medical data of patient samples, such as CT images, MRI images, pathological image data, etc. The training data also includes disease category labels of the patients, such as disease types, lesion degrees, or survival durations, etc. Then, the first modal medical data set in the at least two types of modal medical data is input into the prediction model to obtain the first prediction result and the first confidence level of the patient sample; the second modal medical data set in the at least two types of modal medical data is input into the prediction model to obtain the second prediction result and the second confidence level of the patient sample; where the first modal medical data set is a non-empty proper subset of the second modal medical data set. When the second confidence level is less than the first confidence level, the regularization loss is determined according to the first confidence level and the second confidence level; the objective loss is determined according to the regularization loss, the first prediction result, the second prediction result, and the true label; the parameters of the prediction model are updated according to the objective loss to obtain the trained prediction model. This prediction model can be used to input at least one type of modal medical data of a patient and output the disease prediction result and confidence level of the patient. The disease prediction result can be, for example, the disease type, lesion degree, or survival duration, etc.
[0146] Optionally, the first modal data set can be initialized as a single type of modal medical data, such as one of CT images, MRI images, pathological image data. Optionally, by adding one type of modal medical data to the first modal medical data set, the second modal data set can be obtained. Optionally, the second modal data set after adding the modal medical data can also be determined as the new first modal data set, and further, by adding one type of modal medical data to this new first modal data set, a new second modal data set can be obtained. And so on, to realize obtaining the second modal data set by adding modal medical data to the first model medical data set one by one. Optionally, the second modal data set can include all the modal medical data of the patient sample. Optionally, the modal data with increasing importance levels can be added in sequence according to the importance degree of the modal medical data.
[0147] Figure 8FIG. 800 is a schematic flowchart of a prediction method according to an embodiment of the present application. The method 800 can be executed by any electronic device with data processing capabilities. For example, the electronic device can be implemented as a server or a terminal device. Again, for example, the electronic device can be implemented as Figure 1 the execution device 104 in Figure 8 . The present application does not limit this. As
[0148] shown in FIG. 800, the prediction method 800 includes the following steps 810 to 820.
[0149] Exemplarily, in the medical field, at least one multimodal data of the sample to be predicted can include at least one of CT image data, MRI image data, pathological image data, etc. of a patient; in the field of Internet advertising, at least one multimodal data of the sample to be predicted can include at least one of click data, browsing history data, and profile data of an object.
[0150] 820. Input at least two modalities of data into the prediction model to obtain the prediction result and confidence of the sample to be predicted; wherein, the prediction model is trained according to the above prediction model training method 400.
[0151] Specifically, at least one modality of data obtained in step 810 can be input into the prediction model, and the prediction model outputs the prediction result of the sample to be predicted and the confidence of the prediction result. Among them, the prediction model can include a classification network and a generative network. Specifically, the training process of the prediction model can refer to the relevant description of the prediction model training method above.
[0152] Exemplarily, in the medical field, when the prediction sample includes various modalities of examination data of a patient, the prediction result can include the disease category, or the prediction of the lesion degree, or the prediction of the survival time, etc. In the field of Internet advertising, when the prediction sample includes various modalities of data of an object, the prediction result can include the type of advertisement recommended for the object.
[0153] Therefore, the embodiment of the present application updates the parameters by the regularization loss constraint model, realizes the penalty for the samples whose prediction confidence decreases instead of increasing when adding modality data, thereby improving the reliability and robustness of the prediction confidence of the prediction model, and enhancing the credibility of the prediction result of the prediction model.
[0154] As a specific embodiment, at least one modality of medical data of the patient to be predicted can be obtained, such as CT images, MRI images, pathological image data, etc. Then, input the at least one modality of medical data into the prediction model, and the disease prediction result and confidence of the patient to be predicted can be obtained. Exemplarily, the disease prediction result can include the disease type, the lesion degree, or the survival duration, etc.
[0155] The specific embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, for the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. Again, for example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed by the present application.
[0156] It should also be understood that in various method embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. It should be understood that these sequence numbers can be interchanged under appropriate circumstances, so that the embodiments of the present application described can be implemented in an order other than those illustrated or described.
[0157] The method embodiments of the present application have been described in detail above. Below, in conjunction with Figures 9 to 11 , the device embodiments of the present application will be described in detail.
[0158] Figure 9 is a schematic block diagram of a prediction model training device 10 according to an embodiment of the present application. As Figure 9 shown, the device 10 may include an acquisition unit 11, a prediction unit 12, a determination unit 13, and a parameter update unit 14.
[0159] The acquisition unit 11 is configured to acquire training data, where the training data includes at least two types of modal data of a sample and the true label corresponding to the sample;
[0160] The prediction unit 12 is configured to input a first set of modal data in the at least two types of modal data into a prediction model to obtain a first prediction result and a first confidence level of the sample; and input a second set of modal data in the at least two types of modal data into the prediction model to obtain a second prediction result and a second confidence level of the sample; wherein, the first set of modal data is a non-empty proper subset of the second set of modal data;
[0161] The determination unit 13 is configured to determine a regularization loss according to the first confidence level and the second confidence level when the second confidence level is less than the first confidence level;
[0162] The determining unit 13 is further configured to determine a target loss according to the regularization loss, the first prediction result, the second prediction result, and the true label;
[0163] A parameter updating unit 14, configured to update parameters of the prediction model according to the target loss to obtain the trained prediction model.
[0164] In some embodiments, the determining unit 13 is specifically configured to:
[0165] Determine a first combined loss according to the first prediction result, the second prediction result, and the true label;
[0166] Determine the target loss according to the regularization loss and the first combined loss.
[0167] In some embodiments, the determining unit 13 is specifically configured to:
[0168] Determine a second combined loss according to the first confidence level, the second confidence level, and the true label;
[0169] Determine the target loss according to the regularization loss, the first combined loss, and the second combined loss.
[0170] In some embodiments, the prediction unit 12 is further configured to:
[0171] Add a first modal data to the first modal data set to obtain the second modal data set.
[0172] In some embodiments, the prediction unit 12 is further configured to:
[0173] Initialize the first modal data set as a single modal data.
[0174] In some embodiments, the prediction unit 12 is further configured to:
[0175] Determine the second modal data set after adding the first modal data as the first modal data set.
[0176] In some embodiments, the prediction result of the prediction model depends more on the first modal data than on any modal data in the first modal data set.
[0177] In some embodiments, the prediction model includes a feature extraction network, a prediction network, and a confidence estimation network. Among them, the feature extraction network is used to extract feature data of the input modal data, the prediction network is used to obtain a prediction result of a sample according to the input feature data; the confidence estimation network is used to obtain the confidence of the prediction result according to the input feature data.
[0178] In some embodiments, when the prediction network includes a generative network, the output of the confidence network is determined according to the logical value output by the generative network.
[0179] In some embodiments, when the prediction network includes a classification network, the output of the confidence network is determined according to the probability of the predicted class of the sample.
[0180] In some embodiments, the at least two types of modality data include at least two types of multimodal medical examination data, and the true label includes the diagnosis result corresponding to the sample.
[0181] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 9 The illustrated prediction model training device 10 can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the prediction model training device 10 respectively implement the corresponding processes in the above method 400. For the sake of brevity, it will not be elaborated here.
[0182] Figure 10 is a schematic block diagram of a prediction device 20 according to an embodiment of the present application. As Figure 10 shown, the device 20 may include an acquisition unit 21 and a prediction model 22.
[0183] The acquisition unit 21 is configured to acquire at least one type of modality data of a sample to be predicted;
[0184] The prediction model 22 is configured to input the at least one type of modality data to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the above prediction model training method.
[0185] The device according to the embodiment of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in hardware, can be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read only memory, a programmable read only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0186] Figure 11 It is a schematic block diagram of an electronic device provided by an embodiment of the present application.
[0187] As Figure 11 shown, the electronic device 30 may include:
[0188] A memory 33 and a processor 32. The memory 33 is used to store a computer program 34 and transmit the program code 34 to the processor 32. In other words, the processor 32 can call and run the computer program 34 from the memory 33 to implement the method provided in the embodiment of the present application.
[0189] In some embodiments, the processor 32 can call and run the computer program 34 from the memory 33 to implement the prediction model training method provided in the embodiment of the present application, including:
[0190] Obtain training data, where the training data includes at least two modal data of a sample and the true label corresponding to the sample;
[0191] Input a first set of modal data in the at least two sets of modal data into the prediction model to obtain a first prediction result and a first confidence level of the sample; and input a second set of modal data in the at least two sets of modal data into the prediction model to obtain a second prediction result and a second confidence level of the sample; wherein, the first set of modal data is a non-empty proper subset of the second set of modal data;
[0192] When the second confidence level is less than the first confidence level, determine a regularization loss according to the first confidence level and the second confidence level;
[0193] Determine an objective loss according to the regularization loss, the first prediction result, the second prediction result, and the true label;
[0194] Update the parameters of the prediction model according to the objective loss to obtain the trained prediction model.
[0195] For example, the processor 32 can be used to execute the steps in the above method 400 according to the instructions in the computer program 34.
[0196] In some embodiments, the processor 32 can call and run the computer program 34 from the memory 33 to implement the prediction method provided in the embodiment of the present application, including:
[0197] Obtain at least one modal data of a sample to be predicted;
[0198] Input the at least one modality data into a prediction model to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the above prediction model training method.
[0199] For example, the processor 32 can be used to execute the steps in the above method 800 according to the instructions in the computer program 34.
[0200] In some embodiments of the present application, the processor 32 may include but is not limited to:
[0201] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0202] In some embodiments of the present application, the memory 33 includes but is not limited to:
[0203] A volatile memory and / or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0204] In some embodiments of the present application, the computer program 34 may be divided into one or more units, which are stored in the memory 33 and executed by the processor 32 to complete the method provided by the present application. The one or more units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 34 in the electronic device 30.
[0205] Optionally, as Figure 11 shown, the electronic device 30 may further include:
[0206] a transceiver 33, which may be connected to the processor 32 or the memory 33.
[0207] Among them, the processor 32 may control the transceiver 33 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 33 may include a transmitter and a receiver. The transceiver 33 may further include an antenna, and the number of antennas may be one or more. It should be understood that each component in the electronic device is connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.
[0208] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiments. Or, the embodiments of the present application also provide a computer program product including instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiments.
[0209] When implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0210] It can be understood that in the specific implementation manners of the present application, when the above embodiments of the present application are applied to specific products or technologies and involve relevant data such as user information, user permission or consent needs to be obtained, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0211] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0212] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0213] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0214] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for training a prediction model, characterized in that, Including: Obtaining training data, where the training data includes at least two modalities of data of the samples and the true label corresponding to the samples; Inputting the first modality data set among the at least two modalities of data into the prediction model to obtain a first prediction result and a first confidence level of the sample; and inputting the second modality data set among the at least two modalities of data into the prediction model to obtain a second prediction result and a second confidence level of the sample; wherein, the first modality data set is a non-empty proper subset of the second modality data set; When the second confidence level is less than the first confidence level, determining a regularization loss according to the first confidence level and the second confidence level; Determining an objective loss according to the regularization loss, the first prediction result, the second prediction result and the true label; Updating the parameters of the prediction model according to the objective loss to obtain the trained prediction model.
2. The method according to claim 1, characterized in that, The determining the objective loss according to the regularization loss, the first prediction result, the second prediction result and the true label includes: Determining a first joint loss according to the first prediction result, the second prediction result and the true label; Determining the objective loss according to the regularization loss and the first joint loss.
3. The method according to claim 2, wherein Further including: Determining a second joint loss according to the first confidence level, the second confidence level and the true label; Wherein, the determining the objective loss according to the regularization loss and the first joint loss includes: Determining the objective loss according to the regularization loss, the first joint loss and the second joint loss.
4. The method according to claim 1, characterized in that, Further including: Adding a first modality data to the first modality data set to obtain the second modality data set.
5. The method according to claim 4, wherein Further including: Initializing the first modality data set as a single modality data.
6. The method according to claim 4, wherein Further including: Determining the second modality data set after adding the first modality data as the first modality data set.
7. The method according to claim 4, wherein The prediction result of the prediction model depends more on the first modality data than on any modality data in the first modality data set.
8. The method according to claim 1, wherein The prediction model includes a feature extraction network, a prediction network and a confidence estimation network, wherein the feature extraction network is used to extract feature data of the input modality data, the prediction network is used to obtain a prediction result of the sample according to the input feature data; the confidence estimation network is used to obtain the confidence level of the prediction result according to the input feature data.
9. The method according to claim 8, characterized in that, When the prediction network includes a generative network, the output of the confidence network is determined according to the logical value output by the generative network.
10. The method according to claim 8, wherein When the prediction network includes a classification network, the output of the confidence network is determined according to the probability of the predicted class of the sample.
11. The method according to any one of claims 1 to 10, characterized in that, The at least two modalities of data include at least two multimodal medical examination data, and the true label includes the diagnosis result corresponding to the sample.
12. A prediction method, characterized in that, Including: Obtaining at least one modality of data of the sample to be predicted; Input the at least one modality data into a prediction model to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the method described in any one of claims 1-11.
13. A prediction model training device, characterized in that Comprising: An acquisition unit configured to acquire training data, the training data including at least two modality data of a sample and a true label corresponding to the sample; A prediction unit configured to input a first modality data set among the at least two modality data into the prediction model to obtain a first prediction result and a first confidence level of the sample; and input a second modality data set among the at least two modality data into the prediction model to obtain a second prediction result and a second confidence level of the sample; wherein, the first modality data set is a non-empty proper subset of the second modality data set; A determination unit configured to determine a regularization loss according to the first confidence level and the second confidence level when the second confidence level is less than the first confidence level; The determination unit is further configured to determine an objective loss according to the regularization loss, the first prediction result, the second prediction result and the true label; A parameter update unit configured to update parameters of the prediction model according to the objective loss to obtain the trained prediction model.
14. A prediction device, characterized in that, Comprising: An acquisition unit configured to acquire at least one modality data of a sample to be predicted; A prediction model configured to input the at least one modality data to obtain a prediction result and a confidence level of the sample to be predicted; wherein, the prediction model is trained according to the method described in any one of claims 1-11.
15. An electronic device, characterized in that, Comprising a processor and a memory, wherein instructions are stored in the memory, and when the processor executes the instructions, the processor executes the method described in any one of claims 1-12.
16. A computer storage medium, characterized in that, For storing a computer program, the computer program comprising instructions for executing the method described in any one of claims 1-12.
17. A computer program product, characterized in that, Comprising computer program code, when the computer program code is run by an electronic device, the electronic device is caused to execute the method described in any one of claims 1-12.