Data processing method and related device
By obtaining and blocking the classification accuracy of multimodal samples, and using greedy algorithms to restore important modalities, the problem of too long analysis in the multimodal classification model is solved, and the analysis efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202410076732.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-18
AI Technical Summary
In multimodal classification models, the existing technology needs to traverse the arrangement and combination of sample modal information, resulting in an exponential increase in the number of model training times. How to efficiently analyze the importance of modality has become an urgent problem to be solved.
By obtaining preset samples and labels, masking samples of at least one modal, calculating the classification accuracy rate. When the accuracy rate is less than or equal to the threshold, the modal is restored until the preset conditions are met to generate the target modal. The greedy algorithm is used to reduce the number of traversals and find important modals.
The modal importance analysis of efficient screening models is realized, and the modality that has a greater impact on the model is found, reducing analysis time and resource consumption.
Smart Images

Figure CN120336843A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for data processing and related devices. Background Art
[0002] With the continuous development of artificial intelligence technology, during the process of model training and use, the number of samples required is also increasing. Among these samples, there may be samples of different modalities. In a classification model, the output content of the model is usually a certain category. When multi-modal samples are applied to a classification model, the modalities have different influences on the classification accuracy of multiple categories.
[0003] In a multi-modal classification model, by traversing the permutations and combinations of sample modality information, the influence of different combinations of modality information on the output categories of the model is calculated. However, since this method requires a large amount of verification calculations, if the number of modalities in a multi-modal classification model is N, then the number of model training times required to traverse the permutations and combinations of sample modality information for this multi-modal classification model is N! times. Therefore, the time required for analyzing the modality importance of the model increases with the increase in the number of modalities in the samples. How to efficiently analyze the modality importance of the model has become an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide a method for data processing and related devices for efficiently analyzing the modality importance of a model.
[0005] This application provides a method for data processing in one aspect, including:
[0006] Obtain a preset sample and the label of the preset sample, where the preset sample includes samples of at least two modalities;
[0007] Process the first sample and the label of the first sample to obtain a first correct rate, where the first correct rate is the correct rate of a preset model for classifying the first sample, and the first sample is a sample obtained by masking at least one modality in the preset sample;
[0008] When the first correct rate is less than or equal to a threshold, process the second sample and the label of the second sample to obtain a second correct rate, where the second correct rate is the correct rate of the preset model for classifying the second sample, and the second sample is a sample obtained by restoring the first modality of the sample based on the first sample, the first modality is included in at least one modality, and the sample of the second modality is included in the preset sample;
[0009] When the second correct rate meets a preset condition, generate a target modality, and the target modality is included in at least one modality.
[0010] This application provides a data processing device in a second aspect, including:
[0011] An acquisition unit for acquiring a preset sample and a label of the preset sample, where the preset sample includes samples of at least two modalities;
[0012] A processing unit for processing a first sample and a label of the first sample to obtain a first accuracy rate, where the first accuracy rate is the accuracy rate of a preset model for classifying the first sample, and the first sample is a sample obtained by masking at least one modality in the preset sample;
[0013] The processing unit is further configured to, when the first accuracy rate is less than or equal to a threshold, process a second sample and a label of the second sample to obtain a second accuracy rate, where the second accuracy rate is the accuracy rate of the preset model for classifying the second sample, the second sample is a sample obtained by restoring a first modality based on the first sample, the first modality is included in at least one modality, and the sample of the second modality is included in the preset sample;
[0014] A generating unit for generating a target modality when the second accuracy rate meets a preset condition, where the target modality is included in at least one modality.
[0015] In a possible implementation manner of the second aspect, the acquisition unit is specifically configured to:
[0016] Acquire an input sample, where the input sample includes samples of at least two modalities;
[0017] Input the input sample into a preset model to obtain a label of the input sample;
[0018] Classify the input sample according to the label of the input sample to obtain a preset sample, and each sample in the preset sample includes a preset label.
[0019] In a possible implementation manner of the second aspect, the generating unit is specifically configured to, when the second accuracy rate is less than or equal to the threshold, use the first modality as the target modality.
[0020] In a possible implementation manner of the second aspect, the generating unit is specifically configured to:
[0021] When the difference between the second accuracy rate and the first accuracy rate is less than or equal to the threshold, process a third sample and a label of the third sample to obtain a third accuracy rate, where the third accuracy rate is the accuracy rate of the preset model for classifying the third sample, the third sample is a sample obtained by restoring a second modality based on the second sample, the second modality is included in at least one modality, and the sample of the third modality is included in the preset sample;
[0022] When the difference between the third accuracy rate and the first accuracy rate is greater than the threshold, use the first modality and the second modality as the target modalities.
[0023] In a possible implementation manner of the second aspect, the generating unit is specifically configured to:
[0024] When the difference between the second correct rate and the first correct rate is less than or equal to the threshold, process the third sample and the label of the third sample to obtain the third correct rate, where the third correct rate is the correct rate of the preset model for classifying the third sample, the third sample is a sample obtained by restoring the second modality based on the second sample, the second modality is included in at least one modality, and the samples of the third modality are included in the preset samples;
[0025] When the difference between the third correct rate and the first correct rate is greater than the threshold, use the second modality as the target modality.
[0026] In a possible implementation manner of the second aspect, the processing unit is further configured to process the preset samples to obtain the first samples.
[0027] In a possible implementation manner of the second aspect, the processing unit is specifically configured to:
[0028] Extract features from the preset samples to obtain the feature matrix of the preset samples;
[0029] Set the feature matrices corresponding to at least one modality in the feature matrix of the preset samples to 0 to obtain the feature matrix of the first samples.
[0030] The third aspect of the present application provides a computer device, including:
[0031] A memory, a transceiver, a processor, and a bus system;
[0032] Wherein, the memory is used to store programs;
[0033] The processor is used to execute the programs in the memory, including executing the methods of the above aspects;
[0034] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate.
[0035] The fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is enabled to execute the methods of the above aspects.
[0036] The fifth aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.
[0037] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0038] An embodiment of the present application provides a data processing method and related device, which are used to efficiently analyze the importance of modalities for a model. Obtain a preset sample and the label of the preset sample, where the preset sample includes samples of at least two modalities. Mask samples of at least one modality from the preset sample to obtain a first sample, combine the first sample and the label of the first sample, calculate the correct rate of the preset model for classifying the first sample to obtain a first correct rate. When the first correct rate is less than or equal to a threshold, it is considered that at least one modality includes a target modality. Restore the samples of the first modality in the first sample to obtain a second sample, combine the second sample and the label of the second sample, calculate the correctness of the preset model for classifying the second sample to obtain a second correct rate. When the second correct rate meets a preset condition, generate the target modality. Through simple screening, the modality importance analysis of the model is completed, and the modalities that are relatively important for the model are found. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 FIG. is a schematic structural diagram of a data processing system provided by an embodiment of the present application;
[0040] Figure 2 FIG. is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0041] Figure 3 FIG. is a schematic diagram of modality masking provided by an embodiment of the present application;
[0042] Figure 4 FIG. is a schematic diagram of modality restoration provided by an embodiment of the present application;
[0043] Figure 5 FIG. is a schematic flowchart of the application of the data processing method provided by an embodiment of the present application;
[0044] Figure 6 FIG. is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0045] Figure 7 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] An embodiment of the present application provides a data processing method and related device for efficiently analyzing the importance of modalities for a model.
[0047] In the description, claims and the above-mentioned drawings of this application, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0048] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0049] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0050] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language used by people in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The important technology for model training in the field of artificial intelligence, the pre-trained model, is developed from the large language model (LLM) in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0051] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0052] Multi-modal AI is an artificial intelligence technology that combines multiple perceptual information sources. It uses multiple data modalities such as vision, speech, and text for information processing and analysis to improve the model's understanding and prediction capabilities. For example, multi-modal artificial intelligence can use patient information from multiple sources, including electronic health records, medical imaging, and test results, to compile a more comprehensive patient profile. One modality described in this application can be a medical test result of a patient. For example, a chest X-ray result or an MRI result of a certain part can be regarded as a modality, which can help healthcare practitioners improve patient treatment outcomes and decision-making.
[0053] To facilitate the understanding of the technical solutions provided by the embodiments of this application, some key terms used in the embodiments of this application are explained here:
[0054] Greedy algorithm: Also known as the greedy method, it is a commonly used method for finding the optimal solution problem. This method generally divides the solution process into several steps, but each step applies the greedy principle to select the best / optimal choice (the most favorable choice locally) in the current state, and hopes that the final stacked result is also the best / optimal solution.
[0055] With the continuous development of artificial intelligence technology, during the process of model training and use, the number of samples required is also increasing continuously. Among these samples, there may be samples of different modalities. In a classification model, the output content of the model is usually a certain category. When multi-modal samples are applied to a classification model, the modality has different influences on the classification accuracy of multiple categories.
[0056] In a multi-modal classification model, by traversing the permutations and combinations of sample modal information, the influence of different combinations of modal information on the output categories of the model is calculated. However, since this method requires a large number of verification calculations, if the number of modalities in the multi-modal classification model is N, the number of model training times required for this multi-modal classification model to traverse the permutations and combinations of sample modal information is N! times. For example, during a physical examination, a patient may undergo N different test items to screen for M diseases. Among them, the N different test items can be understood as data samples of N modalities. When a patient undergoes a physical examination, they may hope to focus on screening some diseases with potential risks. If all N different test items are examined simultaneously, it will consume a large amount of time and money.
[0057] In the above example, for any one disease (category), how to obtain accurate screening results (labels) through an appropriate number of tests (modalities). That is, how to efficiently distinguish the influence of different modalities on the classification results, so as to efficiently conduct modal importance analysis has become an urgent problem to be solved.
[0058] This application proposes a solution to the above problems. It can obtain a preset data set, which includes preset samples and labels of the preset samples, and at least two modalities of samples are included in the preset samples. Mask at least one modality of samples from the preset samples to obtain a first sample. Combine the first sample and the label of the first sample, and calculate the accuracy rate of the preset model in classifying the first sample to obtain a first accuracy rate. When the first accuracy rate is less than or equal to the threshold, it is considered that at least one modality includes a target modality. Restore the samples of the first modality in the first sample to obtain a second sample. Combine the second sample and the label of the second sample, and calculate the correctness of the preset model in classifying the second sample to obtain a second accuracy rate. When the second accuracy rate meets the preset conditions, generate the target modality. Through simple screening, the modal importance analysis of the model is completed, and the modalities that are relatively important for the model are found.
[0059] For ease of understanding, please refer to Figure 1 , Figure 1 which is an application environment diagram of the data processing method in the embodiments of this application, as Figure 1As shown in the figure, the data processing method in the embodiment of the present application is applied to a data processing system. The data processing system includes: a server and a terminal device; wherein, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiment of the present application does not limit this here.
[0060] First, the server obtains a preset sample and the label of the preset sample, and the preset sample contains samples of at least two modalities;
[0061] Process the first sample and the label of the first sample to obtain the first correct rate. The label of the first sample is the result obtained by analyzing the first sample by a preset model, the first correct rate is the correct rate of classifying the first sample by the preset model, and the first sample is the sample obtained by masking at least one modality in the preset sample;
[0062] When the correct rate of the first sample is less than or equal to the threshold, process the second sample and the label of the second sample to obtain the second correct rate. The label of the second sample is the result obtained by analyzing the second sample by a preset model, the second correct rate is the correct rate of classifying the second sample by the preset model, the second sample is the sample obtained by restoring the first modality in the first sample, the first modality is included in at least one modality, and the sample of the second modality is included in the preset sample;
[0063] When the second correct rate meets the preset condition, generate a target modality, and the target modality is included in at least one modality.
[0064] Next, from the perspective of the server, the data processing method in the present application will be introduced. Please refer to Figure 2 , the data processing method in the present application includes: step S101 to step S112. Specifically:
[0065] S101. Obtain a preset sample and the label of the preset sample;
[0066] Exemplarily, collect input samples, input the input samples into a preset model to obtain the labels of the input samples, and the input samples contain samples of at least two modalities.
[0067] Classify the input samples according to the labels of the input samples to obtain preset samples and the labels of the preset samples. Each sample in all the preset samples contains a preset label, and the preset samples include samples of at least two modalities.
[0068] In the embodiments of the present application, taking the samples of 16 modalities included in the preset samples as an example, the solution proposed in the present application is introduced.
[0069] It can be understood that the description of the number of modalities included in the preset samples here is only an example. In actual applications, it should be described in combination with specific application scenarios and is not limited here.
[0070] In the embodiments of the present application, by collecting input samples, inputting the input samples into a preset model to obtain the labels of the input samples, classifying the input samples according to the labels of the input samples to obtain preset samples and the labels of the preset samples. Among them, any two labels in the labels of the preset samples are the same. By screening the preset samples corresponding to the preset categories, multi-modal importance analysis is performed using the preset samples and the labels of the preset samples to obtain the target modality, and the target modality is the modality that has a significant impact on the prediction accuracy rate of the preset category of the preset model.
[0071] S102. Mask the samples of at least one modality in the preset samples to obtain the first sample;
[0072] Randomly mask the samples of one modality among the 16 modalities of the preset samples to obtain the first sample. The label of the first sample is the label corresponding to the first sample among the labels of the preset samples.
[0073] When the first accuracy rate is greater than the threshold, randomly mask one modality among the remaining 15 modalities of the first sample to obtain the updated first sample.
[0074] Exemplarily, as Figure 3 shown, first, mask the samples corresponding to modality one in the preset samples. If the first accuracy rate is greater than the threshold, then mask the samples corresponding to modality three in the preset samples to obtain the updated first sample. If the updated first accuracy rate is still greater than the threshold, then mask the samples corresponding to modality ten in the preset samples. Among them, the first accuracy rate is the accuracy rate when the preset model processes the first sample, and the updated first accuracy rate is the accuracy rate when the preset model processes the updated first sample.
[0075] It can be understood that the description of masking the samples corresponding to the modality here is only an example. In actual applications, the specific order of masking the samples corresponding to the modality is randomly masked and is not limited here.
[0076] There are various operations for masking the preset samples here, and these operations will be introduced separately below:
[0077] Method 1, deletion method;
[0078] Delete the samples of any one of the 16 modalities in the preset samples to obtain the first samples. For example, delete the samples corresponding to Modality 1 to obtain the first samples.
[0079] Method 2, noise addition method:
[0080] First, perform feature extraction on the preset samples to obtain the feature matrix of the preset samples;
[0081] Add noise to the feature matrix corresponding to the samples of any one of the 16 modalities in the feature matrix of the preset samples to obtain the feature matrix of the first samples.
[0082] Method 3, zeroing method:
[0083] First, perform feature extraction on the preset samples to obtain the feature matrix of the preset samples;
[0084] Zero out the feature matrix corresponding to the samples of any one of the 16 modalities in the feature matrix of the preset samples to obtain the feature matrix of the first samples.
[0085] In the embodiments of the present application, feature extraction is performed on the preset samples to obtain the feature matrix of the preset samples, and the feature matrix corresponding to the samples of any one modality in the feature matrix of the preset samples is zeroed out to obtain the feature matrix of the first samples, which efficiently realizes the shielding of the samples of any one modality in the preset samples and improves the feasibility of the solution.
[0086] It can be understood that the description of the method for shielding Modality 1 here is only an example. In actual applications, the shielding method should be selected in combination with the specific application scenario, and no limitation is made here.
[0087] In the embodiments of the present application, at least one modality of samples in the preset samples can be shielded by the deletion method, the noise addition method, and the zeroing method, which improves the flexibility of the solution implementation.
[0088] S103. Process the first samples and the labels of the first samples to obtain the first correct rate;
[0089] Among them, the label of the first samples is the label of the preset samples.
[0090] It can be understood that when the first samples are the samples obtained by shielding the samples of Modality A and Modality B, the label of the first samples is the label of the samples obtained by shielding Modality A, or the label of the samples obtained by inputting the samples of shielding Modality B into the preset model to obtain the samples of shielding Modality A, or the label of the samples of shielding Modality B, and no limitation is made here.
[0091] Exemplarily, input the first sample into a preset model to obtain a first label, calculate the probability that the first label is the same as the label of the first sample, and obtain a first accuracy rate.
[0092] S104. Determine whether the first accuracy rate is less than or equal to a threshold value;
[0093] Determine whether the first accuracy rate is less than or equal to 50%. The description of the threshold value here is only an example, and in specific implementation, it can be adjusted according to the specific usage scenario, and no limitation is made here.
[0094] If so, execute step S105;
[0095] If not, execute step S102.
[0096] S105. Recover the sample corresponding to the first modality based on the first sample to obtain a second sample;
[0097] Wherein, the first modality is a modality among at least one modality masked in step S102, and the label of the second sample is the label corresponding to the second sample among the labels of the preset samples.
[0098] Exemplarily, as Figure 4 shown, the currently masked modalities include modality one, modality three, and modality ten. Recover the sample corresponding to modality three to obtain a second sample.
[0099] It should be noted that the first modality here is a modality, which can be any one of the currently masked multiple modalities, and the selection of the first modality is random, and no limitation is made here.
[0100] Combined with the different masking methods in the foregoing S102, the specific recovery means may also be different:
[0101] Case 1. Recovery corresponding to the deletion method:
[0102] When the first sample is obtained by deleting the samples corresponding to at least one modality, add the sample corresponding to the first modality to the first sample to obtain a second sample, and the sample corresponding to the first modality is included in the preset samples.
[0103] Case 2. Recovery corresponding to the noise addition method;
[0104] When the feature vector of the first sample is a feature vector obtained by adding noise to the feature vector of the preset sample, remove the noise of the sample corresponding to the first modality to obtain a second sample.
[0105] Case 3. Recovery corresponding to the zeroing method.
[0106] When the feature vector of the first sample is the feature vector obtained by zeroing the feature vector of the preset sample, restore the feature vector of the sample corresponding to the first modality to its original value to obtain the second sample.
[0107] It can be understood that the description of the method for restoring the first modality here is only an example. In actual applications, the restoration method should be selected according to the specific application scenario, and no limitation is made here.
[0108] S106. Process the second sample and the label of the second sample to obtain the second correct rate.
[0109] Among them, the label of the second sample can be the same as the first label.
[0110] Exemplarily, input the second sample into the preset model to obtain the second label, calculate the probability that the second label is the same as the label of the second sample, and obtain the second correct rate.
[0111] S107. Determine whether the difference between the second correct rate and the first correct rate is less than or equal to the threshold.
[0112] Determine whether the difference between the second correct rate and the first correct rate is less than or equal to 50%. The description of the threshold here is only an example. In specific implementation, it can be adjusted according to the specific usage scenario, and no limitation is made here.
[0113] If not, execute step S108.
[0114] If so, execute step S109.
[0115] S108. Take the first modality as the target modality.
[0116] When the prediction correct rate of the preset model for the second sample is significantly improved compared with the prediction correct rate of the first sample after restoring the sample corresponding to the first modality, take the first modality as the target modality, that is, the first modality is the key modality of the preset model.
[0117] In the embodiment of the present application, after restoring the first sample once, the prediction accuracy of the preset model for the sample is significantly improved. Taking the restored modality as the target modality and using the principle of the greedy algorithm, without traversing all modalities, while ensuring the accuracy of the importance of multiple modalities, the time required for analyzing the importance of multiple modalities is reduced, and the efficiency of analyzing the importance of multiple modalities is improved.
[0118] S109. Based on the second sample, restore the sample corresponding to the second modality to obtain the third sample.
[0119] Among them, the second modality is any modality other than the first modality among the at least one modality masked in step S102, and the label of the second sample is the label corresponding to the second sample in the label of the preset sample.
[0120] Exemplarily, as Figure 4 shown, the currently blocked modalities include Modality 1 and Modality 10. Restore the samples corresponding to Modality 1 to obtain the third sample.
[0121] When the difference between the third accuracy rate and the first accuracy rate is less than or equal to the threshold, restore the samples corresponding to the third modality in the third sample to obtain the updated third sample, where the third modality is any modality other than the first modality and the second modality among the at least one modality blocked in step S102, and there is no restriction here. The label of the updated third sample is the label corresponding to the updated third sample among the labels of the preset samples.
[0122] When the difference between the third accuracy rate and the first accuracy rate is greater than the threshold, stop restoring the samples of the modality in the at least one blocked modality, otherwise, keep restoring the samples of the modality in the at least one blocked modality until all the samples of the modality in the at least one blocked modality are restored, and the specific restoration order is not restricted here.
[0123] It can be understood that the description of restoring the samples corresponding to the modality here is only an example, and in actual applications, the specific order of restoring the samples corresponding to the modality is not restricted here.
[0124] Combined with the different blocking methods in the foregoing S102, the specific restoration means can also be different:
[0125] Case 1: Restoration corresponding to the deletion method
[0126] When the first sample is obtained by deleting the samples corresponding to at least one modality, add the samples corresponding to the first modality to the second sample to obtain the third sample, and the samples corresponding to the second modality are included in the preset samples.
[0127] Case 2: Restoration corresponding to the noise addition method;
[0128] When the feature vector of the first sample is the feature vector obtained by adding noise to the feature vector of the preset sample, remove the noise of the samples corresponding to the second modality to obtain the third sample, and the samples corresponding to the second modality are included in the preset samples.
[0129] Case 3: Restoration corresponding to the zeroing method.
[0130] When the feature vector of the first sample is the feature vector obtained by zeroing the feature vector of the preset sample, restore the feature vector of the samples corresponding to the second modality to its original value to obtain the third sample, and the samples corresponding to the second modality are included in the preset samples.
[0131] It can be understood that the description of the method for restoring the second modality here is only an example, and in actual applications, the restoration method should be selected in combination with the specific application scenario, and there is no restriction here.
[0132] S110. Process the third sample and the label of the third sample to obtain the third correct rate.
[0133] Exemplarily, input the third sample into a preset model to obtain a third label, calculate the probability that the third label is the same as the label of the third sample, and obtain the third correct rate.
[0134] S111. Determine whether the difference between the third correct rate and the first correct rate is less than or equal to a threshold.
[0135] Determine whether the difference between the third correct rate and the first correct rate is less than or equal to 50%. The description of the threshold here is only an example. In specific implementation, it can be adjusted according to the specific usage scenario and is not limited here.
[0136] If not, execute step S112.
[0137] If so, execute step S109.
[0138] S112. Take the first modality and the second modality as the target modalities, or take the second modality as the target modality.
[0139] Since in different application scenarios, the forms of the target modalities are different. For example, neither modality A nor modality B affects the prediction correct rate of the model, but when modality A and modality B appear simultaneously, it will have a greater impact on the prediction correct rate of the model. The target modality includes at least one modality; when multiple modalities are independent of each other, the target modality is a single modality.
[0140] Scenario 1:
[0141] In Scenario 1, the impact of modality A and modality B alone on the prediction correct rate of the preset model can be ignored, but when modality A and modality B appear simultaneously, it will have a greater impact on the prediction correct rate of the preset model. The target modality can include at least one modality.
[0142] In the stage of restoring the samples of at least one modality masked in the preset sample, after restoring the sample corresponding to the first modality, the change in the prediction correct rate of the preset model does not reach the threshold. In the case where after restoring the sample corresponding to the second modality, the change in the prediction correct rate of the preset model reaches the threshold, take the first modality and the second modality as the target modalities, that is, consider the first modality and the second modality as the key modalities of the preset model.
[0143] Based on the embodiments of the present application, the preset samples are all samples of a preset category, that is, the first modality and the second modality are the key modalities of the preset category of the preset model.
[0144] It can be understood that when restoring the samples corresponding to a modality, when the change in the prediction accuracy rate of the preset model reaches the threshold, the restoration of the samples corresponding to the modality stops. Therefore, when the samples corresponding to the Xth modality are restored and the change in the prediction accuracy rate of the preset model reaches the threshold, and when the samples corresponding to the (X - 1)th modality are restored and the change in the prediction accuracy rate of the preset model does not reach the threshold, the X modalities from the first modality to the Xth modality are used as the target modalities. Wherein, X is a positive integer and X is less than or equal to the number of modalities masked in the foregoing step S102.
[0145] In the embodiments of the present application, after restoring the first sample at least twice, the prediction accuracy rate of the preset model for the sample is greatly improved, and all the restored modalities are used as the target modalities. This method is applicable to scenarios where multiple modalities may interact with each other to affect the preset model. The accuracy of multi-modal importance analysis is improved.
[0146] Scenario 2:
[0147] In Scenario 2, multiple modalities are independent of each other, and the target modality is a single modality.
[0148] In the stage of restoring the samples of at least one modality masked in the preset sample, after restoring the samples corresponding to the first modality, the change in the prediction accuracy rate of the preset model does not reach the threshold. In the case where after restoring the samples corresponding to the second modality, the change in the prediction accuracy rate of the preset model reaches the threshold, the second modality is used as the target modality, that is, it is considered that the second modality is the key modality of the preset model.
[0149] Based on the embodiments of the present application, the preset samples are all samples of a preset category, that is, the second modality is the key modality of the preset category of the preset model.
[0150] It can be understood that when restoring the samples corresponding to a modality, when the change in the prediction accuracy rate of the preset model reaches the threshold, the restoration of the samples corresponding to the modality stops. Therefore, when the samples corresponding to the Xth modality are restored and the change in the prediction accuracy rate of the preset model reaches the threshold, and when the samples corresponding to the (X - 1)th modality are restored and the change in the prediction accuracy rate of the preset model does not reach the threshold, the Xth modality is used as the target modality. Wherein, X is a positive integer and X is less than or equal to the number of modalities masked in the foregoing step S102.
[0151] In the embodiments of the present application, after restoring the first sample at least twice, the prediction accuracy rate of the preset model for the sample is greatly improved, and the last restored modality is used as the target modality. This method is applicable to scenarios where multiple modalities independently affect the preset model respectively. The accuracy of multi-modal importance analysis is improved.
[0152] It can be understood that in step S101, when obtaining the preset samples and the labels of the preset samples, it is also possible to collect the preset samples, input the preset samples into the preset model, and obtain the labels of the preset samples. The preset samples include samples of multiple categories.
[0153] When the preset samples include samples of multiple categories, in step S112, the first modality and the second modality, or, the second modality is the key modality of the preset model.
[0154] In the embodiment of the present application, by obtaining a preset data set, the preset data set includes preset samples and the labels of the preset samples, and the preset samples include samples of at least two modalities. Mask at least one modality of samples from the preset samples to obtain the first samples. Combine the first samples and the labels of the first samples, calculate the correct rate of the preset model for classifying the first samples, and obtain the first correct rate. When the first correct rate is less than or equal to the threshold, it is considered that at least one modality includes the target modality. Restore the samples of the first modality in the first samples to obtain the second samples. Combine the second samples and the labels of the second samples, calculate the correctness of the preset model for classifying the second samples, and obtain the second correct rate. When the second correct rate meets the preset conditions, generate the target modality. The modal importance analysis of the model is completed through simple screening, and the modalities that are relatively important for the model are found.
[0155] For the sake of easy understanding, the following will be combined with Figure 5 Introduce a data processing method applied to the medical scenario. Assume that the hospital has N different test items for screening M diseases. The preset model analyzes the N different test items of patient A and obtains 2M classification results. Among them, if the occurrence probability of any one disease is greater than the preset threshold, it is considered that patient A has the disease. If the occurrence probability of any one disease is less than or equal to the preset threshold, it is considered that patient A is not a patient with the disease. Both M and N are positive integers.
[0156] Since a certain test item may be the basis for judging a certain specific disease, but it has no reference value for the judgment of other diseases. Therefore, when there is a certain prediction about the diseases that patient A may have, how to determine which test items patient A needs to perform to assist the judgment of a certain specific disease.
[0157] Combined with the solution provided by the present application, the test results of P patients (such as chest CT results, blood test results, fluoroscopy results, urine test results, etc.) and the screening results of each of the P patients for each of the M diseases can be collected. Taking the occurrence of the i-th disease among the M diseases as the screening classification condition, M classification results are obtained. The test items of the patients in each classification result are used as the preset samples, and the i-th disease is used as the label of the preset samples, where the preset samples include at least two test items.
[0158] For example, the test results of P patients may include: CT results, blood test results, fluoroscopy results, urine test results, etc.; the M diseases may include pneumonia, diabetes, pneumothorax, abnormal liver function, etc.; an input sample includes the test results of patient B in N types of tests such as CT results, blood test results, fluoroscopy results, urine test results, etc.; the label of an input sample includes the test results of patient B in M diseases such as pneumonia, diabetes, pneumothorax, abnormal liver function, etc., that is, whether the patient has M diseases such as pneumonia, diabetes, pneumothorax, abnormal liver function, etc., where patient B is included in P patients, and P is a positive integer.
[0159] In specific applications, the method provided in this application can encode P patients, N types of test results, and M diseases respectively to help distinguish different test results and diseases of different patients.
[0160] First, classify the analysis results of each of the P patients for the M diseases, that is, classify the labels of the preset samples. For example, the i-th type of disease corresponds to diabetes, and collect the N types of test results of the patients with diabetes to obtain p preset samples. Any two samples among the p preset samples have the same label. Here, any two samples in the preset samples have the label that the patient may have diabetes, where p is a positive integer.
[0161] Next, mask the samples of at least one modality in the preset samples to obtain a first sample. Combine the first sample and the label of the first sample, and calculate the accuracy rate of the preset model for classifying the first sample to obtain a first accuracy rate.
[0162] For example, mask the urine test result modality in the p preset samples to obtain a first sample, and input the first sample into the preset model for analysis again to obtain the label of the first sample. When the label of the first sample indicates that patient C is unlikely to have diabetes, it is considered that the preset model misclassifies patient C, and calculate the ratio of the number of labels indicating that the patient may have diabetes in the label of the first sample to p to obtain the first accuracy rate, and the first accuracy rate is less than or equal to 50%.
[0163] Next, when the first accuracy rate is less than or equal to the threshold, it is considered that at least one modality includes the target modality. Restore the samples of the first modality based on the first sample to obtain a second sample. Combine the second sample and the label of the second sample, and calculate the correctness of the preset model for classifying the second sample to obtain a second accuracy rate.
[0164] For example, restore the urine test result modality in p preset samples (that is, add the data of the urine test results of p preset samples to the first sample) to obtain a second sample. Input the second sample into the preset model for analysis again to obtain the label of the second sample. When it is indicated in the label of the second sample that patient C is unlikely to have diabetes, it is considered that the preset model misclassifies patient C, and calculate the ratio of the number of labels indicating that the patient may have diabetes in the label of the second sample to p, then the second correct rate can be obtained, and the difference between the second correct rate and the first correct rate is greater than 50%.
[0165] Finally, when the second correct rate meets the preset condition, generate the target modality. Wherein, the preset condition is that the difference between the second correct rate and the first correct rate is greater than the threshold (that is, the analysis correct rate of the first modality for the label has a significant impact), and the target modality includes the first modality. That is, obtain the test items that play a key role in the judgment of the i-th disease.
[0166] It should be noted that in Figure 5 the application scenario introduced, the test results of the patients and the samples of the disease types of the patients used should be materials authorized by the patients.
[0167] The data processing device in the present application will be described in detail below. Please refer to Figure 6 . Figure 6 FIG. 15 is a schematic diagram of an embodiment of the data processing device 10 in the embodiment of the present application. The data processing device 10 includes:
[0168] An acquisition unit 110, configured to acquire a preset sample and the label of the preset sample, and the preset sample includes samples of at least two modalities;
[0169] A processing unit 120, configured to process the first sample and the label of the first sample to obtain a first correct rate, where the first correct rate is the correct rate of the preset model for classifying the first sample, and the first sample is a sample obtained by masking at least one modality in the preset sample;
[0170] The processing unit 120 is further configured to, when the first correct rate is less than or equal to the threshold, process the second sample and the label of the second sample to obtain a second correct rate, where the second correct rate is the correct rate of the preset model for classifying the second sample, and the second sample is a sample obtained by restoring the first modality based on the first sample, the first modality is included in at least one modality, and the sample of the second modality is included in the preset sample;
[0171] A generation unit 130, configured to generate a target modality when the second correct rate meets the preset condition, and the target modality is included in at least one modality.
[0172] In the embodiments of the present application, by obtaining a preset data set, the preset data set includes preset samples and labels of the preset samples, and the preset samples include samples of at least two modalities. At least one modality of samples is masked from the preset samples to obtain first samples. Combining the first samples and the labels of the first samples, the correct rate of the preset model for classifying the first samples is calculated to obtain a first correct rate. When the first correct rate is less than or equal to a threshold, it is considered that at least one modality includes a target modality. The samples of the first modality in the first samples are restored to obtain second samples. Combining the second samples and the labels of the second samples, the correctness of the preset model for classifying the second samples is calculated to obtain a second correct rate. When the second correct rate meets the preset conditions, the target modality is generated. The modality importance analysis of the model is completed through simple screening, and the modalities that are relatively important for the model are found.
[0173] Optionally, the obtaining unit 110 is specifically configured to:
[0174] Obtain input samples, where the input samples include samples of at least two modalities;
[0175] Input the input samples into a preset model to obtain labels of the input samples;
[0176] Classify the input samples according to the labels of the input samples to obtain preset samples, and the labels of any two samples in the preset samples are the same.
[0177] In the embodiments of the present application, by collecting input samples, inputting the input samples into a preset model to obtain labels of the input samples, classifying the input samples according to the labels of the input samples to obtain preset samples and labels of the preset samples, where any two labels in the labels of the preset samples are the same, by screening the preset samples corresponding to the preset categories, and using the preset samples and the labels of the preset samples for multi-modal importance analysis to obtain the target modality, and the target modality is a modality that has a significant impact on the prediction correct rate of the preset category of the preset model.
[0178] Optionally, when the second correct rate is less than or equal to the threshold, the generating unit 130 is specifically configured to use the first modality as the target modality.
[0179] In the embodiments of the present application, after the first sample is restored once, the prediction accuracy of the preset model for the sample is greatly improved. The modality that has been restored is used as the target modality. Using the principle of the greedy algorithm, not all modalities are traversed, which reduces the time required for multi-modal importance analysis while ensuring the accuracy of multi-modal importance, and improves the efficiency of multi-modal importance analysis.
[0180] Optionally, the generating unit 130 is specifically configured to:
[0181] When the difference between the second correct rate and the first correct rate is less than or equal to the threshold, process the third sample and the label of the third sample to obtain the third correct rate, where the third correct rate is the correct rate of the preset model for classifying the third sample. The third sample is a sample obtained by restoring the second modality based on the second sample. The second modality is included in at least one modality, and the samples of the third modality are included in the preset sample;
[0182] When the difference between the third correct rate and the first correct rate is greater than the threshold, take the first modality and the second modality as the target modalities.
[0183] In the embodiments of the present application, after restoring the first sample at least twice, the prediction accuracy of the preset model for the sample is greatly improved. Take all the restored modalities as the target modalities. This method is applicable to scenarios where multiple modalities may interact with each other and affect the preset model. It improves the accuracy of multi-modal importance analysis.
[0184] Optionally, the generating unit 130 is specifically configured to:
[0185] When the difference between the second correct rate and the first correct rate is less than or equal to the threshold, process the third sample and the label of the third sample to obtain the third correct rate, where the third correct rate is the correct rate of the preset model for classifying the third sample. The third sample is a sample obtained by restoring the second modality based on the second sample. The second modality is included in at least one modality, and the samples of the third modality are included in the preset sample;
[0186] When the difference between the third correct rate and the first correct rate is greater than the threshold, take the second modality as the target modality.
[0187] In the embodiments of the present application, after restoring the first sample at least twice, the prediction accuracy of the preset model for the sample is greatly improved. Take the last restored modality as the target modality. This method is applicable to scenarios where multiple modalities independently affect the preset model respectively. It improves the accuracy of multi-modal importance analysis.
[0188] Optionally, the processing unit 120 is further configured to process the preset sample to obtain the first sample.
[0189] In the embodiments of the present application, at least one modality corresponding to the samples in the preset sample can be masked by the deletion method, the noise addition method, and the zeroing method, which improves the flexibility of the scheme implementation.
[0190] Optionally, the processing unit 120 is specifically configured to:
[0191] Extract the features of the preset sample to obtain the feature matrix of the preset sample;
[0192] Set the feature matrix corresponding to at least one modality in the feature matrix of the preset sample to 0 to obtain the feature matrix of the first sample.
[0193] In the embodiments of the present application, feature extraction is performed on a preset sample to obtain a feature matrix of the preset sample. In the feature matrix of the preset sample, the feature matrix corresponding to the sample of any modality is set to zero to obtain the feature matrix of the first sample, which efficiently realizes the masking of the sample of any modality in the preset sample and improves the feasibility of the solution.
[0194] Figure 7 FIG. 5 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 200 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 222 (for example, one or more processors) and a memory 232, and one or more storage media 230 (for example, one or more mass storage devices) for storing application programs 242 or data 244. Among them, the memory 332 and the storage media 230 may be transient storage or persistent storage. The program stored in the storage media 230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 222 may be configured to communicate with the storage media 230 and execute a series of instruction operations in the storage media 230 on the server 200.
[0195] The server 200 may further include one or more power supplies 226, one or more wired or wireless network interfaces 250, one or more input / output interfaces 258, and / or one or more operating systems 241, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0196] In the above embodiments, the steps executed by the server may be based on the Figure 7 server structure shown.
[0197] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0198] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0199] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0202] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0203] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for data processing, characterized in that, Including: Obtain a preset sample and the label of the preset sample, where the preset sample contains samples of at least two modalities; Process the first sample and the label of the first sample to obtain the first accuracy rate, where the first accuracy rate is the accuracy rate of the preset model classifying the first sample, and the first sample is a sample obtained by masking at least one modality in the preset sample; When the first accuracy rate is less than or equal to a threshold, process the second sample and the label of the second sample to obtain the second accuracy rate, where the second accuracy rate is the accuracy rate of the preset model classifying the second sample, the second sample is a sample obtained by restoring the first modality based on the first sample, the first modality is included in the at least one modality, and the sample of the second modality is included in the preset sample; When the second accuracy rate meets a preset condition, generate a target modality, where the target modality is included in the at least one modality.
2. The method according to claim 1, wherein The obtaining the preset sample and the label of the preset sample includes: Obtain an input sample, where the input sample contains samples of at least two modalities; Input the input sample into the preset model to obtain the label of the input sample; Classify the input sample according to the label of the input sample to obtain the preset sample, where each sample in the preset sample contains a preset label.
3. The method according to claim 1 or 2, characterized in that When the second accuracy rate meets a preset condition, generating the target modality includes: When the second accuracy rate is less than or equal to the threshold, use the first modality as the target modality.
4. The method according to claim 1 or 2, characterized in that, When the second accuracy rate meets a preset condition, generating the target modality includes: When the difference between the second accuracy rate and the first accuracy rate is less than or equal to the threshold, process the third sample and the label of the third sample to obtain the third accuracy rate, where the third accuracy rate is the accuracy rate of the preset model classifying the third sample, the third sample is a sample obtained by restoring the second modality based on the second sample, the second modality is included in the at least one modality, and the sample of the third modality is included in the preset sample; When the difference between the third accuracy rate and the first accuracy rate is greater than the threshold, use the first modality and the second modality as the target modalities.
5. The method according to claim 1 or 2, characterized in that, When the second accuracy rate meets a preset condition, generating the target modality includes: When the difference between the second accuracy rate and the first accuracy rate is less than or equal to the threshold, process the third sample and the label of the third sample to obtain the third accuracy rate, where the third accuracy rate is the accuracy rate of the preset model classifying the third sample, the third sample is a sample obtained by restoring the second modality based on the second sample, the second modality is included in the at least one modality, and the sample of the third modality is included in the preset sample; When the difference between the third accuracy rate and the first accuracy rate is greater than the threshold, use the second modality as the target modality.
6. The method according to any one of claims 1 to 3, characterized in that, Before processing the first sample and the label of the first sample to obtain the first accuracy rate, the method further includes: Process the preset sample to obtain the first sample.
7. The method according to claim 6, wherein Processing the preset sample to obtain the first sample includes: Performing feature extraction on the preset sample to obtain a feature matrix of the preset sample; Setting the feature matrix corresponding to at least one modality in the feature matrix of the preset sample to 0 to obtain the feature matrix of the first sample.
8. A data processing device, characterized in that, Includes: An acquisition unit for acquiring a preset sample and a label of the preset sample, where the preset sample includes samples of at least two modalities; A processing unit for processing the first sample and the label of the first sample to obtain the first accuracy rate, where the first accuracy rate is the accuracy rate of the preset model for classifying the first sample, and the first sample is a sample obtained by masking at least one modality in the preset sample; The processing unit is further configured to, when the first accuracy rate is less than or equal to a threshold, process the second sample and the label of the second sample to obtain the second accuracy rate, where the second accuracy rate is the accuracy rate of the preset model for classifying the second sample, and the second sample is a sample obtained by restoring the first modality based on the first sample, the first modality is included in the at least one modality, and the sample of the second modality is included in the preset sample; A generation unit for generating a target modality when the second accuracy rate meets a preset condition, where the target modality is included in the at least one modality.
9. A computer device, characterized in that, Includes: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the programs in the memory, including executing the data processing method according to any one of claims 1 to 7; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
10. A computer-readable storage medium includes instructions that, when running on a computer, cause the computer to execute the data processing method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, The computer program is executed by the processor to perform the data processing method according to any one of claims 1 to 7.