Training Method, Device, Equipment and Storage Medium of AI Model
By using multiple sample data and the output data of the previous AI model in AI model training for feature extraction and randomly inactivating the hidden state data, the problems of insufficient generalization ability and low resource utilization efficiency of the AI model in the absence of information are solved, and stronger generalization ability and higher resource utilization efficiency are achieved.
Patent Information
- Application Number
- CN202210044173.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing AI models have weak generalization capabilities when information is missing, and the training of each AI model requires collecting a large amount of repetitive user information, which makes resource utilization efficiency low.
By obtaining multiple types of user information, a variety of sample data are generated. If the AI model to be trained is the non-first AI model among multiple continuity AI models, the sample data and the output data of the previous AI model are input to the current AI model for feature extraction, obtain hidden state data, and randomly inactivate the hidden state data, and finally the model is trained based on the updated hidden state data.
The generalization ability of the AI model is improved, so that it can be applied to situations where information is missing, and the resource utilization efficiency is improved by reusing user information multiple times.
Smart Images

Figure CN114398985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning in artificial intelligence, and particularly to a method, device, equipment and storage medium for training an AI model. Background Art
[0002] At present, patient diagnosis and treatment are mainly carried out offline. Even in the Internet era, the part that can assist in diagnosis and treatment online is very small. In recent years, more and more Internet hospitals have emerged. Internet hospitals can provide convenient and fast online AI (Artificial Intelligence) diagnosis and treatment services for patients. For different types of diagnosis and treatment services, corresponding independent AI models are constructed and trained, and online AI diagnosis and treatment services are realized based on various trained AI models. However, the trained AI models can only be used to provide online AI diagnosis and treatment services for patients with complete information, and for some patients with missing information, accurate and reliable online AI diagnosis and treatment services cannot be realized through the AI models, that is, the generalization ability of the AI models is weak. Moreover, online AI diagnosis and treatment services are all independent. For the training of AI models corresponding to each service, it is often necessary to collect various information such as patients' symptoms, allergy history, disease history, diagnosis, etc. That is, the information collected each time is often only used for the training of a certain AI model, and the resource utilization efficiency is low.
[0003] Therefore, how to improve the generalization ability of the AI model and improve the resource utilization efficiency has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a method, device, equipment and storage medium for training an AI model, aiming to improve the generalization ability of the AI model and improve the resource utilization efficiency.
[0005] To achieve the above object, this application provides a method for training an AI model, and the method for training the AI model includes:
[0006] Obtain various types of user information and generate various sample data;
[0007] If the current AI model to be trained is a non-first AI model among multiple successive AI models, then input the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data;
[0008] Perform random inactivation processing on the hidden state data to obtain updated hidden state data;
[0009] Based on the updated hidden state data, train the current AI model to obtain the trained current AI model.
[0010] In addition, to achieve the above object, the present application further provides a training device for an AI model, and the training device for the AI model includes:
[0011] A sample data generation module, configured to obtain various types of user information and generate various sample data;
[0012] A feature extraction module, configured to, if the current AI model to be trained is a non-first AI model among multiple AI models with a succession relationship, input the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data;
[0013] A data inactivation processing module, configured to perform random inactivation processing on the hidden state data to obtain updated hidden state data;
[0014] A model training module, configured to perform model training on the current AI model based on the updated hidden state data to obtain a trained current AI model.
[0015] In addition, to achieve the above object, the present application further provides a computer device, and the computer device includes a memory and a processor;
[0016] The memory is configured to store a computer program;
[0017] The processor is configured to execute the computer program and, when executing the computer program, implement the training method of the AI model as described above.
[0018] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the training method of the AI model as described above are implemented.
[0019] The present application discloses a training method, device, equipment and storage medium for an AI model. By obtaining various types of user information and generating various sample data, if the current AI model to be trained is a non-first AI model among multiple successive AI models, then the various sample data and the output data of at least one AI model preceding the current AI model are input into the current AI model for feature extraction to obtain corresponding hidden state data, and the hidden state data is subjected to dropout processing to obtain updated hidden state data. Based on the updated hidden state data, the current AI model is trained to obtain the trained current AI model, realizing the training of the AI model using data with missing information. Therefore, the trained AI model can be applicable to the situation of missing information, that is, the generalization ability of the AI model is improved. Moreover, based on the various sample data, multiple successive AI models are trained respectively, that is, the user information is reused multiple times, so the resource utilization efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0021] Figure 1 It is a schematic flowchart of the steps of a training method for an AI model provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic flowchart of the steps of inputting various sample data into the current AI model for feature extraction to obtain corresponding hidden state data;
[0023] Figure 3 It is a schematic flowchart of the steps of performing dropout processing on the hidden state data;
[0024] Figure 4 It is a schematic block diagram of a training device for an AI model provided by an embodiment of the present application;
[0025] Figure 5 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0027] The flowcharts shown in the accompanying drawings are only illustrative, not necessarily including all the contents and operations / steps, nor necessarily executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may change according to the actual situation.
[0028] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0029] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0030] The embodiments of the present application provide a training method, device, device and storage medium for an AI model, which are used to realize the monitoring of web pages at the user-perceivable level.
[0031] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a training method for an AI model provided by an embodiment of the present application. This method can be applied to a computer device, and the application scenario of this method is not limited in the present application. Next, taking the application of this training method for the AI model to a computer device as an example, this training method for the AI model will be introduced in detail.
[0032] As Figure 1 shown, the training method of this AI model specifically includes steps S101 to S104.
[0033] S101. Obtain various types of user information and generate various sample data.
[0034] In order to train the AI model, a large amount of user information is obtained. Exemplarily, the user information includes various types such as structured information, text information, and image information.
[0035] Exemplarily, taking the case where an AI model is used for online diagnosis and treatment services as an example, the user information obtained is patient information, which includes structured information, text information, image information, etc. Among them, the structured information includes demographic characteristics, examination and inspection indicators, vital signs, medication history, disease history, new disease conditions within a certain period (such as within 3 years), the diagnosis of the current consultation, the prescriptions for the current consultation, etc. The text information includes the chief complaint and the Q&A dialogue with the doctor, etc. The image information includes images of the affected area, etc.
[0036] In some embodiments, generating a variety of sample data may include:
[0037] Performing data preprocessing on the user information to generate a variety of the sample data, where the data preprocessing includes at least one of data classification and data cleaning.
[0038] Still taking the user information as patient information as an example, performing data preprocessing on the patient information, where the data preprocessing includes at least one of, but is not limited to, data classification and data cleaning. For example, classifying patient information such as vital signs, medication history, disease history, new disease conditions within a certain period, the diagnosis of the current consultation, the prescriptions for the current consultation, the Q&A dialogue with the doctor, and images of the affected area into structured information, text information, image information, etc. And performing data cleaning to remove the invalid patient information among them.
[0039] By performing data preprocessing on various user information, sample data corresponding to various types of user information is obtained.
[0040] For example, if the user information includes multiple types such as structured information, text information, and image information, the corresponding sample data obtained includes structured feature data, text feature data, image feature data, etc.
[0041] S102. If the current AI model to be trained is a non-first AI model among multiple consecutive AI models, then input the variety of the sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain the corresponding hidden state data.
[0042] Exemplarily, for multiple consecutive AI models, each AI model among them is trained in sequence. For example, assume that multiple consecutive AI models include a first AI model, a second AI model, and a third AI model, where the second AI model is preceded by the first AI model and followed by the third AI model. First, the first AI model is trained, then the second AI model is trained, and finally the third AI model is trained.
[0043] In some embodiments, before inputting the multiple types of sample data and the output data of at least one AI model that precedes the current AI model into the current AI model for feature extraction when the current AI model to be trained is a non-first AI model among multiple successive AI models, it may include:
[0044] Determine whether the current AI model is the first AI model among multiple successive AI models;
[0045] If not, obtain the output data of at least one AI model that precedes the current AI model.
[0046] Exemplarily, if the current AI model is the first AI model among multiple successive AI models, input the multiple types of sample data into the current AI model for feature extraction to obtain corresponding hidden state data.
[0047] For example, still taking the above-listed example, if the current AI model is the first AI model as described above, and the first AI model is the first AI model among multiple successive AI models, directly input the obtained multiple types of sample data into the current AI model for feature extraction to obtain corresponding hidden state data, so as to train the current AI model based on the hidden state data.
[0048] Exemplarily, the sample data includes image feature data A_i(1), structured feature data A_i(2), and text feature data A_i(3). Input the image feature data A_i(1) into the current AI model for feature extraction to obtain the hidden state data h_i(1) corresponding to the image feature data; input the structured feature data A_i(2) into the current AI model for feature extraction to obtain the hidden state data h_i(2) corresponding to the structured feature data; input the text feature data A_i(3) into the current AI model for feature extraction to obtain the hidden state data h_i(3) corresponding to the text feature data.
[0049] Exemplarily, the obtained multiple types of hidden state data h_i(1), h_i(2), h_i(3) form a hidden state data set {h_i(1), h_i(2), h_i(3)}.
[0050] Exemplarily, each AI model includes a CNN (Convolutional Neural Network) layer, an FCNN (Full Convolutional Neural Network) layer, and a Text CNN (Text Convolutional Neural Network) layer. If the current AI model is the first AI model among multiple successive AI models, for example, as Figure 2 shown, the image feature data A_i(1) is input into the CNN layer of the current AI model for feature extraction to obtain the hidden state data h_i(1) corresponding to the image feature data A_i(1); the structured feature data A_i(2) is input into the FCNN layer of the current AI model for feature extraction to obtain the hidden state data h_i(2) corresponding to the structured feature data A_i(2); the text feature data A_i(3) is input into the Text CNN layer of the current AI model for feature extraction to obtain the hidden state data h_i(3) corresponding to the text feature data A_i(3). The hidden state data h_i(1), the hidden state data h_i(2), and the hidden state data h_i(3) form a hidden state dataset {h_i(1), h_i(2), h_i(3)}.
[0051] Exemplarily, if the current AI model is not the first AI model among multiple successive AI models, then the output data of at least one AI model preceding the current AI model is obtained. For example, if the current AI model is the second AI model, then the output data of the preceding first AI model corresponding to the input multiple sample data is obtained.
[0052] After that, the multiple sample data and the output data of at least one AI model preceding the current AI model are input into the current AI model for feature extraction to obtain the corresponding hidden state data.
[0053] In some embodiments, inputting the multiple sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain the corresponding hidden state data may include:
[0054] Inputting the image feature data into the CNN layer of the current AI model for feature extraction to obtain the corresponding first hidden state data;
[0055] Inputting the structured feature data and the output data into the FCNN layer of the current AI model for feature extraction to obtain the corresponding second hidden state data;
[0056] Input the text feature data into the Text CNN layer of the current AI model for feature extraction to obtain corresponding third hidden state data.
[0057] For example, still taking the above-listed example, if the current AI model is the second AI model, input the image feature data into the CNN layer of the second AI model for feature extraction to obtain corresponding first hidden state data; input the structured feature data and the output data of the first AI model into the FCNN layer of the second AI model for feature extraction to obtain corresponding second hidden state data; input the text feature data into the Text CNN layer of the second AI model for feature extraction to obtain corresponding third hidden state data.
[0058] S103. Perform dropout processing on the hidden state data to obtain updated hidden state data.
[0059] Exemplarily, still taking the above-listed hidden state data h_i(1), h_i(2), h_i(3) as an example, update the hidden state data by performing dropout processing on at least one of h_i(1), h_i(2), h_i(3).
[0060] In some embodiments, as Figure 3 shown, step S103 may include sub-step S1031 and sub-step S1032.
[0061] S1031. Randomly select at least one hidden state data from multiple types of the hidden state data.
[0062] For example, still taking the hidden state data h_i(1), h_i(2), h_i(3) as an example, randomly select at least one hidden state data from them. For example, randomly select one of h_i(1), h_i(2), h_i(3).
[0063] In some embodiments, each hidden state data corresponds to a unique number. For example, the hidden state data h_i(1) corresponds to number 1, the hidden state data h_i(2) corresponds to number 2, and the hidden state data h_i(3) corresponds to number 3. Randomly selecting at least one hidden state data from multiple types of the hidden state data may include:
[0064] Randomly generate at least one number, and select the hidden state data corresponding to each number from multiple types of the hidden state data.
[0065] For example, still taking the hidden state data h_i(1), h_i(2), and h_i(3) as an example, the hidden state data h_i(1), h_i(2), and h_i(3) correspond to numbers 1, 2, and 3 respectively. At least one number t=Random(1, 2, 3) is randomly generated from 1, 2, and 3, and the hidden state data corresponding to the at least one randomly generated number is selected.
[0066] For example, if the number 1 is randomly generated, the hidden state data h_i(1) corresponding to number 1 is selected; for another example, if the numbers 1 and 2 are randomly generated, the hidden state data h_i(1) and h_i(2) corresponding to numbers 1 and 2 are selected.
[0067] S1032. Replace at least one of the hidden state data with a preset value to inactivate at least one of the hidden state data.
[0068] The preset value may be set to 0. It is understandable that the preset value may also be other values other than 0.
[0069] If the preset value is 0, the hidden state data h_i(t) corresponding to the randomly generated number t=Random(1,2,3) is updated to: h_i(t)=0. For example, when the randomly generated number t=Random(1,2,3) is 1, h_i(1)=0 is set, so that the hidden state data h_i(1) is inactivated.
[0070] S104: Perform model training on the current AI model based on the updated hidden state data to obtain a trained current AI model.
[0071] For example, still taking the hidden state data h_i(1), h_i(2), and h_i(3) as an example, if the randomly generated number t=Random(1, 2, 3) is 1, set h_i(1)=0 to obtain the updated hidden state data h_i(1)=0, as well as h_i(2) and h_i(3). The current AI model is trained based on the updated hidden state data until the current AI model converges to obtain the trained current AI model.
[0072] Exemplarily, a binary cross entropy loss function is used in the training process of the current AI model until the loss function no longer decreases during training, that is, the current AI model converges.
[0073] Since multiple AI models are connected, various types of user information (such as patient information) are applied to the training of multiple AI models without the need for repeated collection and acquisition, thereby improving the efficiency of information acquisition and utilization.
[0074] Exemplarily, in order to provide comprehensive online AI diagnosis and treatment services such as risk assessment -> diagnosis -> treatment recommendation, three corresponding AI models with connectivity are constructed: a risk assessment model, a disease diagnosis model, and a treatment recommendation model.
[0075] Generate corresponding multiple sample data based on multiple types of patient information, where the multiple sample data includes image feature data, structured feature data, and text feature data.
[0076] First, input the image feature data into the CNN layer of the risk assessment model for feature extraction to obtain the hidden state data corresponding to the image feature data, input the structured feature data into the FCNN layer of the risk assessment model for feature extraction to obtain the hidden state data corresponding to the structured feature data, input the text feature data into the Text CNN layer of the risk assessment model for feature extraction to obtain the hidden state data corresponding to the text feature data, perform dropout processing on the obtained multiple hidden state data, and based on the multiple hidden state data after dropout processing, train the risk assessment model to obtain the trained risk assessment model and obtain the first output data of the risk assessment model.
[0077] After that, input the image feature data into the CNN layer of the disease diagnosis model for feature extraction to obtain the hidden state data corresponding to the image feature data, input the structured feature data and the first output data into the FCNN layer of the disease diagnosis model for feature extraction to obtain the corresponding hidden state data, input the text feature data into the Text CNN layer of the disease diagnosis model for feature extraction to obtain the hidden state data corresponding to the text feature data, perform dropout processing on the obtained multiple hidden state data, and based on the multiple hidden state data after dropout processing, train the disease diagnosis model to obtain the trained disease diagnosis model and obtain the second output data of the disease diagnosis model.
[0078] Then, input the image feature data into the CNN layer of the treatment recommendation model for feature extraction to obtain the hidden state data corresponding to the image feature data, input the structured feature data, the first output data, and the second output data into the FCNN layer of the treatment recommendation model for feature extraction to obtain the corresponding hidden state data, input the text feature data into the Text CNN layer of the treatment recommendation model for feature extraction to obtain the hidden state data corresponding to the text feature data, perform dropout processing on the obtained multiple hidden state data, and based on the multiple hidden state data after dropout processing, train the treatment recommendation model to obtain the trained treatment recommendation model.
[0079] After training the risk assessment model, disease diagnosis model, and treatment recommendation model, during the model usage stage, for patients coming for online consultations, first, collect the patient information of the patients. The patient information includes structured information, text information, image information, etc. Among them, the structured information includes demographic characteristics, test and examination indicators, vital signs, medication history, disease history, etc. The text information includes the chief complaint and the questions and answers during the doctor's consultation. The image information includes images of the affected area, etc. Then, use the risk assessment model to prompt the patient's disease risk; subsequently, use the disease diagnosis model to accurately diagnose the patient; finally, use the treatment recommendation model to recommend treatment for the patient. For patients lacking a certain type of information among various types of patient information such as structured information, text information, and image information, due to the strong generalization ability of the model, it is also applicable to the situation of missing information, and accurate and reliable online AI diagnosis and treatment services can also be obtained through the model.
[0080] In the above embodiments, by obtaining various types of user information and generating various sample data, if the current AI model to be trained is a non-first AI model among multiple successive AI models, then input the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain the corresponding hidden state data, and perform a random inactivation process on the hidden state data to obtain the updated hidden state data. Based on the updated hidden state data, train the current AI model to obtain the trained current AI model, realizing the training of the AI model using data with missing information. Therefore, the trained AI model can be applicable to the situation of missing information, that is, the generalization ability of the AI model is improved. And, based on the various sample data, train multiple successive AI models respectively, that is, the user information is reused multiple times. Therefore, the resource utilization efficiency is improved.
[0081] Please refer to Figure 4 , Figure 4 FIG. is a schematic block diagram of a training device for an AI model provided by an embodiment of the present application. The training device for the AI model can be configured in a computer device and is used to execute the aforementioned AI model training method.
[0082] As Figure 4 shown, the training device 1000 for the AI model includes: a sample data generation module 1001, a feature extraction module 1002, a data inactivation processing module 1003, and a model training module 1004.
[0083] The sample data generation module 1001 is used to obtain various types of user information and generate various sample data;
[0084] A feature extraction module 1002, which is configured to, if the current AI model to be trained is a non-first AI model among multiple AI models with a sequential relationship, input the multiple types of sample data and the output data of at least one AI model that precedes the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data;
[0085] A data inactivation processing module 1003, which is configured to perform a random inactivation process on the hidden state data to obtain updated hidden state data;
[0086] A model training module 1004, which is configured to perform model training on the current AI model based on the updated hidden state data to obtain a trained current AI model.
[0087] In one embodiment, the multiple types of sample data include image feature data, structured feature data, and text feature data; the feature extraction module 1002 is further configured to:
[0088] Input the image feature data into the CNN layer of the current AI model for feature extraction to obtain corresponding first hidden state data;
[0089] Input the structured feature data and the output data into the FCNN layer of the current AI model for feature extraction to obtain corresponding second hidden state data;
[0090] Input the text feature data into the Text CNN layer of the current AI model for feature extraction to obtain corresponding third hidden state data.
[0091] In one embodiment, the data inactivation processing module 1003 is further configured to:
[0092] Randomly select at least one type of hidden state data from the multiple types of hidden state data;
[0093] Replace at least one type of hidden state data with a preset value to inactivate at least one type of hidden state data.
[0094] In one embodiment, each type of hidden state data corresponds to a unique number, and the data inactivation processing module 1003 is further configured to:
[0095] Randomly generate at least one number, and select the hidden state data corresponding to each number from the multiple types of hidden state data.
[0096] In one embodiment, the training device for the AI model further includes an acquisition module, which is configured to:
[0097] Determine whether the current AI model is the first AI model among multiple AI models with a sequential relationship;
[0098] If not, obtain the output data of at least one AI model that precedes the current AI model.
[0099] In one embodiment, the feature extraction module 1002 is further configured to:
[0100] If the current AI model is the first AI model among multiple successive AI models, input the multiple types of sample data into the current AI model for feature extraction to obtain corresponding hidden state data.
[0101] In one embodiment, the sample data generation module 1001 is further configured to:
[0102] Perform data preprocessing on the user information to generate the multiple types of sample data, where the data preprocessing includes at least one of data classification and data cleaning.
[0103] Among them, each module in the above AI model training device 1000 corresponds to each step in the above AI model training method embodiment, and its functions and implementation processes will not be elaborated here one by one.
[0104] The methods and devices of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0105] Exemplarily, the above methods and devices can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 5 shown.
[0106] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.
[0107] Please refer to Figure 5, the computer device includes a processor and a memory connected by a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0108] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0109] The internal memory provides an environment for the operation of computer programs in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be made to execute any one of the AI model training methods.
[0110] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0111] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:
[0112] Obtain various types of user information and generate various sample data;
[0113] If the current AI model to be trained is a non-first AI model among multiple successive AI models, then input the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data;
[0114] Perform dropout processing on the hidden state data to obtain updated hidden state data;
[0115] Based on the updated hidden state data, train the current AI model to obtain a trained current AI model.
[0116] In one embodiment, the various sample data include image feature data, structured feature data, and text feature data;
[0117] When the processor implements inputting the multiple sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data, it is used to implement:
[0118] Input the image feature data into the CNN layer of the current AI model for feature extraction to obtain corresponding first hidden state data;
[0119] Input the structured feature data and the output data into the FCNN layer of the current AI model for feature extraction to obtain corresponding second hidden state data;
[0120] Input the text feature data into the Text CNN layer of the current AI model for feature extraction to obtain corresponding third hidden state data.
[0121] In one embodiment, when the processor implements performing dropout processing on the hidden state data, it is used to implement:
[0122] Randomly select at least one hidden state data from the multiple hidden state data;
[0123] Replace at least one hidden state data with a preset value to inactivate at least one hidden state data.
[0124] In one embodiment, each hidden state data corresponds to a unique number. When the processor implements randomly selecting at least one hidden state data from the multiple hidden state data, it is used to implement:
[0125] Randomly generate at least one number, and select the hidden state data corresponding to each number from the multiple hidden state data.
[0126] In one embodiment, before the processor implements, if the current AI model to be trained is a non-first AI model among multiple successive AI models, inputting the multiple sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction, it is used to implement:
[0127] Determine whether the current AI model is the first AI model among multiple successive AI models;
[0128] If not, obtain the output data of at least one AI model preceding the current AI model.
[0129] In one embodiment, after the processor implements determining whether the current AI model is the first AI model among multiple successive AI models, it is used to implement:
[0130] If the current AI model is the first AI model among multiple successive AI models, then multiple types of the sample data are input into the current AI model for feature extraction to obtain corresponding hidden state data.
[0131] In one embodiment, when the processor implements the generation of multiple types of sample data, it is used to implement:
[0132] Perform data preprocessing on the user information to generate multiple types of the sample data, where the data preprocessing includes at least one of data classification and data cleaning.
[0133] The embodiment of the present application further provides a computer-readable storage medium.
[0134] A computer program is stored on the computer-readable storage medium of the present application, and when the computer program is executed by a processor, the steps of the training method of the AI model as described above are implemented.
[0135] Among them, the computer-readable storage medium may be an internal storage unit of the training device or computer device of the AI model described in the foregoing embodiment, such as the hard disk or memory of the training device or computer device of the AI model. The computer-readable storage medium may also be an external storage device of the training device or computer device of the AI model, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital Card (SD Card), a Flash Card, etc. equipped on the training device or computer device of the AI model.
[0136] Furthermore, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.
[0137] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0138] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or system including a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including such element.
[0139] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application.
Claims
1. A training method for an AI model, characterized in that, the training method for the AI model includes: obtaining various types of user information and generating various sample data; the various sample data include image feature data, structured feature data, and text feature data; if the current AI model to be trained is a non-first AI model among multiple successive AI models, then inputting the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data; performing dropout processing on the hidden state data to obtain updated hidden state data; training the current AI model based on the updated hidden state data to obtain a trained current AI model; the step of inputting the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data includes: inputting the image feature data into the CNN layer of the current AI model for feature extraction to obtain corresponding first hidden state data; inputting the structured feature data and the output data into the FCNN layer of the current AI model for feature extraction to obtain corresponding second hidden state data; inputting the text feature data into the Text CNN layer of the current AI model for feature extraction to obtain corresponding third hidden state data.
2. The training method for an AI model according to claim 1, characterized in that, the step of performing dropout processing on the hidden state data includes: randomly selecting at least one hidden state data from the various hidden state data; replacing at least one of the hidden state data with a preset value to inactivate at least one of the hidden state data.
3. The training method for an AI model according to claim 2, characterized in that, each hidden state data corresponds to a unique number, and the step of randomly selecting at least one hidden state data from the various hidden state data includes: randomly generating at least one number and selecting the hidden state data corresponding to each number from the various hidden state data.
4. The training method for an AI model according to claim 1, characterized in that, before the step of, if the current AI model to be trained is a non-first AI model among multiple successive AI models, then inputting the various sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction, includes: determining whether the current AI model is the first AI model among multiple successive AI models; if not, then obtaining the output data of at least one AI model preceding the current AI model.
5. The training method for an AI model according to claim 4, characterized in that, after the step of determining whether the current AI model is the first AI model among multiple successive AI models, includes: If the current AI model is the first AI model among multiple successive AI models, then input multiple types of the sample data into the current AI model for feature extraction to obtain corresponding hidden state data.
6. The method for training an AI model according to any one of claims 1 to 5, wherein, the generating multiple types of sample data includes: performing data preprocessing on the user information to generate multiple types of the sample data, wherein the data preprocessing includes at least one of data classification and data cleaning.
7. An apparatus for training an AI model, wherein, the apparatus for training an AI model includes: a sample data generation module, configured to obtain multiple types of user information and generate multiple types of sample data; the multiple types of sample data include image feature data, structured feature data, and text feature data; a feature extraction module, configured to, if the current AI model to be trained is not the first AI model among multiple successive AI models, input the multiple types of the sample data and the output data of at least one AI model preceding the current AI model into the current AI model for feature extraction to obtain corresponding hidden state data; a data inactivation processing module, configured to perform random inactivation processing on the hidden state data to obtain updated hidden state data; a model training module, configured to perform model training on the current AI model based on the updated hidden state data to obtain a trained current AI model; the feature extraction module is further configured to input the image feature data into the CNN layer of the current AI model for feature extraction to obtain corresponding first hidden state data; input the structured feature data and the output data into the FCNN layer of the current AI model for feature extraction to obtain corresponding second hidden state data; and input the text feature data into the Text CNN layer of the current AI model for feature extraction to obtain corresponding third hidden state data.
8. A computer device, wherein, the computer device includes a memory and a processor; the memory is configured to store a computer program; the processor is configured to execute the computer program and, when executing the computer program, implement the method for training an AI model according to any one of claims 1 to 6.
9. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for training an AI model according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model training method and device and electronic equipment
CN111523663A
Multi-modal model training method, device and equipment and storage medium
CN112464993A