A method and system for analyzing medical evaluations using an adversarial network through knowledge distillation
Medical evaluation is analyzed through the adversarial network of knowledge distillation, combined with the MASH module of the generation adversarial network and transformer model, training samples are generated, and knowledge distillation is performed through the GRU model, BERT model and LSTM network, and student models are refined and trained to analyze medical evaluation data. The dependence on a large amount of high-quality data and time in the existing technology is solved, and efficient and accurate medical evaluation analysis is achieved.
Patent Information
- Application Number
- CN202510221222.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art relies on huge language models in the emotional classification of medical evaluation, requires a large amount of high-quality data and time, and is difficult to implement under real-world restrictions, affecting the development of medical services.
Medical evaluation is analyzed through the adversarial network of knowledge distillation, combined with the MASH module of the generation adversarial network and transformer model, training samples are generated, and knowledge distillation is performed through the GRU model, BERT model and LSTM network, and student models are refined and trained to analyze medical evaluation data.
It reduces the demand for a large amount of high-quality data and time, improves the efficiency and accuracy of medical evaluation data analysis, and promotes the development of medical services.
Smart Images

Figure CN119724522B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and system for analyzing medical evaluations through adversarial networks with knowledge distillation. Background Art
[0002] Medical evaluations of medical institutions, such as: treatment quality, facility evaluations, and patient experiences, help patients make informed choices when selecting medical institutions, and at the same time help medical institutions improve their medical services. An artificial intelligence-driven system can quickly collect medical evaluations, and then use NLP for analysis, extract the attitudes of medical evaluations, and classify the treatment quality, facility evaluations, and patient experiences according to the attitudes. Among them, sentiment classification is an important issue in NLP, which attempts to classify text into sentiment levels, and this is crucial for evaluating medical evaluations in various situations. However, currently, sentiment classification largely relies on large language models, requires a large amount of high-quality data and a large amount of time, and is difficult to implement under the limited conditions of the real world, affecting the development of medical services. Summary of the Invention
[0003] The technical problem to be solved by the present invention is: The present invention provides a method and system for analyzing medical evaluations through adversarial networks with knowledge distillation, which can analyze medical evaluation data while reducing the need for a large amount of high-quality data and a large amount of time.
[0004] To solve the above technical problem, the technical solution adopted by the present invention is:
[0005] In a first aspect, the present invention provides a method for analyzing medical evaluations through adversarial networks with knowledge distillation, including:
[0006] Obtain medical evaluation data, perform emotion annotation on the medical evaluation data to obtain the medical evaluation data after emotion annotation, and the emotion annotation includes positive emotion, negative emotion, and mixed emotion;
[0007] Combine a generative adversarial network with the MASH module of a transformer model to obtain an improved generative adversarial network, input the medical evaluation data after emotion annotation into the improved generative adversarial network to generate training samples, and input the training samples into a teacher model for initialization to obtain an initialized teacher model;
[0008] Combine the GRU model, the first BERT model, and the first convolutional layer to obtain the first student model. Combine the LSTM network, the second BERT model, and the second convolutional layer to obtain the second student model. Through knowledge distillation, extract the knowledge learned from the initialized teacher model into the first student model and the second student model respectively, and train the first student model and the second student model with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model. Based on the trained first student model and the trained second student model, obtain the analysis result of the medical evaluation data.
[0009] The beneficial effects of the present invention are as follows: Generate training samples through a generative adversarial network without relying on a large amount of high-quality data in the real world, and combine the generative adversarial network with the MASH module of the transformer model, so that the improved generative adversarial network can more comprehensively and accurately capture the dependency relationships existing in the medical evaluation data with emotion annotations input, so as to improve the authenticity of the generated training samples and the quality of the training samples. Extract the knowledge learned from the teacher model into the first student model and the second student model through knowledge distillation, and train the first student model and the second student model with the goal of minimizing the distillation loss. The first student model and the second student model integrate the advantages of the GRU model, the BERT model, and the LSTM network, improving the accuracy and robustness of the trained first student model and the trained second student model. And using the compression and acceleration technology of knowledge distillation, the trained first student model and the trained second student model can analyze medical evaluations while reducing computational resource consumption and inference time, improving the analysis efficiency of medical evaluations and promoting the development of medical services.
[0010] Optionally, the obtaining of the medical evaluation data includes:
[0011] Construct a web crawler through the Selenium module of Python and the Chromium driver;
[0012] Generate a URL corresponding to the first page for each medical institution in the preset list of medical institutions through the web crawler, and automatically navigate to the corresponding first page to obtain medical evaluation data based on the URL using the Selenium module. The first page includes a map page, a social network page, and a medical institution website page;
[0013] Parse the medical evaluation data through Beautiful Soup to obtain the parsed medical evaluation data, which includes medical evaluation text, medical score, and timestamp;
[0014] Perform format conversion on the parsed medical evaluation data to obtain the medical evaluation data after format conversion. The format conversion includes CSV format conversion and JSON format conversion.
[0015] According to the above description, the web crawler generates a URL corresponding to the first page for each medical institution. Based on this URL, the Selenium module of the web crawler can automatically perform navigation to obtain medical evaluation data without manual acquisition, improving the efficiency of medical data acquisition. Moreover, the first page includes a map page, a social network page, and a medical institution website page, improving the integrity and comprehensiveness of the obtained medical evaluation data. Additionally, the obtained medical evaluation data will be parsed and format-converted to ensure the usability of the final medical evaluation data.
[0016] Optionally, the improved generative adversarial network includes an improved generator and an improved discriminator. The combination of the generative adversarial network and the MASH module of the transformer model to obtain the improved generative adversarial network includes:
[0017] Construct an improved generator. Combine the generator of the generative adversarial network with the MHSA module of the transformer model, and introduce the MHSA module as the initial layer to apply linear projection to modify the input tensor of the dimension to adapt to the actual number of attention heads. The MHSA module will perform calculations on all the attention heads according to the scaled dot product to generate attention weights applied to the corresponding values of the tensor, and concatenate and linearly process all the attention weights to obtain the final attention output.
[0018] Combine the MHSA module with an additional layer of spectral normalization linear transformation, a Leaky ReLU activation function, and Dropout filtering regularization, and add an attention layer in the ending block to obtain the improved generator.
[0019] Construct an improved discriminator. Combine each linear layer in the discriminator of the generative adversarial network with an additional layer of spectral normalization linear transformation to obtain the combined linear layer, and add residual connections between the combined linear layers. Optionally select a combined linear layer and combine it with a Softmax activation function as the output layer to obtain the improved discriminator.
[0020] According to the above description, the improved generative adversarial network not only improves the generator but also the discriminator. When improving the generator, the MHSA module is introduced as the initial layer, and the MHSA module is used to generate attention weights to improve the potential of the generator to capture subtle dependencies in the input noise. The MHSA module is combined with additional layers of spectral normalization linear transformation, the Leaky ReLU activation function, and Dropout filtering regularization to improve the generator's ability to capture subtle features and gradients within the latent space, thereby improving the authenticity, consistency, and diversity of the generated training samples. When improving the discriminator, the linear layer is combined with additional layers of spectral normalization linear transformation to stabilize the training process and improve convergence, and residual connections are added between the combined linear layers to help the gradient move in the network and reduce the problem of gradient disappearance in the generative adversarial network. The combined linear layer and the Softmax activation function are combined as the output layer to improve the discriminator's ability to distinguish between real medical evaluation data and generated medical evaluation data.
[0021] Optionally, the inputting of the sentiment-labeled medical evaluation data into the improved generative adversarial network to generate training samples includes:
[0022] Preprocessing the sentiment-labeled medical evaluation data to obtain preprocessed medical evaluation data, and inputting the preprocessed medical evaluation data into the improved generative adversarial network to generate training samples. The preprocessing includes: removing invalid text, lemmatization, and text vectorization.
[0023] According to the above description, the medical evaluation data input into the improved generative adversarial network is preprocessed, including removing invalid text, word reduction, and text vectorization. Removing invalid text can reduce the noise in the medical evaluation data and improve the quality of the medical evaluation data. Lemmatization improves the normalization of the medical evaluation data. Text vectorization makes the medical evaluation data more adaptable to the improved generative adversarial network and improves the usability of the medical evaluation data.
[0024] Optionally, the distilling of the knowledge learned in the initialized teacher model into the first student model and the second student model respectively through knowledge distillation, and training the first student model and the second student model with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model includes:
[0025] Introduce a first max pooling layer into the first student model, and use the first max pooling layer and the first convolutional layer as the first feature extractor to capture the first feature state of the knowledge through the first feature extractor;
[0026] Introduce a second max pooling layer into the second student model, and use the second max pooling layer and the second convolutional layer as a second feature extractor to capture the second feature state of the knowledge through the second feature extractor;
[0027] Train the first student model and the second student model with the goal of minimizing the distillation loss based on the first feature state and the second feature state to obtain the trained first student model and the trained second student model.
[0028] According to the above description, a first max pooling layer and a second max pooling layer are respectively introduced into the first student model and the second student model, and they are used together with the corresponding convolutional layers as feature extractors, which improves the ability of the first student model and the second student model to capture local patterns and features. That is, it ensures the accuracy of the captured first feature state and second feature state, and further improves the analysis ability of the first student model and the second student model trained based on the first feature state and the second feature state.
[0029] Optionally, the training of the first student model and the second student model with the goal of minimizing the distillation loss based on the first feature state and the second feature state to obtain the trained first student model and the trained second student model includes:
[0030] The first student model outputs an embedding e through the first BERT model t , and calculates the first current hidden state of the knowledge according to the embedding e t , the first feature state, and the recursive update equation of the GRU model. The recursive update equation is:
[0031]
[0032] where h t represents the first current hidden state at position t, h t-1 represents the first feature state at position t - 1, and e t represents the vector generated by the first BERT model after processing the input text at position t;
[0033] Calculate the first update gate through the first update gate mechanism formula. The first update gate mechanism formula is:
[0034]
[0035] where z t represents the first update gate at position t, represents the sigmoid activation function, W z represents the weight matrix related to the first update gate, and bz represents the bias vector related to the first update gate, h t-1 represents the first feature state at position t - 1, e t represents the vector generated after the first BERT model processes the input text at position t;
[0036] When the first update gate is lower than the update threshold, retain the first current hidden state; otherwise, calculate the reset gate through the reset gate mechanism formula, and the reset gate mechanism formula is:
[0037]
[0038] where r t represents the reset gate at position t, W r represents the weight matrix related to the reset gate, b r represents the bias vector related to the reset gate, h t-1 represents the first feature state at position t - 1, e t represents the vector generated after the first BERT model processes the input text at position t;
[0039] Calculate the candidate gate according to the reset gate and the candidate gate mechanism formula, and the candidate gate mechanism formula is:
[0040]
[0041] where represents the candidate gate at position t, tanh represents the hyperbolic tangent function, W h represents the weight matrix related to the candidate gate, b h represents the bias vector related to the candidate gate, r t represents the reset gate at position t, h t-1 represents the first feature state at position t - 1, e t represents the vector generated after the first BERT model processes the input text at position t;
[0042] Recalculate the first current hidden state of the knowledge according to the candidate gate, the reset gate and the first hidden formula to obtain the latest first current hidden state, and the first hidden formula is:
[0043]
[0044] where H t represents the latest first current hidden state at position t, z t represents the first update gate at position t, h t-1 represents the first feature state at position t - 1, represents the candidate gate at position t;
[0045] The second student model outputs an embedding through the second BERT model , and based on the embedding , the input gate of the knowledge is calculated according to the embedding, the second feature state, and the input gate mechanism formula of the LSTM network. The forgetting gate of the knowledge is calculated according to the second feature state and the forgetting mechanism formula of the LSTM. The second update gate of the knowledge is calculated according to the second feature state and the second update gate mechanism formula of the LSTM. The input gate mechanism formula is:
[0046]
[0047] where represents the input gate at position t, represents the sigmoid activation function, W i represents the weight matrix related to the input gate, b i represents the bias vector related to the input gate, represents the second feature state at position t-1, represents the vector generated by the second BERT model after processing the input text at position t;
[0048] The forgetting mechanism formula is:
[0049]
[0050] where represents the forgetting gate at position t, represents the sigmoid activation function, W f represents the weight matrix related to the forgetting gate, b f represents the bias vector related to the forgetting gate, represents the second feature state at position t-1, represents the vector generated by the second BERT model after processing the input text at position t
[0051] The second update gate mechanism formula is:
[0052]
[0053] where represents the second update gate at position t, tanh represents the hyperbolic tangent function, W g represents the weight matrix related to the second update gate, b g represents the bias vector related to the second update gate, represents the second feature state at position t-1, Denotes the vector generated after the second BERT model processes the input text at position t
[0054] Calculate the candidate memory cell according to the second update gate, the input gate and the candidate memory cell formula, and the candidate memory cell formula is:
[0055]
[0056] Wherein, Denotes the candidate memory cell at position t, Denotes the second update gate at position t, Denotes the input gate at position t;
[0057] Calculate the latest memory cell according to the candidate memory cell, the forget gate and the memory update formula, and the memory update formula is:
[0058]
[0059] Wherein, Denotes the latest memory cell at position t, Denotes the forget gate at position t, Denotes the candidate memory cell at position t, Denotes the latest memory cell at position t-1;
[0060] Calculate the output gate of the knowledge according to the second feature state and the output gate mechanism formula of the LSTM, and the output gate mechanism formula is:
[0061]
[0062] Wherein, Denotes the output gate at position t, Denotes the sigmoid activation function, W o Denotes the weight matrix related to the output gate, b o Denotes the bias vector related to the output gate, Denotes the second feature state at position t-1, Denotes the vector generated after the second BERT model processes the input text at position t
[0063] Calculate the second current hidden state of the knowledge according to the output gate, the latest memory cell and the second hidden formula, and the second hidden formula is:
[0064]
[0065] Wherein, represents the second current hidden state at position t, represents the output gate at position t, represents the latest memory cell at position t, where tanh represents the hyperbolic tangent function;
[0066] Input the first current hidden state and the second current hidden state into a dense layer for classification to obtain the predicted probability distribution for each category. Calculate the distillation loss based on the predicted probability distribution, and train the first student model and the second student model with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model.
[0067] According to the above description, the first student model uses the input embedding e of the first BERT model and the recursive update equation of the GRU model to calculate the first current hidden state of knowledge, improving the ability of the first student model to capture the potential temporal dependence of the input knowledge. And according to the first update gate, it decides the balance between retaining the previous first current hidden state and incorporating the candidate gate to recalculate the first current hidden state, enabling the first student model to adaptively update the first current hidden state representation according to the input embedding e and the previous context, thereby improving the analysis ability of the first student model. The second student model updates the memory cell according to the forget gate of the forgetting mechanism to obtain the latest memory cell. The output gate can regulate the information flow from the latest memory cell to the second current hidden state, and the second current hidden state is calculated based on the activation of the output gate and the latest memory unit, improving the ability of the second student model to capture the context and sequential information contributed by the LSTM in the second BERT model, and thus improving the analysis ability of the second student model.
[0068] Optionally, the inputting the first current hidden state and the second current hidden state into a dense layer for classification includes:
[0069] Input the first current hidden state and the second current hidden state into a dense layer for classification to obtain the predicted scores for each category, and use the Softmax activation function to convert the predicted scores into a predicted probability distribution.
[0070] According to the above description, the linear transformation of the dense layer can convert the first current hidden state and the second current hidden state into the same dimension as the number of categories, thereby ensuring the accuracy of the predicted scores for each category obtained. Using the Softmax activation function to convert the predicted scores into a predicted probability distribution ensures the accuracy of the obtained predicted probability distribution.
[0071] Optionally, calculating the distillation loss according to the predicted probability distribution, and training the first student model and the second student model with the goal of minimizing the distillation loss, obtaining the trained first student model and the trained second student model includes:
[0072] Calculating the distillation loss of each training sample according to the predicted probability distribution by using sparse categorical cross-entropy, and calculating the total distillation loss according to the distillation loss and the total distillation loss formula, the total distillation loss formula is:
[0073]
[0074] Wherein, represents the total distillation loss, represents the distillation loss of the β-th training sample, and N represents the total number of training samples;
[0075] Training the first student model and the second student model by using the Adam algorithm with the goal of minimizing the total distillation loss, obtaining the trained first student model and the trained second student model.
[0076] According to the above description, calculating the distillation loss by using sparse categorical cross-entropy can effectively handle the class imbalance problem, optimize the predicted probability distribution of the student model, take the average distillation loss of all training samples as the total distillation loss, and train with the goal of minimizing the total distillation loss to ensure the rationality of the trained first student model and the trained second student model. When training with the goal of minimizing the total distillation loss, the Adam algorithm is used, so that the advantages of momentum and adaptive learning rate can be combined during training, quickly converge during the training process and avoid falling into local optima, and improve the analysis ability of the trained first student model and the trained second student model.
[0077] In a second aspect, the present invention provides a system for analyzing medical evaluations through a knowledge distillation adversarial network, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for analyzing medical evaluations through a knowledge distillation adversarial network described in the first aspect.
[0078] Wherein, the technical effects corresponding to the system for analyzing medical evaluations through a knowledge distillation adversarial network provided in the second aspect refer to the relevant descriptions of the method for analyzing medical evaluations through a knowledge distillation adversarial network provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a flowchart of the method for analyzing medical evaluations through a knowledge distillation adversarial network provided in this embodiment;
[0080] Figure 2 Schematic diagram of the overall process of a method for analyzing medical evaluations through a knowledge distillation-based adversarial network provided in this embodiment;
[0081] Figure 3 Schematic diagram of the process for training the first student model and the second student model involved in this embodiment;
[0082] Figure 4 Schematic diagram of the structure of a system for analyzing medical evaluations through a knowledge distillation-based adversarial network provided in this embodiment.
[0083] Explanation of reference numerals
[0084] 1. A system for analyzing medical evaluations through a knowledge distillation-based adversarial network;
[0085] 2. Processor;
[0086] 3. Memory. Detailed implementation manners
[0087] To better understand the above technical solutions, the exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more clear and thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0088] Embodiment 1
[0089] Please refer to Figures 1 to 3 , the present invention provides a method for analyzing medical evaluations through a knowledge distillation-based adversarial network, including the steps of:
[0090] S1. Obtain medical evaluation data, perform emotion annotation on the medical evaluation data to obtain the medical evaluation data after emotion annotation, and the emotion annotation includes positive emotion, negative emotion, and mixed emotion;
[0091] In this embodiment, as Figure 2As shown, perform emotion annotation on the obtained medical evaluation data, including positive emotions, negative emotions, and mixed emotions. Mixed emotions refer to medical evaluation data that includes both positive and negative emotions, as well as medical evaluation data with a neutral attitude. After emotion annotation, count the number of medical evaluation data with the same emotion annotation, that is, obtain the number of positive emotion data, the number of negative emotion data, and the number of mixed emotion data. When the number of any one type of emotion data is lower than the emotion threshold, an oversampling strategy is used to increase the number of emotion data corresponding to the emotion threshold that is lower than the threshold. For example, if the emotion threshold is 50 and the number of mixed emotion data is 35, then the oversampling strategy is used to increase the number of mixed emotion data so that it is not lower than the emotion threshold. The specific emotion threshold is set according to the actual situation.
[0092] At this time, the obtaining of medical evaluation data described in step S1 includes:
[0093] S11. Construct a web crawler through the Selenium module of Python and the Chromium driver;
[0094] S12. Generate a URL corresponding to the first page for each medical institution in the preset medical institution list through the web crawler, and use the Selenium module to automatically navigate to the corresponding first page based on the URL to obtain medical evaluation data. The first page includes a map page, a social network page, and a medical institution website page;
[0095] S13. Parse the medical evaluation data through Beautiful Soup to obtain the parsed medical evaluation data. The parsed medical evaluation data includes medical evaluation text, medical score, and timestamp;
[0096] S14. Perform format conversion on the parsed medical evaluation data to obtain the medical evaluation data after format conversion. The format conversion includes CSV format conversion and JSON format conversion.
[0097] In this embodiment, as Figure 2As shown, a web crawler is built using Python's Selenium module and Chromium driver. Through this web crawler, a URL corresponding to the first page of each medical institution in a preset list of medical institutions is generated. Based on this URL, the Selenium module of the web crawler can automatically navigate to the corresponding first page and obtain medical evaluation data from it. The first page includes a map page, a social network page, and a medical institution website page. That is, the obtained medical evaluation data is sourced from the map page, the social network page, and the medical institution website page. The preset list of medical institutions can be a pre-built and stored list of medical institutions or a currently customized list of medical institutions. Beautiful Soup is a parsing library based on the HTML DOM. Therefore, the obtained medical evaluation data is parsed using BeautifulSoup to obtain medical evaluation data including medical evaluation text, medical scores, and timestamps, and it is converted into CSV format and JSON format.
[0098] S2. Combine the generative adversarial network with the MASH module of the transformer model to obtain an improved generative adversarial network. Input the medical evaluation data with emotion annotations into the improved generative adversarial network to generate training samples, and input the training samples into the teacher model for initialization to obtain an initialized teacher model;
[0099] In this embodiment, as Figure 2As shown, the generative adversarial network is improved by combining the generative adversarial network with the MASH module of the transformer model. The MASH module refers to the Multi-Head Attention module, so as to obtain an improved generative adversarial network. Then, the medical evaluation data with emotion annotation obtained in step S1 is input into the improved generative adversarial network to generate training samples, and the training samples are input into the teacher model for initialization. At this time, the teacher model is a large and powerful neural network model, such as: DistilBERT (knowledge distillation type BERT model), XLNet (Exploring the Limits of Pre-trained Language Models, natural language processing model), ERNIE3.0 (Enhanced Representation through kNowledge IntEgration 3.0, Wenxin 3.0), ALBERT (lightweight BERT model), m-BERT (Multilingual BERT, bilingual training model), Indic-BERT (variant of BERT model optimized for Indian languages), and XLM-RoBERTa (Cross-Lingual Language Model RoBERTa, multilingual pre-training model), thus obtaining the initialized teacher model.
[0100] At this time, the improved generative adversarial network described in step S2 includes an improved generator and an improved discriminator. The combination of the generative adversarial network and the MASH module of the transformer model to obtain the improved generative adversarial network includes:
[0101] S21. Construct an improved generator by combining the generator of the generative adversarial network with the MHSA module of the transformer model, introducing the MHSA module as the initial layer, and applying linear projection to modify the input tensor of the dimension to adapt to the actual number of attention heads. The MHSA module will perform calculations on all the attention heads according to the scaled dot product, generating attention weights applied to the corresponding values of the tensor, and concatenating and linearly processing all the attention weights to obtain the final attention output;
[0102] S22. Combine the MHSA module with the additional layer of spectral normalization linear transformation, the Leaky ReLU activation function, and Dropout filtering regularization, and add an attention layer in the ending block to obtain the improved generator;
[0103] In this embodiment, as Figure 2As shown, the improved generative adversarial network includes an improved generator and an improved discriminator. The improved generator is obtained by combining the original generator of the generative adversarial network with the MHSA module of the transformer model. By introducing the MHSA module as the initial layer, the potential of continuously focusing on the subtle dependencies in the input noise can be exploited. The MHSA module applies linear projections to modify the input tensor of dimensions to adapt to the actual number of attention heads. Specifically, the MHSA module applies linear projections Q, K, and V to modify the input tensor X of dimensions, where Q represents the Query vector, K represents the Key vector, V represents the Value vector, B represents the Batch Size, i.e., the number of samples in a batch of input data, L represents the Length or Sequence Length, i.e., the length of the sequence data, C represents the Channels or Features, i.e., the number of features per time step or per word, and the input tensor X represents a way of representing the shape. And the MHSA module calculates according to the scaled dot product over all attention heads to generate the attention weights applied to the corresponding values of the tensor, i.e., the attention weights of V(X). The concatenation process and linear process are performed on all attention weights to obtain the final attention output, which can be represented by Y. The MHSA module is combined with an additional layer of spectral normalization linear transformation, the Leaky ReLU activation function, and Dropout regularization. At this time, the additional layer of spectral normalization linear transformation selects a spectral normalization linear transformation with a shape of (512, 256). The Leaky ReLU activation function actually introduces a small negative slope in the negative input region to limit the complete suppression of information, thereby improving the ability of the generator to capture subtle features and gradients inside the latent space. And an attention layer is added in the ending block to obtain the improved generator. That is, the improved generator adds two stages: 1 is the initial layer: the MHSA module, and 2 is the ending block: the attention layer. The input tensor X of dimensions, where Q represents the Query vector, K represents the Key vector, V represents the Value vector, B represents the Batch Size, i.e., the number of samples in a batch of input data, L represents the Length or Sequence Length, i.e., the length of the sequence data, C represents the Channels or Features, i.e., the number of features per time step or per word. The input tensor X represents a way of representing the shape. And the MHSA module calculates according to the scaled dot product over all attention heads to generate the attention weights applied to the corresponding values of the tensor, i.e., the attention weights of V(X). The concatenation process and linear process are performed on all attention weights to obtain the final attention output, which can be represented by Y. The MHSA module is combined with an additional layer of spectral normalization linear transformation, the Leaky ReLU activation function, and Dropout regularization. At this time, the additional layer of spectral normalization linear transformation selects a spectral normalization linear transformation with a shape of (512, 256). The Leaky ReLU activation function actually introduces a small negative slope in the negative input region to limit the complete suppression of information, thereby improving the ability of the generator to capture subtle features and gradients inside the latent space. And an attention layer is added in the ending block to obtain the improved generator. That is, the improved generator adds two stages: 1 is the initial layer: the MHSA module, and 2 is the ending block: the attention layer.
[0104] S23. Construct an improved discriminator. Combine each linear layer in the discriminator of the generative adversarial network with an additional layer of spectral normalization linear transformation to obtain the combined linear layer, and add residual connections between the combined linear layers. Arbitrarily select a combined linear layer and combine it with a Softmax activation function as the output layer to obtain the improved discriminator.
[0105] In this embodiment, as Figure 2As shown, the original discriminator of the generative adversarial network consists of multiple layers, such as convolutional layers, fully connected layers, and linear layers, and each layer is followed by a Leaky ReLU activation function and Dropout regularization. The improved discriminator combines each linear layer with an additional layer of spectral normalization linear transformation on the basis of the original discriminator, adds residual connections between each combined linear layer except the last one, and selects any combined linear layer and a Softmax activation function to be combined as the output layer. The output layer mainly generates logits, that is, the numerical values of the probability distribution between the actual category and the false category, so as to obtain the improved discriminator. Among them, the Leaky ReLU activation function behind each layer is the Leaky ReLU activation function with a slope of 0.2.
[0106] At this time, the input of the medical evaluation data with emotion annotation into the improved generative adversarial network to generate training samples in step S2 includes:
[0107] S24. Preprocess the medical evaluation data with emotion annotation to obtain the preprocessed medical evaluation data, and input the preprocessed medical evaluation data into the improved generative adversarial network to generate training samples. The preprocessing includes: removing invalid text, lemmatization, and text vectorization.
[0108] In this embodiment, the medical evaluation data with emotion annotation is preprocessed including removing invalid text, lemmatization, and text vectorization to obtain the preprocessed medical evaluation data. Among them, removing invalid text is to delete all forms of emojis, punctuation marks, and duplicate data, lemmatization is to implement the normalization process of the text by retaining the original form of the lemma, and text vectorization is to convert the medical evaluation data into a numerical representation.
[0109] S3. Combine the GRU model, the first BERT model, and the first convolutional layer to obtain the first student model, combine the LSTM network, the second BERT model, and the second convolutional layer to obtain the second student model, and distill the knowledge learned from the initialized teacher model into the first student model and the second student model respectively through knowledge distillation. Train the first student model and the second student model with the minimum distillation loss as the goal to obtain the trained first student model and the trained second student model, and obtain the analysis result of the medical evaluation data based on the trained first student model and the trained second student model.
[0110] In this embodiment, as Figure 2As shown, different from directly using a small and lightweight neural network as the student model in the traditional way, two improved student models are combined in this application. Among them, the first student model is formed by combining a GRU model, a first BERT model, and a first convolutional layer, and the second student model is formed by combining an LSTM network, a second BERT model, and a second convolutional layer. The knowledge learned from the initialized teacher model is respectively distilled into the first student model and the second student model through knowledge distillation. The first student model and the second student model are trained with the goal of minimizing the distillation loss, and the analysis result of medical evaluation data is obtained based on the trained first student model and the trained second student model.
[0111] At this time, in step S3, the knowledge learned from the initialized teacher model is respectively distilled into the first student model and the second student model through knowledge distillation, and the first student model and the second student model are trained with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model, including:
[0112] S31. Introduce a first max-pooling layer into the first student model, and use the first max-pooling layer and the first convolutional layer as a first feature extractor to capture the first feature state of the knowledge through the first feature extractor;
[0113] S32. Introduce a second max-pooling layer into the second student model, and use the second max-pooling layer and the second convolutional layer as a second feature extractor to capture the second feature state of the knowledge through the second feature extractor;
[0114] S33. Train the first student model and the second student model with the goal of minimizing the distillation loss based on the first feature state and the second feature state to obtain the trained first student model and the trained second student model.
[0115] In this embodiment, as Figure 2 shown, a first max-pooling layer and a second max-pooling layer are respectively introduced into the first student model and the second student model, and the first max-pooling layer and the first convolutional layer are used together as the first feature extractor of the first student model, and the second max-pooling layer and the second convolutional layer are used together as the second feature extractor of the second student model. The first feature state of the knowledge distilled from the teacher model is captured through the first feature extractor, and the second feature state of the knowledge distilled from the teacher model is captured through the second feature extractor. Among them, when capturing the first feature state through the first convolutional layer or capturing the second feature state through the second convolutional layer, the captured first feature state or second feature state can be represented by a convolutional feature formula, and the convolutional feature formula is as follows:
[0116]
[0117] Among them, C αt represents the feature state output by the convolutional layer for the α-th filter at position t, and w αj represents the weight matrix related to the α-th filter, represents the value of the refined knowledge at the (j - 1)-th position of the α-th filter, k represents the number of all positions, and θ α represents the bias vector related to the α-th filter;
[0118] The first max pooling layer and the second max pooling layer are used to downsample the first feature state and the second feature state. Specifically: the first max pooling layer and the second max pooling layer select the maximum activation in each window with a specified pool size of 2, so as to reduce the dimensions of the first feature state and the second feature state while retaining important information. The first feature state and the second feature state after the operations of the first max pooling layer and the second max pooling layer can be represented by the pooling feature formula. The pooling feature formula is:
[0119]
[0120] Among them, represents the feature state output by the max pooling layer for the α-th filter at position t, C α2t represents the feature state output by the α-th filter at position t in the window with a pool size of 2, C α2t+1 represents the feature state output by the α-th filter at position t + 1 in the window with a pool size of 2;
[0121] Based on the finally obtained first feature state and second feature state, the first student model and the second student model are trained with the goal of minimizing the refinement loss to obtain the trained first student model and the trained second student model.
[0122] In this embodiment, step S33 includes the training of the first student model and the training of the second student model. Among them, the specific steps of training the first student model are as follows:
[0123] The first student model outputs the embedding e t through the first BERT model. According to the embedding e t and the first feature state and the recursive update equation of the GRU model, the first current hidden state of the knowledge is calculated. The recursive update equation is:
[0124]
[0125] Among them, h t represents the first current hidden state at position t, ht-1 Represents the first feature state at position t-1, e t Represents the vector generated by the first BERT model after processing the input text at position t;
[0126] Calculate the first update gate through the first update gate mechanism formula, and the first update gate mechanism formula is:
[0127]
[0128] where z t Represents the first update gate at position t, Represents the sigmoid activation function, W z Represents the weight matrix related to the first update gate, b z Represents the bias vector related to the first update gate, h t-1 Represents the first feature state at position t-1, e t Represents the vector generated by the first BERT model after processing the input text at position t;
[0129] When the first update gate is lower than the update threshold, retain the first current hidden state. Otherwise, calculate the reset gate through the reset gate mechanism formula, and the reset gate mechanism formula is:
[0130]
[0131] where r t Represents the reset gate at position t, W r Represents the weight matrix related to the reset gate, b r Represents the bias vector related to the reset gate, h t-1 Represents the first feature state at position t-1, e t Represents the vector generated by the first BERT model after processing the input text at position t;
[0132] Calculate the candidate gate according to the reset gate and the candidate gate mechanism formula, and the candidate gate mechanism formula is:
[0133]
[0134] where, Represents the candidate gate at position t, tanh represents the hyperbolic tangent function, W h Represents the weight matrix related to the candidate gate, b h Represents the bias vector related to the candidate gate, r t Represents the reset gate at position t, h t-1 Represents the first feature state at position t-1, e tDenotes the vector generated after the first BERT model processes the input text at position t;
[0135] Recalculate the first current hidden state of the knowledge according to the candidate gate, the reset gate and the first hidden formula to obtain the latest first current hidden state, and the first hidden formula is:
[0136]
[0137] Where, H t Denotes the latest first current hidden state at position t, z t Denotes the first update gate at position t, h t-1 Denotes the first feature state at position t-1, Denotes the candidate gate at position t;
[0138] In this embodiment, as Figure 3 Shown, when training the first student model to calculate the first current hidden state, not only the output embedding e of the first BERT model is utilized t , but also the recursive update equation of the GRU model is utilized, and a series of mathematical operations of the GRU model are introduced in the process of integrating the GRU model into the first BERT model, such as: the first update gate mechanism formula, the reset gate mechanism formula, the candidate gate mechanism formula and the first hidden formula, and according to the calculated first update gate, it is determined whether to retain the previous first current hidden state or recalculate the latest first current hidden state, so that the first student model can adaptively update the hidden state representation according to the input embedding e t And the previous context.
[0139] Among them, the specific training steps of the second student model are as follows:
[0140] The second student model outputs an embedding through the second BERT model , according to the embedding , the input gate of the knowledge is calculated according to the input gate mechanism formula of the second feature state and the LSTM network, the forget gate of the knowledge is calculated according to the forget mechanism formula of the second feature state and the LSTM, and the second update gate of the knowledge is calculated according to the second update gate mechanism formula of the second feature state and the LSTM. The input gate mechanism formula is:
[0141]
[0142] Where, Denotes the input gate at position t, Denotes the sigmoid activation function, W i Denotes the weight matrix related to the input gate, bi represents the bias vector related to the input gate, represents the second feature state at position t - 1, represents the vector generated after the second BERT model processes the input text at position t
[0143] The formula of the forgetting mechanism is:
[0144]
[0145] where, represents the forgetting gate at position t, represents the sigmoid activation function, W f represents the weight matrix related to the forgetting gate, b f represents the bias vector related to the forgetting gate, represents the second feature state at position t - 1, represents the vector generated after the second BERT model processes the input text at position t
[0146] The formula of the second update gate mechanism is:
[0147]
[0148] where, represents the second update gate at position t, tanh represents the hyperbolic tangent function, W g represents the weight matrix related to the second update gate, b g represents the bias vector related to the second update gate, represents the second feature state at position t - 1, represents the vector generated after the second BERT model processes the input text at position t
[0149] Calculate the candidate storage unit according to the second update gate, the input gate and the candidate storage unit formula, and the candidate storage unit formula is:
[0150]
[0151] where, represents the candidate storage unit at position t, represents the second update gate at position t, represents the input gate at position t;
[0152] Calculate the latest storage unit according to the candidate storage unit, the forgetting gate and the storage update formula, and the storage update formula is:
[0153]
[0154] Among them, represents the latest storage unit at position t, represents the forget gate at position t, represents the candidate storage unit at position t, represents the latest storage unit at position t-1;
[0155] Calculate the output gate of the knowledge according to the second feature state and the output gate mechanism formula of the LSTM. The output gate mechanism formula is:
[0156]
[0157] Among them, represents the output gate at position t, represents the sigmoid activation function, W o represents the weight matrix related to the output gate, b o represents the bias vector related to the output gate, represents the second feature state at position t-1, represents the vector generated by the second BERT model after processing the input text at position t
[0158] Calculate the second current hidden state of the knowledge according to the output gate, the latest storage unit and the second hidden formula. The second hidden formula is:
[0159]
[0160] Among them, represents the second current hidden state at position t, represents the output gate at position t, represents the latest storage unit at position t, and tanh represents the hyperbolic tangent function;
[0161] In this embodiment, as Figure 3 shown, when training the second student model to calculate the second current hidden state, not only the output embedding of the second BERT model is utilized , but also the forgetting and updating mechanism of the LSTM network is utilized. A series of mathematical operations of the LSTM network are introduced in the process of integrating the LSTM network into the second BERT model, such as: the input gate mechanism formula, the forgetting mechanism formula, the second update gate mechanism formula, the candidate storage unit formula, the storage update formula, the output gate mechanism formula, and the second hidden formula. The output gate can adjust the information flow from the storage unit to the second hidden state.
[0162] Input the first current hidden state and the second current hidden state into a dense layer for classification to obtain the predicted probability distribution for each category. Calculate the distillation loss based on the predicted probability distribution, and train the first student model and the second student model with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model.
[0163] At this time, the inputting the first current hidden state and the second current hidden state into a dense layer for classification includes:
[0164] Input the first current hidden state and the second current hidden state into a dense layer for classification to obtain the predicted scores for each category, and use the Softmax activation function to convert the predicted scores into the predicted probability distribution.
[0165] In this embodiment, as Figure 3 shown, input the obtained first current hidden state and second current hidden state into a dense layer for classification to obtain the predicted probability distribution for each category. The category here refers to the category corresponding to the emotion annotation in step S1. The dense layer can map the first current hidden state and the second current hidden state to the category space of the emotion classification task, thereby obtaining the predicted scores for each category, and use the Softmax activation function to convert the obtained predicted scores into the predicted probability distribution.
[0166] At this time, the calculating the distillation loss based on the predicted probability distribution and training the first student model and the second student model with the goal of minimizing the distillation loss to obtain the trained first student model and the trained second student model includes:
[0167] Use sparse categorical cross-entropy to calculate the distillation loss for each training sample based on the predicted probability distribution, and calculate the total distillation loss according to the distillation loss and the total distillation loss formula. The total distillation loss formula is:
[0168]
[0169] where represents the total distillation loss, represents the distillation loss of the β-th training sample, and N represents the total number of training samples;
[0170] Train the first student model and the second student model with the goal of minimizing the total distillation loss using the Adam algorithm to obtain the trained first student model and the trained second student model.
[0171] In this embodiment, as Figure 3As shown, the refined loss of each training sample is calculated according to the predicted probability distribution using sparse categorical cross-entropy, and the total refined loss is calculated according to the total refined loss formula. The Adam algorithm is used to train the first student model and the second student model with the goal of minimizing the total refined loss, so as to obtain the trained first student model and the trained second student model.
[0172] Embodiment 2
[0173] Please refer to Figure 4 , the present invention provides a system 1 for analyzing medical evaluations through a knowledge distillation-based adversarial network, including a memory 3, a processor 2, and a computer program stored on the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in Embodiment 1 are implemented.
[0174] Since the system / apparatus described in the above embodiments of the present invention is the system / apparatus adopted for implementing the method in the above embodiments of the present invention, based on the method described in the above embodiments of the present invention, those skilled in the art can understand the specific structure and variations of the system / apparatus, and thus will not be elaborated here. Any system / apparatus adopted for the method in the above embodiments of the present invention falls within the scope of protection of the present invention.
[0175] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.
[0177] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a claim listing several means, several of these means can be embodied by the same hardware element. The use of the terms first, second, third, etc. is for convenience only and does not denote any order. These terms can be construed as part of the name of the element.
[0178] In addition, it should be noted that in the description of this specification, the descriptions of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0179] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concept. Therefore, the claims should be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0180] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention should also include these modifications and variations.
Claims
1. A method for analyzing medical evaluation through adversarial networks using knowledge distillation, characterized in that: include: Acquire medical evaluation data, and perform emotion annotation on the medical evaluation data to obtain emotion-annotated medical evaluation data, wherein the emotion annotation includes positive emotion, negative emotion, and mixed emotion; The generative adversarial network is combined with the MASH module of the transformer model to obtain an improved generative adversarial network, the medical evaluation data with emotion annotation is input into the improved generative adversarial network to generate training samples, and the training samples are input into the teacher model for initialization to obtain an initialized teacher model; A first student model is obtained by combining a GRU model, a BERT model, and a first convolutional layer, and a second student model is obtained by combining an LSTM network, a BERT model, and a second convolutional layer, and the knowledge learned in the initialized teacher model is respectively refined to the first student model and the second student model through knowledge distillation, and the first student model and the second student model are trained with the minimum extraction loss as the goal to obtain a trained first student model and a trained second student model, and an analysis result of the medical evaluation data is obtained based on the trained first student model and the trained second student model; The improved generative adversarial network includes an improved generator and an improved discriminator. The improved generative adversarial network is obtained by combining the generative adversarial network with the MASH module of the transformer model, including: Construct an improved generator, combining the generator of the generative adversarial network with the MHSA module of the transformer model, introducing the MHSA module as the initial layer to apply linear projection to modify the dimensional input tensor to adapt to the actual number of attention heads, and the MHSA module will calculate the proportional dot product on all the attention heads to generate the attention weights applied to the corresponding values of the tensor, and all the attention weights will be processed in series and linearly to obtain the final attention output; Combining the MHSA module with an additional layer of spectral normalization linear transformation, Leaky ReLU activation function and Dropout filtering regularization, and adding an attention layer in the end block to obtain an improved generator; An improved discriminator is constructed by combining each linear layer in the discriminator of the generative adversarial network with an additional layer of spectral normalization linear transformation to obtain a combined linear layer, and adding residual connections between each combined linear layer. An arbitrary combined linear layer is selected and combined with a Softmax activation function as the output layer to obtain an improved discriminator.
2. The method for analyzing medical evaluation through adversarial networks using knowledge distillation as claimed in claim 1, characterized in that: The obtaining of medical evaluation data includes: Build a web search engine using Python's Selenium module and Chromium driver; Generate a URL corresponding to a first page for each medical institution in a preset list of medical institutions by using the network search engine, and automatically navigate to the corresponding first page based on the URL using the Selenium module to obtain medical evaluation data, wherein the first page includes a map page, a social network page, and a medical institution website page; Parsing the medical evaluation data by using Beautiful Soup to obtain parsed medical evaluation data, wherein the parsed medical evaluation data includes a medical evaluation text, a medical score, and a timestamp; The parsed medical evaluation data is format-converted to obtain format-converted medical evaluation data, wherein the format conversion includes CSV format conversion and JSON format conversion.
3. The method for analyzing medical evaluation through adversarial networks using knowledge distillation as claimed in claim 1, characterized in that: The step of inputting the emotion-labeled medical evaluation data into the improved generative adversarial network to generate training samples includes: The medical evaluation data after emotion annotation is preprocessed to obtain preprocessed medical evaluation data, and the preprocessed medical evaluation data is input into an improved generative adversarial network to generate training samples, wherein the preprocessing includes: removing invalid text, word form restoration and text vectorization.
4. The method for analyzing medical evaluation through adversarial networks using knowledge distillation as claimed in claim 1, characterized in that: The method of extracting the knowledge learned from the initialized teacher model into the first student model and the second student model respectively through knowledge distillation, training the first student model and the second student model with the minimum extraction loss as the goal, and obtaining the trained first student model and the trained second student model includes: Introducing a first maximum pooling layer into the first student model, and using the first maximum pooling layer and the first convolutional layer as a first feature extractor, and capturing a first feature state of the knowledge through the first feature extractor; Introducing a second maximum pooling layer into the second student model, and using the second maximum pooling layer and the second convolutional layer as a second feature extractor, and capturing a second feature state of the knowledge through the second feature extractor; Based on the first feature state and the second feature state, the first student model and the second student model are trained with the minimum refinement loss as the goal to obtain a trained first student model and a trained second student model.
5. The method for analyzing medical evaluation through adversarial network using knowledge distillation as claimed in claim 4, characterized in that: The step of training the first student model and the second student model based on the first feature state and the second feature state with a minimum extraction loss as a goal to obtain the trained first student model and the trained second student model comprises: The first student model is embedded by the first BERT model output t , according to the embedding e t , the first feature state and the recursive update equation of the GRU model calculate the first current hidden state of the knowledge, and the recursive update equation is: Among them, h t represents the first current hidden state at position t, h t-1 represents the first characteristic state at position t-1, e t represents the vector generated by the first BERT model after processing the input text at position t; The first update gate is calculated by the first update gate mechanism formula, wherein the first update gate mechanism formula is: Among them, z t represents the first update gate at position t, represents the sigmoid activation function, W z represents the weight matrix associated with the first update gate, b z represents the bias vector associated with the first update gate, h t-1 represents the first characteristic state at position t-1, e t represents the vector generated by the first BERT model after processing the input text at position t; When the first update gate is lower than the update threshold, the first current hidden state is retained, otherwise, the reset gate is calculated by the reset gate mechanism formula, and the reset gate mechanism formula is: Among them, r t represents the reset gate at position t, W r represents the weight matrix associated with the reset gate, b r represents the bias vector associated with the reset gate, h t-1 represents the first characteristic state at position t-1, e t represents the vector generated by the first BERT model after processing the input text at position t; The candidate gate is calculated according to the reset gate and the candidate gate mechanism formula, and the candidate gate mechanism formula is: in, represents the candidate gate at position t, tanh represents the hyperbolic tangent function, W h represents the weight matrix associated with the candidate gate, b h represents the bias vector associated with the candidate gate, r t represents the reset gate at position t, h t-1 represents the first characteristic state at position t-1, e t represents the vector generated by the first BERT model after processing the input text at position t; The first current hidden state of the knowledge is recalculated according to the candidate gate, the reset gate and the first hidden formula to obtain the latest first current hidden state, wherein the first hidden formula is: = Among them, H t represents the latest first current hidden state at position t, z t represents the first update gate at position t, h t-1 represents the first characteristic state at position t-1, represents the candidate gate at position t; The second student model outputs embeddings through the second BERT model , according to the embedding , the input gate of the knowledge is calculated according to the second feature state and the input gate mechanism formula of the LSTM network, the forget gate of the knowledge is calculated according to the second feature state and the forget mechanism formula of the LSTM, and the second update gate of the knowledge is calculated according to the second feature state and the second update gate mechanism formula of the LSTM, and the input gate mechanism formula is: in, represents the input gate at position t, represents the sigmoid activation function, W i represents the weight matrix associated with the input gate, b i represents the bias vector associated with the input gate, represents the second characteristic state at position t-1, Represents the vector generated by the second BERT model after processing the input text at position t The forgetting mechanism formula is: in, represents the forget gate at position t, represents the sigmoid activation function, W f represents the weight matrix associated with the forget gate, b f represents the bias vector associated with the forget gate, represents the second characteristic state at position t-1, Represents the vector generated by the second BERT model after processing the input text at position t The second update gate mechanism formula is: in, represents the second update gate at position t, tanh represents the hyperbolic tangent function, W g represents the weight matrix associated with the second update gate, b g represents the bias vector associated with the second update gate, represents the second characteristic state at position t-1, Represents the vector generated by the second BERT model after processing the input text at position t The candidate storage unit is calculated according to the second update gate, the input gate and the candidate storage unit formula, and the candidate storage unit formula is: in, represents the candidate storage unit at position t, represents the second update gate at position t, represents the input gate at position t; The latest storage unit is calculated according to the candidate storage unit, the forget gate and the storage update formula, and the storage update formula is: in, represents the latest storage unit at position t, represents the forget gate at position t, represents the candidate storage unit at position t, represents the latest storage unit at position t-1; The output gate of the knowledge is calculated according to the second feature state and the output gate mechanism formula of LSTM, and the output gate mechanism formula is: in, represents the output gate at position t, represents the sigmoid activation function, W o represents the weight matrix associated with the output gate, b o represents the bias vector associated with the output gate, represents the second characteristic state at position t-1, Represents the vector generated by the second BERT model after processing the input text at position t The second current hidden state of the knowledge is calculated according to the output gate, the latest storage unit and the second hidden formula, where the second hidden formula is: in, represents the second current hidden state at position t, represents the output gate at position t, represents the latest storage unit at position t, tanh represents the hyperbolic tangent function; The first current hidden state and the second current hidden state are input into a dense layer for classification to obtain a predicted probability distribution for each category, a refinement loss is calculated according to the predicted probability distribution, and the first student model and the second student model are trained with the minimum refinement loss as the goal to obtain a trained first student model and a trained second student model.
6. The method for analyzing medical evaluation through adversarial networks using knowledge distillation as claimed in claim 5, characterized in that: The inputting the first current hidden state and the second current hidden state into a dense layer for classification comprises: The first current hidden state and the second current hidden state are input into a dense layer for classification to obtain a prediction score for each category, and a Softmax activation function is used to convert the prediction score into a prediction probability distribution.
7. The method for analyzing medical evaluation through adversarial networks using knowledge distillation as claimed in claim 5, characterized in that: The step of calculating the refinement loss according to the predicted probability distribution, training the first student model and the second student model with the minimum refinement loss as the goal, and obtaining the trained first student model and the trained second student model comprises: The sparse classification cross entropy is used to calculate the refinement loss of each training sample according to the predicted probability distribution, and the total refinement loss is calculated according to the refinement loss and the total refinement loss formula. The total refinement loss formula is: in, represents the total refining loss, represents the refinement loss of the βth training sample, and N represents the total number of training samples; The first student model and the second student model are trained using the Adam algorithm with the goal of minimizing the total refinement loss to obtain a trained first student model and a trained second student model.
8. A system for analyzing medical evaluations through adversarial networks using knowledge distillation, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Emotion analysis method based on domain adversarial training
CN114997175A
Model compression method and system based on meta learning
CN117151173A