Tumor grading model establishment method capable of being deployed at desktop end and tumor grading system

By adopting the combination of knowledge distillation and large language models in resource-constrained hospital environments, a tumor grading model that can be deployed on the desktop is solved, and the problems of tumor grading in resource-constrained and complex text processing are achieved, and efficient and accurate tumor grading is achieved.

CN120164565APending Publication Date: 2025-06-17XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195208.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

How to achieve efficient and accurate tumor grading in resource-constrained hospital environments, especially under the complexity of pathological texts and the needs of multitasking.

Method used

A knowledge distillation method is used to establish a tumor grading model that can be deployed on the desktop. Through joint training of the teacher model and the student model, a large language model is used to extract named entities in the pathological text and classify them, and an attention mechanism is introduced in the feature extraction stage to improve the accuracy of the model.

Benefits of technology

It realizes efficient and accurate extraction of named entities related to tumor grading from pathological texts and classifications under resource constraints, improving the accuracy and efficiency of tumor grading and reducing the need for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164565A_ABST
    Figure CN120164565A_ABST
Patent Text Reader

Abstract

The invention discloses a tumor grading model establishment method capable of being deployed at a desktop end and a tumor grading system, and belongs to the field of medical text processing, and the method comprises the steps: respectively establishing a teacher model and a student model based on a large language model; the scale of the large language model in the student model is suitable for being deployed at a desktop end, and the large language model in the teacher model has a larger scale and a stronger text understanding ability; obtaining a medical text marked with named entities and classification information related to tumor classification as training data, and training a teacher model and a student model through a knowledge distillation mode; the training loss comprises output layer loss (including identification loss and classification loss of named entities) of the teacher model and the student model and loss among features output by a middle layer of a large language model part; and after the training is finished, taking the student model as a tumor grading model. According to the method, efficient and accurate tumor grading can be realized by means of the large language model under the condition that resources are limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical text processing, and more specifically, relates to a tumor grading model establishment method and a tumor grading system that can be deployed on a desktop. Background Art

[0002] During the treatment of tumors, reasonable and standardized grading of tumors is very important for the formulation of subsequent treatment plans. Pathology reports are the gold standard for tumor diagnosis and therefore have important clinical guiding significance in the diagnosis and treatment of patients.

[0003] Traditional tumor diagnosis mainly relies on manual diagnosis by doctors. Faced with the overloaded clinical pathology workload and limited time, the cumbersome and complex tumor grading standards occupy a lot of working time of pathologists, and the manual grading process may make mistakes. The key to tumor grading is to accurately extract key information related to tumor grading from pathology texts such as pathology reports and electronic medical records. Large Language Models (LLMs) are a revolutionary development in modern artificial intelligence technology. Large language models can capture complex patterns and semantic relationships in language by training on massive text data, thus having a deeper level of language understanding ability. Specifically, in the field of named entity recognition, large language models can not only identify specific entities, such as names, places or professional terms, but also understand the context and extract the relationships and hidden information between these entities. The above characteristics of large language models make large language models promising for efficient and accurate tumor grading.

[0004] However, unlike ordinary text, the entity relationships that need to be extracted in pathological text are more complex, and there may be multiple tumor-related information in the same pathological text. In order to ensure the large language model's ability to understand pathological text, the model scale is often very large. In the hospital environment, the storage and computing resources of diagnostic equipment are very limited, and for privacy protection needs, related calculations are limited to the desktop and cannot be assisted by large servers. This limits the application of large language models in tumor grading.

[0005] Overall, how to achieve efficient and accurate tumor grading under resource-limited conditions remains a difficult problem that needs to be solved urgently in clinical diagnosis. Summary of the invention

[0006] In view of the defects and improvement needs of the prior art, the present invention provides a tumor grading model establishment method and a tumor grading system that can be deployed on a desktop. The purpose is to achieve efficient and accurate tumor grading with the help of a large language model under resource-constrained conditions.

[0007] To achieve the above object, according to one aspect of the present invention, a method for establishing a tumor grading model deployable on a desktop is provided, including:

[0008] Based on a large language model, a teacher model and a student model are respectively established to extract and classify named entities related to tumor grading from pathological texts; the scale of the large language model in the student model is suitable for deployment on a desktop, and the large language model in the teacher model has a larger scale and stronger text understanding ability;

[0009] Obtain medical texts annotated with named entities related to tumor grading and classification information as training data to obtain a training data set;

[0010] Train the teacher model and the student model using the training data set through knowledge distillation, and the training loss function is L total =αL mix +βL intermediate ; after the training is completed, use the student model as the tumor grading model;

[0011] wherein, L total represents the overall loss; L mix represents the output layer loss between the teacher model and the student model, including the recognition loss and classification loss of named entities; L intermediate represents the loss between the features of partial intermediate layer outputs in the large language models of the teacher model and the student model; α and β respectively represent the weights of L mix and L intermediate .

[0012] Furthermore,

[0013]

[0014] wherein, represents a task set, including a named entity recognition task and a named entity classification task; L t represents the loss between the output result of the teacher model for task t and the output result of the student model for task t, and γ t represents the weight of L t .

[0015] Furthermore, the teacher model includes:

[0016] The first large language model is used to identify and classify named entities related to tumor grading in medical texts;

[0017] The student model includes:

[0018] A feature extraction network is used to extract text features from medical texts;

[0019] And a second large language model for identifying and classifying named entities related to tumor grading in medical texts according to the text features extracted by the feature extraction network;

[0020] Among them, the scale of the second large language model is suitable for deployment on the desktop, and the first large language model has a larger scale and stronger text understanding ability.

[0021] Furthermore, the feature extraction network includes:

[0022] An embedding module for tokenizing the pathological text and embedding it into the word vector space to obtain an embedding matrix E;

[0023] A feature extraction module for extracting the text features F of the pathological text according to the embedding matrix E;

[0024] A soft attention module for modeling the context features H according to the text features F and calculating the weights of each token in the context features H;

[0025] And a weighted fusion module for multiplying each token in the context features H by the corresponding weight to obtain the text features F'.

[0026] Furthermore, the feature extraction module includes multiple feature extraction branches and a feature aggregation sub-module;

[0027] Each feature extraction branch includes a convolutional layer and a max-pooling layer connected in sequence, for performing convolutional operations and max-pooling operations on the embedding matrix E in sequence; in different feature extraction branches, the convolutional kernel sizes of the convolutional layers are different;

[0028] The feature aggregation sub-module is used to splice the outputs of each feature extraction branch together to obtain the text features F.

[0029] Furthermore, the soft attention module includes a context modeling layer and an attention weight calculation sub-module connected in sequence;

[0030] The context modeling layer is used to take the text features F as the input sequence and perform context modeling to obtain the context features H;

[0031] The attention weight calculation sub-module is used to calculate the weight of the i-th token in the context features H according to where H i represents the i-th token in H;

[0032] where i ∈ {1, 2, ……, n}, n represents the number of tokens in the context features H; score(H i ) = W a H i + b a ,Wa and b a are learnable parameters.

[0033] Furthermore, the context modeling layer is a long short-term memory network.

[0034] Furthermore, after the training is completed, it also includes: converting the weight parameters of the tumor grading model into 16-bit floating-point numbers.

[0035] According to another aspect of the present invention, there is provided a tumor grading system that can be deployed on the desktop, including:

[0036] An input module for inputting pathological texts;

[0037] A tumor grading module for extracting and classifying tumor grading-related named entities from the pathological texts input by the input module by using the tumor grading model;

[0038] And a grading determination module for processing the output result of the tumor grading model by using a preset regularization algorithm to obtain a tumor grading result;

[0039] Among them, the tumor grading model is established by the above-mentioned method for establishing a tumor grading model that can be deployed on the desktop provided by the present invention.

[0040] Furthermore, the tumor grading system that can be deployed on the desktop provided by the present invention further includes:

[0041] An interaction module for inputting tumor grading-related named entities and classification results;

[0042] And, the grading determination module also uses a preset rule algorithm to process the named entities and classification results input by the interaction module to obtain a tumor grading result.

[0043] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0044] (1) The present invention uses the method of knowledge distillation to establish a tumor grading model that can be deployed on the desktop, and a hierarchical distillation method is proposed during the model training process. Specifically, the designed loss function includes not only the loss of the model output layer but also the loss of the model intermediate layer, so that during the model training process, knowledge distillation is performed not only on the output layer of the model but also on the intermediate layer representation of the model, ensuring that the student model deeply learns the hierarchical semantic information of the teacher model in medical text understanding. As a result, the student model only needs to have a scale that can be deployed on the desktop to have a strong pathological text understanding ability. Finally, the trained tumor grading model can be deployed on the desktop to efficiently and accurately extract and classify tumor grading-related named entities from pathological texts, thereby realizing efficient and accurate tumor grading.

[0045] (2) In tumor pathological texts, multiple tumor-related information may appear simultaneously. In the loss function designed in the present invention, the output layer loss simultaneously includes the loss of the named entity recognition task and the loss of the named entity classification task, enabling the effective fusion of the named entity recognition and entity relationship classification tasks during knowledge distillation. Thus, multi-task distillation is achieved, enabling the student model to better understand the entity relationships in the text and simultaneously possess the processing capabilities of multiple tasks under the limited parameter scale, thereby further improving the accuracy of tumor grading.

[0046] (3) In the model established in the present invention, an attention mechanism is introduced in the text feature extraction stage. During the knowledge distillation process, the attention weights of different tokens in the learned features are learned to reflect the importance of the features of different tokens in the tumor grading problem, and the features are weighted based on the learned attention weights before being input into the large language model. Thereby, the model can more accurately focus on key content when processing text, further improving the accuracy of tumor grading.

[0047] (4) In the model established in the present invention, multiple convolutional layers with different kernel sizes are combined with the max pooling layer to extract features of different scales, which are then concatenated as the original text features, enabling the extracted text features to more accurately reflect the features of tumors. In a further preferred embodiment, a long short-term memory network is used to perform context modeling on the text features, accurately capturing context features and facilitating the improvement of the calculation accuracy of attention weights.

[0048] (5) In traditional model training methods, the weight parameters of the model are mostly 32-bit floating-point numbers. In the final stage of knowledge distillation in the present invention, the weight parameters of the tumor grading model are all converted to 16-bit floating-point numbers, enabling lightweight optimization of the model without significantly affecting the model performance, reducing the model size and computational requirements, and better meeting the requirements of local deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Schematic diagram of a method for establishing a tumor grading model deployable on a desktop provided by an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the structure of a tumor grading model provided by an embodiment of the present invention;

[0051] Figure 3 Block diagram of a tumor grading system deployable on a desktop provided by an embodiment of the present invention;

[0052] Figure 4 Tumor grading process provided by an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the graphical interface in the automatic grading mode provided by the embodiment of the present invention;

[0054] Figure 6 Schematic diagram of the graphical interface in the manual grading mode provided by the embodiment of the present invention. Detailed implementation manners

[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0056] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0057] In order to achieve efficient and accurate tumor grading under resource-constrained conditions, the present invention provides a method for establishing a tumor grading model that can be deployed on a desktop and a tumor grading system. The overall idea is to further propose hierarchical distillation on the basis of knowledge distillation, greatly reducing the requirements for the scale of the student model, so that the student model only needs to have a scale that can be deployed on a desktop to have strong pathological text understanding ability. Finally, the trained tumor grading model can be deployed on a desktop to efficiently and accurately extract and classify the named entities related to tumor grading from pathological texts, thereby achieving efficient and accurate tumor grading.

[0058] The following are embodiments.

[0059] Embodiment 1:

[0060] A method for establishing a tumor grading model that can be deployed on a desktop, as Figure 1 shown, includes:

[0061] Based on a large language model, a teacher model and a student model are respectively established for extracting and classifying the named entities related to tumor grading from pathological texts; the scale of the large language model in the student model is suitable for being deployed on a desktop, and the large language model in the teacher model has a larger scale and stronger text understanding ability;

[0062] Medical texts annotated with the named entities related to tumor grading and classification information are obtained as training data to obtain a training data set;

[0063] Train the teacher model and the student model using the training dataset through knowledge distillation, and the training loss function is L total = αL mix + βL intermediate ; After the training is completed, use the student model as the tumor grading model;

[0064] Among them, L total represents the overall loss; L mix represents the output layer loss between the teacher model and the student model, including the recognition loss and classification loss of named entities; L intermediate represents the loss between the features of partial intermediate layer outputs in the large language models of the teacher model and the student model; α and β respectively represent the weights of L mix and L intermediate .

[0065] In this embodiment, a tumor grading model that can be deployed on the desktop is established by means of knowledge distillation, and a hierarchical distillation method is proposed during the model training process. Specifically, the designed loss function includes both the loss of the model output layer and the loss of the model intermediate layer, so that during the model training process, not only the knowledge distillation of the model output layer is carried out, but also the distillation of the model intermediate layer representation is carried out, ensuring that the student model deeply learns the hierarchical semantic information of the teacher model in medical text understanding. As a result, the student model only needs to have a scale that can be deployed on the desktop to have a strong pathological text understanding ability. Finally, the trained tumor grading model can be deployed on the desktop to efficiently and accurately extract named entities related to tumor grading from pathological texts and classify them, so as to achieve efficient and accurate tumor grading.

[0066] It is easy to understand that in practical applications, the large language model in the teacher model can preferably select a large language model with strong medical text understanding ability, and the large language model in the student model can select a large language model with as high medical text understanding ability as possible under the condition of ensuring that its scale is suitable for deployment on the desktop. Optionally, in this embodiment, the large language model in the teacher model selects QwenLM-72B, and the large language model in the student model selects ChatLM-0.2B. It should be noted that the selection of the large language model here is only an optional implementation manner and should not be understood as the only limitation of the present invention. In some other embodiments of the present invention, other suitable large language models can also be selected according to the requirements of the model scale and text understanding ability.

[0067] It should be noted that for different tumors, different key information needs to be extracted to achieve tumor grading. In practical applications, it can be determined accordingly with reference to specific grading rules. Optionally, in this embodiment, the target tumor is gastrointestinal stromal tumor, and the entities to be extracted from the pathological text include tumor location, tumor size, and mitotic count.

[0068] Considering that in some specific task scenarios, there may be multiple sets of target entities, that is, in the same pathological text, there are relevant information about multiple tumors. Therefore, the model needs to have the processing ability for multiple tasks at the same time. Based on this, in order to perform more effective distillation on the output layer of the model, as an optional implementation manner, in this embodiment, for the calculation of the output layer loss, multi-task distillation is specifically adopted. Specifically, when calculating the output layer loss, the loss L of the named entity recognition task is calculated respectively entity and the loss L of the named entity classification task classification , and weights γ1 and γ2 are assigned to these two losses respectively. Then the expression of the output layer loss L mix can be expressed as:

[0069]

[0070] where represents the task set, including the named entity recognition task and the named entity classification task; L t represents the loss between the output result of the teacher model for task t and the output result of the student model for task t, and γ t represents the weight of L t .

[0071] Optionally, the loss L of the named entity recognition task entity and the loss L of the named entity classification task classification are both calculated using cross-entropy loss.

[0072] In the loss function designed in this embodiment, the output layer loss simultaneously includes the loss of the named entity recognition task and the loss of the named entity classification task, so that the named entity recognition and the entity relationship classification tasks are effectively integrated during knowledge distillation, thereby realizing multi-task distillation, making the output probability distribution of the student model as close as possible to that of the teacher model, and realizing the learning of the prediction results of the teacher model. Finally, the student model can better understand the entity relationships in the text and have the processing ability for multiple tasks while the parameter scale is limited, so as to further improve the accuracy of tumor grading.

[0073] When calculating the intermediate layer loss, it is only necessary to select the same number of intermediate layers from the large language models of the teacher model and the student model respectively, and in descending order, make the feature representations of the corresponding layers of the student model match those of the teacher model. Through intermediate layer distillation, the transfer of deep semantic information can be achieved. Optionally, in this embodiment, the mean squared error is used to calculate the loss of the corresponding intermediate layers in the two models. Accordingly, the intermediate layer loss L intermediate is calculated as follows:

[0074]

[0075] where represents the selected intermediate layer number; represents the feature output by the l-th selected intermediate layer in the large language model of the teacher model, represents the feature output by the l-th selected intermediate layer in the large language model of the student model.

[0076] In this embodiment, the teacher model includes:

[0077] A first large language model for identifying and classifying named entities related to tumor grading in medical texts;

[0078] The student model includes:

[0079] A feature extraction network for extracting text features from medical texts;

[0080] And a second large language model for identifying and classifying named entities related to tumor grading in medical texts according to the text features extracted by the feature extraction network;

[0081] wherein, the scale of the second large language model is suitable for deployment on the desktop, and the first large language model has a larger scale and stronger text understanding ability.

[0082] Considering that when realizing tumor grading based on medical texts, the entity relationships to be extracted are often relatively complex. To further improve the accuracy of the large language model in performing named entity recognition and classification tasks, in this embodiment, the feature extraction network in the model is further optimized and designed.

[0083] As Figure 2 shown, in this embodiment, the feature extraction network includes:

[0084] An embedding module for tokenizing the pathological text and embedding it into the word vector space to obtain an embedding matrix E;

[0085] A feature extraction module for extracting text features F of the pathological text according to the embedding matrix E;

[0086] A soft attention module, which is used to model the context feature H based on the text feature F and calculate the weights of each token in the context feature H;

[0087] And a weighted fusion module, which is used to multiply each token in the context feature H by the corresponding weight to obtain the text feature F'.

[0088] Based on the above feature extraction network, in this embodiment, an attention mechanism is introduced in the text feature extraction stage. During the knowledge distillation process, the attention weights of different tokens in the features are learned to reflect the importance of the features of different tokens in the tumor grading problem, and the features are weighted based on the learned attention weights before being input into the large language model. Thus, the model can more accurately focus on the key content when processing text, thereby further improving the accuracy of tumor grading.

[0089] Furthermore, in this embodiment, the feature extraction module includes multiple feature extraction branches and a feature aggregation sub-module;

[0090] Each feature extraction branch includes a convolutional layer and a max pooling layer connected in sequence, which are used to perform convolutional operations and max pooling operations on the embedding matrix E in sequence; in different feature extraction branches, the convolutional kernel sizes of the convolutional layers are different;

[0091] The feature aggregation sub-module is used to splice the outputs of each feature extraction branch together to obtain the text feature F.

[0092] Furthermore, the soft attention module includes a context modeling layer and an attention weight calculation sub-module connected in sequence;

[0093] The context modeling layer is used to use the text feature F as the input sequence to perform context modeling to obtain the context feature H;

[0094] The attention weight calculation sub-module is used to calculate the weight of the i-th token in the context feature H according to where H i represents the i-th token in H;

[0095] where i ∈ {1, 2, ……, n}, n represents the number of tokens in the context feature H; score(H i ) = W a H i + b a , W a and b a are learnable parameters.

[0096] Furthermore, the context modeling layer is a long short-term memory network.

[0097] In this embodiment, multiple convolutional layers with different-sized convolutional kernels are combined with a max pooling layer to extract features of different scales, which are then concatenated as the original text features, enabling the extracted text features to more accurately reflect the characteristics of tumors. On this basis, a long short-term memory network is further used to perform context modeling on the text features, which can accurately capture context features and is conducive to improving the calculation accuracy of attention weights.

[0098] To make the established tumor grading model more compliant with the requirements of local deployment, as a preferred implementation, after the training is completed, this embodiment further includes: converting all the weight parameters of the tumor grading model into 16-bit floating-point numbers.

[0099] In traditional model training methods, the weight parameters of the model are often stored using 32-bit floating-point numbers. In this embodiment, all the weight parameters of the tumor grading model are converted into 16-bit floating-point numbers, which can achieve lightweight optimization of the model without significantly affecting the model performance, reduce the model size and computational requirements, and better meet the requirements of local deployment.

[0100] Finally, the scale of the tumor grading model established in this embodiment is 0.2B, which can be deployed and run on the internal computers of hospitals without additional computing power support. At the same time, local deployment also provides a certain degree of guarantee for patients' personal privacy and hospital data security.

[0101] Inputting the named entities extracted by the tumor grading model and the corresponding classification information into the preset tumor grading rules can obtain the corresponding tumor grading results. As shown in Table 1, it is the information extraction accuracy rate of the tumor grading model established based on this embodiment in the gastrointestinal stromal tumor scenario.

[0102] Table 1 Information extraction accuracy rate in the gastrointestinal stromal tumor scenario

[0103]

[0104] Based on the results shown in Table 1, this embodiment effectively transfers the powerful medical text understanding and entity extraction capabilities of large-scale pre-trained models to smaller-scale models, enabling tumor grading to be achieved with the help of the text understanding and entity extraction capabilities of large language models even under resource constraints; based on the text understanding and entity extraction capabilities of large language models, this embodiment reduces the workload of pathologists in grading, significantly improving the efficiency and accuracy of tumor grading. In addition, due to the use of a generative named entity recognition method, this embodiment has higher recognition accuracy in the case of multiple tumors in a single pathology report compared to traditional machine learning methods. At the same time, this embodiment supports the grading judgment of various tumors with clear grading rules and has good versatility.

[0105] Embodiment 2:

[0106] A tumor grading system that can be deployed on the desktop, as Figure 3 shown, includes:

[0107] An input module for inputting pathological texts;

[0108] A tumor grading module for extracting and classifying tumor grading-related named entities from the pathological texts input by the input module using a tumor grading model; The classification results include the affiliated tumor and index type;

[0109] And a grading determination module for processing the output results of the tumor grading model using a preset regularization algorithm to obtain the tumor grading results;

[0110] Among them, the tumor grading model is established by the method for establishing a tumor grading model that can be deployed on the desktop provided in the above Embodiment 1.

[0111] As a preferred implementation manner, the tumor grading system that can be deployed on the desktop provided in this embodiment further includes:

[0112] An interaction module for inputting tumor grading-related named entities and classification results;

[0113] And, the grading determination module also processes the named entities and classification results input by the interaction module using a preset rule algorithm to obtain the tumor grading results.

[0114] Correspondingly, the tumor grading system that can be deployed on the desktop provided in this embodiment has two working modes, namely the automatic grading mode and the manual grading mode. In the automatic grading mode, inputting the original pathological text can complete tumor grading; In the manual grading mode, inputting the tumor attributes can complete tumor grading, and the corresponding process is as Figure 4 shown.

[0115] To facilitate the interaction between medical staff and the model, the tumor grading system that can be deployed on the desktop provided in this embodiment further includes a graphical interface. In the automatic grading mode, this interface will display the original pathological text, and display the extraction results and grading results of the named entities, and will also highlight the relevant text content in the input pathological text for easy review and confirmation, as Figure 5 shown. In the manual grading mode, this interface allows medical staff to input the corresponding indicators and display the grading results, as Figure 6 shown.

[0116] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for establishing a tumor grading model that can be deployed on a desktop, characterized in that: include: A teacher model and a student model are respectively established based on the large language model to extract and classify named entities related to tumor grading from pathological texts; the scale of the large language model in the student model is suitable for deployment on a desktop, and the large language model in the teacher model has a larger scale and stronger text comprehension capability; Obtaining medical texts annotated with named entities and classification information related to tumor grading as training data to obtain a training data set; The teacher model and the student model are trained using the training data set by means of knowledge distillation, and the training loss function is L total =αL mix +βL intermediate ; After the training is completed, the student model is used as the tumor grading model; Among them, L total Indicates the overall loss; L mix represents the output layer loss of the teacher model and the student model, including the recognition loss and classification loss of named entities; L intermediate represents the loss between the features of some intermediate layer outputs in the large language model in the teacher model and the student model; α and β represent L mix and L intermediate The weight of .

2. The method for establishing a tumor grading model that can be deployed on a desktop according to claim 1, characterized in that: in, represents a set of tasks, including named entity recognition tasks and named entity classification tasks; L t represents the loss between the output of the teacher model for task t and the output of the student model for task t, γ t Indicates L t The weight of .

3. The method for establishing a tumor grading model that can be deployed on a desktop as claimed in claim 2, characterized in that: The teacher model includes: The first language model for identifying and classifying named entities related to tumor grade in medical text; The student model includes: Feature extraction network, used to extract text features from medical text; and a second language model for identifying and classifying named entities related to tumor grading in medical text based on text features extracted by the feature extraction network; Among them, the scale of the second largest language model is suitable for deployment on the desktop, and the first largest language model has a larger scale and stronger text comprehension ability.

4. The method for establishing a tumor grading model that can be deployed on a desktop as claimed in claim 3, characterized in that: The feature extraction network includes: The embedding module is used to segment the pathological text and embed it into the word vector space to obtain the embedding matrix E; A feature extraction module, used for extracting text features F of pathological text according to the embedding matrix E; A soft attention module is used to obtain a context feature H based on the text feature F and calculate the weight of each token in the context feature H; And a weighted fusion module is used to multiply each token in the context feature H by the corresponding weight to obtain the text feature F'.

5. The method for establishing a tumor grading model that can be deployed on a desktop according to claim 4, characterized in that: The feature extraction module includes a plurality of feature extraction branches and a feature aggregation submodule; Each feature extraction branch includes a convolution layer and a maximum pooling layer connected in sequence, which are used to perform convolution operations and maximum pooling operations on the embedding matrix E in sequence; in different feature extraction branches, the convolution kernel size of the convolution layer is different; The feature aggregation submodule is used to splice the outputs of each feature extraction branch together to obtain the text feature F.

6. The method for establishing a tumor grading model that can be deployed on a desktop as claimed in claim 5, characterized in that: The soft attention module includes a context modeling layer and an attention weight calculation submodule connected in sequence; The context modeling layer is used to take the text feature F as an input sequence, perform context modeling, and obtain context feature H; The attention weight calculation submodule is used to calculate the Calculate the weight of the i-th token in the context feature H, H i represents the i-th token in H; Where i∈{1,2,……,n}, n represents the number of tokens of the context feature H; score(H i )=W a H i +b a , W a and b a are learnable parameters.

7. The method for establishing a tumor grading model that can be deployed on a desktop as claimed in claim 6, characterized in that: The context modeling layer is a long short-term memory network.

8. The method for establishing a tumor grading model that can be deployed on a desktop according to any one of claims 1 to 7, characterized in that: After the training is completed, the method further includes: converting the weight parameters of the tumor grading model into 16-bit floating point numbers.

9. A tumor grading system that can be deployed on a desktop, characterized in that: include: Input module, used to input pathological text; A tumor grading module, used for extracting and classifying tumor grading-related named entities from the pathological text inputted by the input module using a tumor grading model; and a grading determination module, which is used to process the output result of the tumor grading model using a preset regularization algorithm to obtain a tumor grading result; Wherein, the tumor grading model is established by the method for establishing a tumor grading model that can be deployed on a desktop according to any one of claims 1 to 8.

10. The tumor grading system deployable on a desktop according to claim 9, characterized in that: Also includes: An interactive module for inputting tumor grade-related named entities and classification results; Furthermore, the grading determination module also processes the named entities and classification results inputted by the interaction module using a preset rule algorithm to obtain a tumor grading result.