Pathological question and answer model training method, pathological question and answer method, electronic equipment and program product
By generating a feature database and extracting feature information, combining a similarity matching model and a question-and-answer model, a pathological question-and-answer model is formed, which solves the problems of high labor costs and incomplete knowledge of traditional pathological knowledge popularization methods, and achieves more comprehensive and efficient acquisition of pathological knowledge.
Patent Information
- Application Number
- CN202510086834.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional way of popularizing pathological knowledge has problems such as high labor costs and insufficient comprehensive knowledge.
By obtaining literature data containing leukemia pathology, generating a feature database, extracting feature information, inputting preset similarity matching model and initial question-and-answer model, iterative training is carried out to form a pathological question-and-answer model to provide pathological knowledge.
It has improved the labor cost and insufficient knowledge of traditional pathological knowledge popularization methods, and provided a more comprehensive and efficient way to obtain pathological knowledge.
Smart Images

Figure CN120179767A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a pathology question-and-answer model training method, a pathology question-and-answer method, an electronic device, and a program product. Background Art
[0002] With the rapid development of AI (Artificial Intelligence) technology, service and knowledge popularization tasks are no longer limited to traditional manual communication and paper material exhibitions. Intelligent, refined services and popular science have become the development and pursuit goals of various industries.
[0003] Acute lymphoblastic leukemia (ALL) is a malignant tumor disease that originates from the abnormal proliferation of B or T lymphocytes in the bone marrow. The abnormally proliferating primitive cells can accumulate in the bone marrow and inhibit normal hematopoietic function, and can also invade tissues outside the bone marrow, such as the meninges, lymph nodes, gonads, liver, etc.
[0004] In order to reduce the workload of medical staff and improve the public's awareness and prevention of ALL, hospitals usually need to provide consulting services and popular science on ALL pathology, prevention, treatment and other related information (pathology, preventive measures, treatment methods, estimated efficacy and other information are summarized as pathology here). In the prior art, pathology consulting services or popular science are usually provided by professional medical staff to patients or their families, or publicized through paper material display boards. However, as people pay more attention to their physical health, the number of consultants has gradually increased and the consulting questions are complicated, resulting in the high labor cost and insufficient knowledge of traditional pathology knowledge popularization methods. Summary of the invention
[0005] In view of this, the purpose of the embodiments of the present application is to provide a pathology question and answer model training method, a pathology question and answer method, an electronic device and a program product, which can improve the problems of high labor costs and insufficient knowledge coverage in traditional pathology knowledge popularization methods.
[0006] In order to achieve the above technical objectives, the technical solutions adopted in this application are as follows:
[0007] In a first aspect, an embodiment of the present application provides a pathology question answering model training method, the method comprising:
[0008] Obtain literature data documenting leukemia pathology;
[0009] Based on the entity relationship between entities in the literature data, a feature database is generated, wherein the feature database includes a question set representing pathological problems and summary information representing answers corresponding to the pathological problems;
[0010] Extract the feature information corresponding to the abstract information according to the feature database, where the feature information includes at least one of word importance feature, semantic relationship feature, semantic information feature, and format feature;
[0011] Input the problem set and the feature information into a preset similarity matching model to obtain a similarity matching result, where the similarity matching result includes the paragraph information in the literature where the target abstract corresponding to the pathological problem in the problem set is located, and the target abstract represents the abstract information corresponding to the feature information with the highest similarity to the pathological problem;
[0012] Input the problem set and the paragraph information into an initial question - answering model to predict answers to the problem set through the initial question - answering model, and obtain a prediction result corresponding to the problem set;
[0013] Iteratively train the initial question - answering model based on the prediction result until the evaluation index of the initial question - answering model meets the preset conditions, and obtain the trained initial question - answering model as a pathological question - answering model.
[0014] Combined with the first aspect, in some alternative embodiments, extracting the feature information corresponding to the abstract information according to the feature database includes:
[0015] Convert the abstract information into word text;
[0016] Extract features from the word text based on a preset feature extraction strategy to obtain the feature information.
[0017] Combined with the first aspect, in some alternative embodiments, converting the abstract information into word text includes
[0018] Call the PKUSEG word - segmentation tool to perform domain word - segmentation on the abstract information and remove the stop words in the abstract information, so as to segment the abstract information into multiple words to obtain the word text.
[0019] Combined with the first aspect, in some alternative embodiments, extracting features from the word text based on a preset feature extraction strategy to obtain the feature information includes:
[0020] Use the TF - IDF statistical strategy to determine the word importance feature, where the word importance feature represents the importance of each word in the word text in the literature data;
[0021] Use the Word2Vec word - vector generation strategy to determine the semantic relationship feature, where the semantic relationship feature represents the semantic relationship between the current word and the context in the word text;
[0022] Using the Sentence Transformer semantic vector generation strategy, determine the semantic information features, where the semantic information features represent the embedding vectors carrying the semantic information of each word in the word text;
[0023] Using the correlation weighting strategy, determine the format features, where the format features represent the lengths of the respective words in the word text.
[0024] In combination with the first aspect, in some alternative embodiments, using the correlation weighting strategy to determine the format features includes:
[0025] Using the TF-IDF statistical strategy and the BM25 statistical strategy to respectively determine the first similarity score corresponding to the TF-IDF statistical strategy and the second similarity score corresponding to BM25;
[0026] Normalize the first similarity score and the second similarity score to obtain the normalized first similarity score and the normalized second similarity score;
[0027] Perform weighted averaging on the normalized first similarity score and the normalized second similarity score to obtain the format features.
[0028] In combination with the first aspect, in some alternative embodiments, before inputting the question set and the feature information into a preset similarity matching model to obtain a similarity matching result, the method further includes:
[0029] Construct an initial similarity matching model;
[0030] Add weight information to different features in the feature information to obtain weighted feature information;
[0031] Divide the weighted feature information into a test set and a training set according to a preset ratio;
[0032] Train the initial similarity matching model through the test set and the training set to obtain a trained initial similarity matching model as the preset similarity matching model.
[0033] In combination with the first aspect, in some alternative embodiments, the initial question-answering model includes the Chines-macbert-large model and the MedBERT-base-wwm-Chinese model;
[0034] Based on the prediction result, perform iterative training on the initial question-answering model until the evaluation index of the initial question-answering model meets a preset condition, and obtain a trained initial question-answering model as a pathological question-answering model, including:
[0035] Obtain labeled data, where the labeled data includes the true answers pre-labeled for the pathological problems in the problem set;
[0036] Take the cross-entropy loss as the loss function of the initial Q&A model, and calculate the loss value between the prediction result and the labeled data through the loss function;
[0037] Based on the loss value, use the AdamW optimizer to update the hyperparameters of the initial Q&A model;
[0038] Loop to obtain labeled data, calculate the loss value between the prediction result and the true answer through the loss function, and based on the loss value, use the AdamW optimizer to update the hyperparameters of the initial Q&A model until the loss function converges, and obtain the trained initial Q&A model as the pathological Q&A model.
[0039] In a second aspect, an embodiment of the present application further provides a pathological Q&A method, and the method includes:
[0040] Obtain question data, where the question data represents the pathological questions input by the user;
[0041] Input the question data into the above-mentioned preset similarity matching model to obtain a paragraph matching result, where the paragraph matching result includes the paragraphs of the literature where the answer corresponding to the question data is located;
[0042] Input the question data and the paragraph matching result into the above-mentioned pathological Q&A model to obtain an answer result, where the answer result includes the starting position of the answer corresponding to the question data.
[0043] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a processor and a memory coupled to each other. The memory stores a computer program. When the computer program is executed by the processor, the electronic device is enabled to execute the above-mentioned pathological Q&A model training method, or when the computer program is executed by the processor, the electronic device is enabled to execute the above-mentioned pathological Q&A method.
[0044] In a fourth aspect, an embodiment of the present application further provides a computer program product, including a computer program, where the computer program implements the above-mentioned pathological Q&A model training method when executed by a processor, or implements the above-mentioned pathological Q&A method when executed by the processor.
[0045] The invention adopting the above technical solution has the following advantages:
[0046] In the technical solution provided by this application, first, literature data recording leukemia pathology is obtained. Then, based on the entity relationships between entities in the literature data, a feature database is generated, and feature information is extracted according to the feature database. Then, the question set and the feature information are input into a preset similarity matching model to obtain a similarity matching result. Then, the question set and the similarity matching result are input into an initial question-answering model to predict answers to the question set through the initial question-answering model, obtaining a prediction result corresponding to the question set. Finally, the initial question-answering model is iteratively trained based on the prediction result until the evaluation index of the initial question-answering model meets the preset conditions, obtaining a trained initial question-answering model as the pathology question-answering model. In this way, by using a large number of leukemia-related literatures as the backend database, when a patient inputs a relevant pathology question, the paragraph where the answer is located is matched through the preset similarity matching model, and then the specific starting position of the answer is determined through the initial question-answering model, and finally the pathology knowledge required by the patient is obtained as the answer. In this way, the problems of high labor costs and incomplete knowledge in the traditional pathology knowledge popularization method are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] This application can be further illustrated by the non-limiting embodiments given in the drawings. It should be understood that the following drawings only show some embodiments of this application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is a structural block diagram of an electronic device provided by an embodiment of this application.
[0049] Figure 2 It is a schematic flowchart of a pathology question-answering model training method provided by an embodiment of this application.
[0050] Figure 3 It is an example diagram of sub-entity classification representation provided by an embodiment of this application.
[0051] Figure 4 It is an example diagram of a multi-level question set provided by an embodiment of this application.
[0052] Figure 5 It is an example diagram of weight information provided by an embodiment of this application.
[0053] Figure 6 It is a structural block diagram of an abstract recall module provided by an embodiment of this application.
[0054] Figure 7 It is an example diagram of the evaluation result of a pathology question-answering model provided by an embodiment of this application.
[0055] Figure 8Schematic flowchart of the pathological Q&A method provided by the embodiments of the present application.
[0056] Figure 9 Block diagram of the structure of the pathological Q&A system provided by the embodiments of the present application.
[0057] Icons: 100 - electronic device; 101 - processor; 102 - memory. Detailed implementation manners
[0058] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that in the description of the drawings or the specification, similar or identical parts are all denoted by the same reference numerals, and the implementation manners not illustrated or described in the drawings are forms known to those of ordinary skill in the art. In the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0059] Please refer to Figure 1 , an electronic device 100 provided by the embodiments of the present application may include a processor 101 and a memory 102. A computer program is stored in the memory 102. When the computer program is executed by the processor 101, the electronic device 100 can execute the corresponding steps in the following pathological Q&A model training method, or when the computer program is executed by the processor 101, the electronic device 100 can execute the corresponding steps in the following pathological Q&A method.
[0060] In this embodiment, the electronic device 100 may be a personal computer, a laptop computer, a cloud server, etc. It is used to obtain literature data recording leukemia pathology. Then, based on the entity relationships between the entities in the literature data, a feature database is generated, and feature information is extracted according to the feature database. Then, the question set and the feature information are input into a preset similarity matching model to obtain a similarity matching result. Then, the question set and the similarity matching result are input into an initial Q&A model to predict answers for the question set through the initial Q&A model, and a prediction result corresponding to the question set is obtained. Finally, the initial Q&A model is iteratively trained based on the prediction result until the evaluation index of the initial Q&A model meets the preset conditions, and the trained initial Q&A model is obtained as the pathological Q&A model.
[0061] Alternatively, the electronic device 100 may also be used to obtain question data and input the question data into the above-mentioned preset similarity matching model to obtain a paragraph matching result. Then, the question data and the paragraph matching result are input into the above-mentioned pathological Q&A model to obtain an answer result, and the answer result includes an answer corresponding to the question data.
[0062] In this embodiment, the processor 101 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 101 may be a general-purpose processor. For example, the processor 101 may be a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0063] The memory 102 may be, but is not limited to, a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc. In this embodiment, the memory 102 may be used to store literature data, a feature database, feature information, a preset similarity matching model, an initial question-and-answer model, a prediction result, a preset condition, a pathological question-and-answer model, question data, a paragraph matching result, an answer result, etc. Of course, the memory 102 may also be used to store a program, and after receiving an execution instruction, the processor 101 executes the program.
[0064] It can be understood that Figure 1 the structure of the electronic device 100 shown in Figure 1 is only a schematic structural diagram, and the electronic device 100 may also include more Figure 1 components than those shown.
[0065] Please refer to Figure 2 , the present application also provides a method for training a pathological question-and-answer model, which can be applied to the above-mentioned electronic device 100 and executed or implemented by the electronic device 100 for each step in the method. Among them, the method for training a pathological question-and-answer model may include the following steps:
[0066] Step 210, obtaining literature data recording leukemia pathology;
[0067] Step 220, generating a feature database based on the entity relationships between entities in the literature data, where the feature map database includes a question set representing pathological questions and summary information representing the answers corresponding to the pathological questions;
[0068] Step 230: Extract the feature information corresponding to the abstract information according to the feature database, where the feature information includes at least one of word importance feature, semantic relationship feature, semantic information feature, and format feature;
[0069] Step 240: Input the question set and the feature information into a preset similarity matching model to obtain a similarity matching result, where the similarity matching result includes the paragraph information in the literature where the target abstract corresponding to the pathological question in the question set is located, and the target abstract represents the abstract information corresponding to the feature information with the highest similarity to the pathological question;
[0070] Step 250: Input the question set and the paragraph information into an initial question-answering model to predict answers to the question set through the initial question-answering model, and obtain a prediction result corresponding to the question set;
[0071] Step 260: Iteratively train the initial question-answering model based on the prediction result until the evaluation index of the initial question-answering model meets the preset conditions, and obtain the trained initial question-answering model as a pathological question-answering model.
[0072] In the above embodiment, first, literature data recording leukemia pathology is obtained. Then, based on the entity relationships between entities in the literature data, a feature database is generated, and feature information is extracted according to the feature database. Then, the question set and the feature information are input into a preset similarity matching model to obtain a similarity matching result. Then, the question set and the similarity matching result are input into an initial question-answering model to predict answers to the question set through the initial question-answering model, and obtain a prediction result corresponding to the question set. Finally, the initial question-answering model is iteratively trained based on the prediction result until the evaluation index of the initial question-answering model meets the preset conditions, and the trained initial question-answering model is obtained as a pathological question-answering model. In this way, by using a large number of leukemia-related literatures as the backend database, when a patient inputs a relevant pathological question, the paragraph where the answer is located in the literature is matched through a preset similarity matching model, and then the specific starting position of the answer is determined through the initial question-answering model, and finally the pathological knowledge required by the patient is obtained as the answer. In this way, the problems of high labor cost and insufficient comprehensive knowledge in the traditional pathological knowledge popularization method are improved.
[0073] The following will elaborate on each step of the pathological question-answering model training method in detail as follows:
[0074] In step 210, the literature data can be literature related to acute lymphoblastic leukemia retrieved by developers or users from various literature websites in advance. Exemplarily, the literature data can be from three core databases, namely CNKI, Wanfang, and VIP, with the time limit from January 1, 2000 to August 6, 2024. The specific retrieval process is as follows: The retrieval formula for CNKI is: "Article Title = Acute Lymphoblastic Leukemia", and the literature retrieval type is "Academic Journals"; the retrieval formula for Wanfang is: "Title = Acute Lymphoblastic Leukemia", and the literature retrieval type is "Journal Papers"; the retrieval formula for VIP is: "Title = Acute Lymphoblastic Leukemia", and the journal range is all journals.
[0075] In this embodiment, the acquisition of literature data can be based on the above retrieval formulas to download relevant literature from various literature database websites in real time and upload it to the processor 101 of the above electronic device 100, facilitating the construction of a database for the subsequent training of the pathological Q&A model; alternatively, the acquisition of literature data can also be that the user downloads relevant literature from various literature database websites based on the above retrieval formulas, pre-stores it in the memory 102 of the above electronic device 100, and during the subsequent training process of the pathological Q&A model, the processor 101 calls it based on the operation instructions issued by the user. The specific manner of acquiring literature data is not specifically limited here.
[0076] In step 220, key entities are identified and classified from the literature related to acute lymphoblastic leukemia (ALL). The key entities focus on five main entity types, including disease types, disease characteristics, genes, prescriptions, and treatment effects. The definitions of these entity types refer to Table 1. To deeply mine and refine these entity types, they are further divided into more specific sub-entity categories. The specific classification information refers to Figure 3 the sub-entity classification table. Through this detailed entity recognition and classification, a comprehensive and multi-level ALL knowledge system is constructed, providing a solid knowledge foundation for the subsequent training of the pathological Q&A model.
[0077] Table 1, Entity Definition Table:
[0078]
[0079] In this embodiment, the construction of the question set can be based on the relationships between entities (the user pre-constructs a hierarchical table of the relationships between entities according to medical experience to constrain the rationality of the random permutation and combination of entities. For example, a treatment effect must be based on a prescription as a prerequisite, and a disease overview must be based on a disease type as a prerequisite, etc.). According to the above entities, a multi-level question set is constructed through random permutation and combination to ensure the diversity and breadth of pathological questions, thereby ensuring the comprehensiveness of subsequent answer prediction. Exemplarily, refer to Figure 4, the levels of the question set can include disease types - disease profiles, genes - diseases - disease profiles, prescriptions - diseases - treatment effects, prescriptions - prescriptions - diseases - treatment effects, prescriptions - genes - diseases - treatment effects, etc., including but not limited to multi-level question sets of binary groups, ternary groups, and quaternary groups.
[0080] In this embodiment, relevant literature on acute lymphoblastic leukemia is retrieved using three core databases, CNKI, Wanfang, and VIP, and the files are batch exported in NoteExpress format. The file content includes key information such as literature titles, authors, keywords, and abstracts, and the target content required by the user in the file is extracted as abstract information.
[0081] In step 230, since the construction of the question set is generated by random permutations between entities, although a wide and comprehensive question set is generated, there are cases where answers cannot be directly obtained from the abstract information. Therefore, it is necessary to construct the connection between the abstract information and different pathological questions in the question set by extracting the feature information in the abstract information.
[0082] In this embodiment, according to the feature database, extracting the feature information corresponding to the abstract information may include:
[0083] Converting the abstract information into a word text;
[0084] Based on a preset feature extraction strategy, performing feature extraction on the word text to obtain the feature information.
[0085] In this embodiment, converting the abstract information into a word text may include
[0086] Invoking the PKUSEG word segmentation tool to perform domain word segmentation on the abstract information and removing the stop words in the abstract information to split the abstract information into multiple words to obtain the word text.
[0087] In this embodiment, PKUSEG(medicine) is a word segmentation tool designed specifically for the medical field, which can use a model trained in a specific field to accurately identify and segment professional terms, significantly improving the quality of word segmentation. Users can further enhance its accuracy through a custom dictionary. In specific implementation, the model for word segmentation can be defined as a dedicated model in the medical field in the PKUSEG tool by setting the parameter model_name = "medicine" in Python.
[0088] In this embodiment, based on a preset feature extraction strategy, performing feature extraction on the word text to obtain the feature information may include:
[0089] Using the TF-IDF statistical strategy, determine the importance feature of the word, where the importance feature of the word characterizes the importance of each word in the literature data in the word text;
[0090] Using the Word2Vec word vector generation strategy, determine the semantic relationship feature, where the semantic relationship feature characterizes the semantic relationship between the current word and the context in the word text;
[0091] Using the Sentence Transformer semantic vector generation strategy, determine the semantic information feature, where the semantic information feature characterizes the embedding vector carrying the semantic information of each word in the word text;
[0092] Using the correlation weighting strategy, determine the format feature, where the format feature characterizes the length of each word in the word text.
[0093] In this embodiment, TF-IDF is a statistical method for measuring the importance of words in literature data, which consists of two parts: term frequency (TF) and inverse document frequency (IDF). Term frequency reflects the frequency of a word's occurrence in a single literature, while inverse document frequency measures the universality of a word across all literatures. The TF-IDF value is calculated by multiplying the two, expressed as: TF-IDF(t, d) = TF(t, d) × IDF(t), where t represents the word, d represents the literature, and TF-IDF(t, d) represents the importance of the word in the literature data, serving as the above-mentioned importance feature of the word.
[0094] Word2Vec is a word embedding technique in NLP (Natural Language Processing). It uses a neural network to learn the low-dimensional vector representation of words, captures the semantic relationships between words, and thus generates the low-dimensional word vector representation of each word, serving as the above-mentioned semantic relationship feature.
[0095] Sentence Transformer is an NLP model based on the Transformer architecture, which is specifically designed to generate embedding vectors that capture the semantic information of text, serving as the semantic information feature.
[0096] In this embodiment, using the correlation weighting strategy to determine the format feature, where the format feature characterizes the length of each word in the word text, may include:
[0097] Using the TF-IDF statistical strategy and the BM25 statistical strategy to respectively determine the first similarity score corresponding to the TF-IDF statistical strategy and the second similarity score corresponding to BM25;
[0098] Normalize the first similarity score and the second similarity score to obtain the normalized first similarity score and the normalized second similarity score;
[0099] Perform a weighted average on the normalized first similarity score and the normalized second similarity score to obtain the format feature.
[0100] In this embodiment, the term frequency-inverse document frequency (TF-IDF) is used to measure the importance of a word for a collection of documents or a single document in the document data. It consists of two parts: TF (Term Frequency), which represents the frequency of a word appearing in a single document, and the calculation formula is where n x,j is the number of times the word x appears in document j, and ∑ k n k,j is the total number of all words in document j; while the inverse document frequency index (IDF) is used to measure the generality of a word in the entire document data, and the calculation formula is where N is the total number of documents in the document data, and N(X) is the number of documents containing the word x. Then the TF-IDF calculation formula:
[0101]
[0102] BM25 is an algorithm widely used in the field of information retrieval to determine the similarity score between a query and a document. "BM" represents Best Match, and the number "25" indicates that this is the 25th iterative version of the algorithm. Q represents the query, q i represents each token in the query, d represents the document to be matched, W i is the weight assigned to each token, and R measures the relevance between the token and the document. The calculation formula:
[0103]
[0104] Normalize the scores of the two methods to between 0 and 1 by dividing by the corresponding maximum score value in their respective strategies. If the maximum score value is 0, the original score is kept;
[0105] Perform a weighted average on the normalized TF-IDF and BM25 scores to obtain the above format feature. During the weighting process, use 0.5 as the same weight value.
[0106] Between step 230 and step 240, the method may further include:
[0107] Construct an initial similarity matching model;
[0108] Add weight information to different features in the feature information to obtain weighted feature information;
[0109] Divide the weighted feature information into a test set and a training set according to a preset ratio;
[0110] Train the initial similarity matching model through the test set and the training set to obtain a trained initial similarity matching model, which is used as the preset similarity matching model.
[0111] In this embodiment, an initial similarity matching model for similarity calculation is pre-constructed (this model can be any convolutional neural network model, deep neural network model, recurrent neural network model, etc. with similarity calculation function), and then each pathological problem in the problem set is input into the initial similarity matching model to obtain the 5 pieces of abstract information with the highest similarity between the feature information and the pathological problem. Then, the matched abstract information is manually screened and cleaned. The correctly matched answer paragraphs are used as positive samples, and the incorrectly matched answer paragraphs are used as negative samples. The ratio of the positive and negative samples after screening and cleaning is 1:2. Then, according to the actual experience of users or developers, the importance of each feature in the feature information is evaluated, and thus weight information is added to each feature to obtain weighted feature information. Exemplarily, referring to Figure 5 ,the weight of the word importance feature is 0.175958, the weight of the semantic relationship feature is 0.032221, the weight of the semantic information feature is 0.105193, and the weight of the format feature is 0.686628. For the weighted feature information, 300 pathological problems and the weighted feature information of the corresponding answer paragraphs are randomly selected from the binary groups, ternary groups, and quaternary groups in the above problem set, and the weighted feature information of the pathological problems and the corresponding answer paragraphs is divided into a training set and a test set according to a preset ratio (such as 7:3, 8:2, etc.). Then, the initial similarity matching model is trained through the training set and the test set to obtain a trained initial similarity matching model, which is used as the preset similarity matching model in step 240 below.
[0112] In this embodiment, the calculation of similarity can be the Euclidean distance, Mahalanobis distance, cosine similarity, etc. between the pathological problem and each feature information.
[0113] Referring to Figure 6, in the above steps 220 to 230, first, according to the entity relationships between entities in the literature data, the literature data is converted into a question set and abstract information (i.e., the literature abstract). Then, the PKUSEG tokenization tool is called to perform domain tokenization on the abstract information, and stop words in the abstract information are removed to preprocess the text of the abstract information and split the abstract information into multiple words. Then, feature information in the abstract information is extracted through TF-IDF, Word2Vec, Sentence Transformer, and TF-IDF+BM25 respectively. Finally, the initial similarity matching model is trained with the feature information and the question set to obtain a preset similarity matching model (i.e., a multi-dimensional similarity fusion model). In this way, through multi-dimensional feature extraction and weighted processing of features in different dimensions, the accuracy of answer passage matching for pathological questions is enhanced.
[0114] Among them, for each question, a correctly matched passage and two incorrectly matched passages can be obtained. The correctness of the passage is used as the output label, where the correct passage is used as the positive sample, and the incorrect passages are used as negative samples. The TF-IDF statistical strategy score, Word2Vec word vector generation strategy score, Sentence Transformer semantic vector generation strategy score, and relevance weighting strategy score corresponding to the matched passage are used as its input features. A matching passage correctness classification model is established using the random forest algorithm, and based on the feature importance ranking generated by the final model, the corresponding weights of different feature extraction strategies are generated:
[0115] y = 0.687*(TF-IDF+BM25)+0.176*TF-IDF+0.105*Sentence Transformer
[0116] +0.032*Word2Vec.
[0117] In step 240, the question set and the feature information corresponding to each abstract information are input into the preset similarity matching model to calculate the similarity between each pathological question in the question set and each piece of feature information through the preset similarity matching model, and the piece of feature information with the maximum similarity is selected as the target feature information. Then, the passage information of the literature where the abstract information corresponding to the target feature information is located is used as the similarity matching result and is used as the output of the preset similarity matching model.
[0118] In step 250, an initial question-answering model is pre-constructed. The initial question-answering model is composed of the Chines-macbert-large pre-trained model and the MedBERT-base-wwm-Chinese pre-trained model. The question set and the paragraph information of the answers corresponding to each pathological question in the question set are input into the initial question-answering model, and the specific start positions start_logits and end positions end_logits of the answers corresponding to each pathological question in the paragraphs where they are located are obtained as the prediction results.
[0119] In this embodiment, both the Chines-macbert-large and MedBERT-base-wwm-Chinese models will generate their respective start_logits and end_logits, which respectively represent the probability distributions of the start and end positions of the answers. The top 20 positions with the highest scores in start_logits and end_logits will be used as the candidate start and end positions of the answers. By traversing all possible combinations of these candidate positions through a double loop, it is ensured that the end position cannot be before the start position, and the answer length cannot exceed 512. Then, the total score of each valid position combination is calculated, and the combination with the highest score is continuously updated and recorded. Finally, the answer range with the highest score composed of the optimal start and end positions will be used as the final output.
[0120] Using Chines-macbert-large and MedBERT-base-wwm-Chinese, the prediction results of the two models are respectively obtained; 0.5 is used as the weight of the two prediction results to linearly combine the prediction scores, and the weighted start_logits and end_logits are obtained. Then, repeat the steps (the top 20 positions with the highest scores in start_logits and end_logits will be used as the candidate start and end positions of the answers. By traversing all possible combinations of these candidate positions through a double loop, it is ensured that the end position cannot be before the start position, and the answer length cannot exceed 512. Then, the total score of each valid position combination is calculated, and the combination with the highest score is continuously updated and recorded. Finally, the answer range with the highest score composed of the optimal start and end positions will be used as the final output).
[0121] In this embodiment, the Chines-macbert-large pre-trained model makes the model better understand the context semantic association by improving the masking strategy and introducing a similar word replacement mechanism. The model adopts the Masked Language Model (MLM) task in the pre-training stage and learns the bidirectional representation of the text by predicting the masked words, thus showing excellent performance in multiple Chinese natural language processing tasks; the MedBERT-base-wwm-Chinese pre-trained model is a professional model pre-trained on a large-scale Chinese clinical text. By pre-training on the text corpus in the medical field, the model not only learns general language knowledge but also obtains rich medical field knowledge representation. It adopts the Whole Word Masking (WWM) strategy, which can maintain the integrity of words when processing Chinese medical terms and improves the model's understanding ability of medical professional vocabulary. In addition, the model integrates tasks such as medical entity recognition in the pre-training process, enhances the grasp of medical concepts and relationships, and makes it particularly suitable for processing text understanding tasks in the medical field.
[0122] In step 260, based on the prediction result, the initial question-answering model is iteratively trained until the evaluation index of the initial question-answering model meets the preset condition, and the trained initial question-answering model is obtained as the pathological question-answering model, which may include:
[0123] Obtain labeled data, where the labeled data includes the true answers pre-labeled for the pathological questions in the question set;
[0124] Take the cross-entropy loss as the loss function of the initial question-answering model, and calculate the loss value between the prediction result and the labeled data through the loss function;
[0125] Based on the loss value, use the AdamW optimizer to update the hyperparameters of the initial question-answering model;
[0126] Loop to obtain labeled data, calculate the loss value between the prediction result and the true answer through the loss function, and based on the loss value, use the AdamW optimizer to update the hyperparameters of the initial question-answering model until the loss function converges, and obtain the trained initial question-answering model as the pathological question-answering model.
[0127] In this embodiment, the start_logits and end_logits of the answer output by the foregoing initial Q&A model are used. These logits represent the scores of each position as the start position and end position of the answer. Then, the true answer annotated for the pathological question is obtained as the annotated data, and the cross-entropy loss is used as the loss function to calculate the loss value between the annotated data and the foregoing prediction result, which is used to measure the matching degree between the model prediction probability distribution (i.e., the prediction result) and the true answer. Then, based on the loss value, the hyperparameters of the initial Q&A model are updated through the AdamW optimizer. In practical applications, the softmax function can be used to perform softmax processing on each logit to obtain the probability distribution, and the position with the highest probability is selected as the start and end positions of the answer. The annotated data, loss value calculation, and hyperparameter update are looped until the loss function converges, and the trained initial Q&A model is obtained as the pathological Q&A model.
[0128] Understandably, in practical applications, the pathological Q&A model obtained through a single complete iterative training may still not meet the user's requirements for the stability of the answer prediction result. Therefore, the ROUGE series of evaluation metrics can be used to evaluate the trained initial Q&A model, and the similarity between the model's predicted answer and the true answer can be measured from different dimensions. This metric includes ROUGE-1, ROUGE-2, and ROUGE-L. Among them, ROUGE-1 is used to measure the overlap of words between the prediction result and the true answer; ROUGE-2 is used to measure the overlap of word pairs between the prediction result and the true answer; ROUGE-L is used to measure the length of the longest common subsequence between the prediction result and the true answer, which helps to evaluate the consistency of the logical structure of the prediction result and the true answer. In this embodiment, the average value of the ROUGE metric values corresponding to multiple pathological questions is used to measure whether the trained pathological Q&A model meets the user's requirements. Exemplarily, referring to Figure 7 , the average value of the ROUGE metric for this training is 0.69, which is less than the preset metric threshold of 0.8. Then, this pathological Q&A model does not meet the user's requirements, and the above step 260 is repeated until the average value of the ROUGE metric of the trained pathological Q&A model is not less than the user's preset metric threshold.
[0129] Please refer to Figure 8 , this embodiment of the present application also provides a pathological Q&A method, which can be applied to the above electronic device 100 and executed or implemented by the electronic device 100 for each step of the method. The method may include:
[0130] Step 310, obtaining question data, where the question data represents a pathological question input by the user;
[0131] Step 320: Input the problem data into the above-mentioned preset similarity matching model to obtain a paragraph matching result, where the paragraph matching result includes the paragraphs of the literature where the answer corresponding to the problem data is located.
[0132] Step 330: Input the problem data and the paragraph matching result into the above-mentioned pathological Q&A model to obtain an answer result, where the answer result includes the answer corresponding to the problem data.
[0133] In this embodiment, first, obtain the pathological problem input by the user as the problem data. In practical applications, the input method of the pathological problem can be text input, voice input, or gesture input. Then, by calling a large language model with voice and gesture conversion functions, convert the audio and video of the voice input and gesture input into text format for subsequent answering. Then, input the problem data into the above-mentioned preset similarity matching model to obtain the paragraphs of the literature where the answer corresponding to the problem data is located as the paragraph matching result. Then, input the problem data and the paragraph matching result into the above-mentioned pathological Q&A model to obtain the starting position of the answer corresponding to the problem data as the answer result. In practical applications, after determining the starting position of the literature where the answer is located, through a simple literature search, the corresponding answer text can be called from the above-mentioned literature data for display, facilitating the user to intuitively understand the answer corresponding to the pathological problem and realizing visual knowledge popularization.
[0134] Please refer to Figure 9 , for the convenience of understanding and implementing the above-mentioned pathological Q&A model training method and pathological Q&A method, the present application uses the pathological Q&A system installed in the electronic device 100 as the implementation environment to elaborate in detail the implementation process of the above-mentioned pathological Q&A model training method and pathological Q&A method as follows:
[0135] 1. Data acquisition
[0136] Conduct a literature search in three databases, namely China National Knowledge Infrastructure (CNKI), Wanfang, and VIP. The time limit is from January 1, 2000, to August 6, 2024. The search process is as follows: The search formula for CNKI is: "Article Title = Acute Lymphoblastic Leukemia", and the literature search type is "Academic Journals"; the search formula for Wanfang is: "Title = Acute Lymphoblastic Leukemia", and the literature search type is "Journal Papers"; the search formula for VIP is: "Title = Acute Lymphoblastic Leukemia", and the journal scope is all journals.
[0137] 2. Knowledge extraction
[0138] Identify and classify key entities from the literature related to acute lymphoblastic leukemia (ALL). The core focuses on five main entity types, including disease categories, disease characteristics, genes, prescriptions, and treatment effects. To deeply explore and refine these entity types, they are further divided into more specific sub-entity categories. Through this detailed entity identification and classification, a comprehensive and multi-level ALL knowledge system is constructed, providing a solid knowledge foundation for the subsequent pathological question-answering system.
[0139] 3. Problem set construction
[0140] When constructing the pathological question-answering system for acute leukemia, a problem set is generated based on entity relationship analysis. By deeply analyzing the internal relationships among key entities such as disease categories, disease characteristics, genes, prescriptions, and treatment effects, a multi-level problem set of binary groups, triple groups, and quadruple groups such as disease category - disease overview, gene - disease - disease category, prescription - disease - treatment effect, prescription - prescription - disease - treatment effect, prescription - gene - disease - treatment effect is constructed. The problem set is automatically generated in a random permutation and combination manner to ensure the diversity and comprehensiveness of the questions.
[0141] 4. Construction of the abstract recall module
[0142] 4.1. Construction of the abstract recall module
[0143] Since the problem set is automatically generated by random permutation among entities, although the questions are extensive, there are cases where answers cannot be directly obtained from the abstract information. Therefore, it is necessary to construct an abstract recall module. The framework of the abstract recall module refers to Figure 6 .
[0144] In this module, by identifying the bibliographic references of each abstract information, the paragraphs of the literature where each abstract information is located are obtained. Specifically, by extracting the characteristic information of each abstract information (i.e., ALL field keywords), and retrieving and calculating the similarity between each pathological question and each characteristic information, the abstract information corresponding to the characteristic information with the highest similarity, as well as the paragraph information (i.e., bibliographic reference information) of the literature where the abstract information is located, are used as the recalled correct bibliographic reference information.
[0145] Specifically, the recall process of the bibliographic reference information is as follows:
[0146] (1) Convert the abstract into text and preprocess the text using word segmentation and stop word removal in the PKUSEG(medicine) domain. PKUSEG(medicine) is a word segmentation tool designed specifically for the medical field. It can use models trained in specific domains to accurately identify and segment professional terms, significantly improving the quality of word segmentation. Users can further enhance its accuracy through a custom dictionary and can set the parameter model_name="medicine" in Python.
[0147] (2) Use the TF-IDF, Word2Vec, Sentence Transformer, TF-IDF+BM25 weighted combination algorithm to extract multiple feature information that matches the question and the abstract. The specific extraction process refers to the description of feature information extraction in step 230 above and will not be elaborated here.
[0148] 4.2. Training and evaluation of the multi-dimensional similarity fusion model (i.e., the preset similarity matching model)
[0149] The system first identifies and recalls the 5 abstracts with the highest matching rate, manually screens and cleans the matched data, uses the correctly matched abstract segments as positive samples, and the incorrectly matched abstract segments as negative samples. After screening, a dataset with a positive-negative sample ratio of 1:2 is formed. Use the random forest algorithm to establish a classification model for the correctness of the matched paragraphs, and generate the corresponding weights of different feature extraction strategies based on the feature importance ranking generated by the final model to train the multi-dimensional similarity fusion model. Among them, the weight of the word importance feature is 0.175958, the weight of the semantic relationship feature is 0.032221, the weight of the semantic information feature is 0.105193, and the weight of the format feature is 0.686628.
[0150] Randomly select 300 pieces of data from the binary groups, ternary groups, and quaternary groups respectively to form a balanced and diverse test set to evaluate the multi-dimensional similarity fusion model. According to the actual training results, the accuracy of the multi-dimensional similarity fusion model reaches 85.45%.
[0151] 5. Pathology Q&A model training
[0152] 5.1. Model selection and training
[0153] Describe the correctly matched abstract segments to generate question-answer pairs, which are used as the dataset for the pathology Q&A model. The model uses Chines-macbert-large and MedBERT-base-wwm-Chinese as the pre-trained models in the pathology Q&A model. The Chines-macbert-large pre-trained model is further optimized based on MacBERT. By improving the masking strategy and introducing a similar word replacement mechanism, the model can better understand the context semantic associations. This model adopts the Masked Language Model (MLM) task in the pre-training stage and learns the bidirectional representation of the text by predicting the masked words, thus showing excellent performance in multiple Chinese natural language processing tasks. The above MedBERT-base-wwm-Chinese pre-trained model is a professional model pre-trained on a large-scale Chinese clinical text. By pre-training on the text corpus in the medical field, this model not only learns general language knowledge but also obtains rich medical domain knowledge representations. It adopts the Whole Word Masking (WWM) strategy, which can maintain the integrity of words when processing Chinese medical terms, improving the model's understanding ability of medical professional vocabulary. In addition, the model integrates tasks such as medical entity recognition during the pre-training process, enhancing the grasp of medical concepts and relationships, making it particularly suitable for processing text understanding tasks in the medical field.
[0154] The basic idea of model training is to input the encodings of pathology questions and corresponding answer paragraphs and output the starting position information of the answers to complete the extraction task. Specifically, the Chines-macbert-large and MedBERT-base-wwm-Chinese models directly output start_logits and end_logits, and these logits represent the scores of each position as the starting and ending positions of the answer. The model uses CrossEntropyLoss to calculate the loss, measure the matching degree between the model's predicted probability distribution and the true labels, and update the parameters through the AdamW optimizer. Finally, the logits are processed by softmax to obtain the probability distribution, and the position with the highest probability is selected as the starting and ending positions of the answer.
[0155] 5.2. Model Evaluation
[0156] Evaluate the performance of the pathological Q&A model. The evaluation metrics adopt the ROUGE series of metrics, including ROUGE-1, ROUGE-2, and ROUGE-L, which can measure the similarity between the model's prediction results and the true answers from different dimensions. The ROUGE-1 metric mainly focuses on the overlap of words between the prediction results and the true answers. ROUGE-2 further considers the overlap of word pairs between the prediction results and the true answers, while ROUGE-L focuses on the length of the longest common subsequence between the prediction results and the true answers, which helps to evaluate the consistency of the logical structure of the predicted answers with the true answers.
[0157] 6. User Interaction Layer
[0158] This pathological Q&A system uses the Flask framework to develop a Web application interface. Users can access the pathological Q&A system through a browser, input the pathological questions they want to query on the interface, and the system will call the Chines-macbert-large and MedBERT-base-wwm-Chinese models in the backend for processing and return the prediction results. According to the starting position of the answers recorded in the prediction results, the corresponding answer texts will be queried from the literature data for display, and functions such as error prompts will also be provided.
[0159] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-mentioned electronic device 100 can refer to the corresponding processes of each step in the foregoing method, and will not be elaborated here too much.
[0160] The embodiment of the present application also provides a computer program product, including a computer program, and the computer program, when executed by the processor 101, implements the above-mentioned pathological Q&A model training method or implements the above-mentioned pathological Q&A method.
[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0162] In summary, the embodiments of the present application provide a method for training a pathological Q&A model, a pathological Q&A method, an electronic device 100, and a program product. In this technical solution, first, literature data recording leukemia pathology is obtained. Then, based on the entity relationships between entities in the literature data, a feature database is generated, and feature information is extracted according to the feature database. Then, the question set and the feature information are input into a preset similarity matching model to obtain a similarity matching result. Then, the question set and the similarity matching result are input into an initial Q&A model to predict answers for the question set through the initial Q&A model, obtaining a prediction result corresponding to the question set. Finally, the initial Q&A model is iteratively trained based on the prediction result until the evaluation index of the initial Q&A model meets the preset conditions, obtaining the trained initial Q&A model as the pathological Q&A model. In this way, by using a large number of leukemia-related literatures as the backend database, when a patient inputs a relevant pathological question, the paragraph where the answer is located is matched through the preset similarity matching model, and then the starting position of the answer is determined through the initial Q&A model, and finally the pathological knowledge required by the patient is obtained as the answer. In this way, the problems of high labor cost and incomplete knowledge in the traditional way of popularizing pathological knowledge are improved.
[0163] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, each functional module in the embodiments of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0164] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A pathology question answering model training method, characterized in that: The method comprises: Obtain literature data documenting leukemia pathology; Based on the entity relationship between entities in the literature data, a feature database is generated, wherein the feature database includes a question set representing pathological problems and summary information representing answers corresponding to the pathological problems; Extracting feature information corresponding to the summary information according to the feature database, wherein the feature information includes at least one of a word importance feature, a semantic relationship feature, a semantic information feature, and a format feature; Inputting the question set and the feature information into a preset similarity matching model to obtain a similarity matching result, wherein the similarity matching result includes paragraph information in the document where the target abstract corresponding to the pathological question in the question set is located, and the target abstract represents the abstract information corresponding to the feature information having the highest similarity to the pathological question; Inputting the question set and the paragraph information into an initial question-answering model, so as to predict answers to the question set through the initial question-answering model, and obtain prediction results corresponding to the question set; The initial question-answering model is iteratively trained based on the prediction result until the evaluation index of the initial question-answering model meets the preset conditions, thereby obtaining a trained initial question-answering model as a pathology question-answering model.
2. The method according to claim 1, characterized in that Extracting feature information corresponding to the summary information according to the feature database includes: Converting the summary information into word text; Based on a preset feature extraction strategy, feature extraction is performed on the word text to obtain the feature information.
3. The method according to claim 2, characterized in that Convert the summary information into word text, including The PKUSEG word segmentation tool is called to perform field word segmentation on the summary information, and stop words in the summary information are removed to segment the summary information into multiple words to obtain the word text.
4. The method according to claim 2, characterized in that: Based on a preset feature extraction strategy, feature extraction is performed on the word text to obtain the feature information, including: Using the TF-IDF statistical strategy, determining the word importance feature, wherein the word importance feature represents the importance of each word in the word text in the document data; Using the Word2Vec word vector generation strategy, the semantic relationship feature is determined, and the semantic relationship feature represents the semantic relationship between the current word and the context in the word text; Determine the semantic information feature by using the Sentence Transformer semantic vector generation strategy, wherein the semantic information feature represents an embedded vector carrying the semantic information of each word in the word text; The format feature is determined by using a relevance weighting strategy, wherein the format feature represents the length of each word in the word text.
5. The method according to claim 4, characterized in that The format features are determined using a relevance weighting strategy, including: Using the TF-IDF statistical strategy and the BM25 statistical strategy, respectively determine a first similarity score corresponding to the TF-IDF statistical strategy and a second similarity score corresponding to the BM25; Normalizing the first similarity score and the second similarity score to obtain a normalized first similarity score and a normalized second similarity score; A weighted average is performed on the normalized first similarity score and the normalized second similarity score to obtain the format feature.
6. The method according to claim 1, characterized in that Before inputting the question set and the feature information into a preset similarity matching model to obtain a similarity matching result, the method further includes: Construct an initial similarity matching model; Adding weight information to different features in the feature information to obtain weighted feature information; Dividing the weighted feature information into a test set and a training set according to a preset ratio; The initial similarity matching model is trained by using the test set and the training set to obtain a trained initial similarity matching model as the preset similarity matching model.
7. The method according to claim 1, characterized in that The initial question-answering model includes a Chinese-macbert-large model and a MedBERT-base-wwm-Chinese model; Iteratively training the initial question-answering model based on the prediction result until the evaluation index of the initial question-answering model meets the preset conditions, and obtaining a trained initial question-answering model as a pathology question-answering model, including: Acquire annotated data, wherein the annotated data includes real answers annotated in advance to the pathological questions in the question set; Using cross entropy loss as the loss function of the initial question-answering model, and calculating the loss value between the prediction result and the labeled data through the loss function; Based on the loss value, update the hyperparameters of the initial question-answering model using the AdamW optimizer; The labeled data is acquired cyclically, the loss value between the predicted result and the true answer is calculated through the loss function, and based on the loss value, the hyperparameters of the initial question-answering model are updated using the AdamW optimizer until the loss function converges, thereby obtaining a trained initial question-answering model as the pathological question-answering model.
8. A pathology question-answering method, characterized in that: The method comprises: Acquire problem data, wherein the problem data represents a pathological problem input by a user; Inputting the question data into a preset similarity matching model according to any one of claims 1 to 7 to obtain a paragraph matching result, wherein the paragraph matching result includes a paragraph of the document where the answer corresponding to the question data is located; The question data and the paragraph matching result are input into the pathology question-answering model as described in any one of claims 1 to 7 to obtain an answer result, wherein the answer result includes a starting position of the answer corresponding to the question data.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory coupled to each other, wherein the memory stores a computer program. When the computer program is executed by the processor, the electronic device executes the method as claimed in any one of claims 1 to 7, or executes the method as claimed in claim 8.
10. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7, or implements the method according to claim 8.