Unstructured temporal bone image report semantic tag modeling and extracting method based on large model instruction fine tuning

By constructing the ontology of the temporal bone image report and using the big model instruction fine-tuning method, combined with the low-rank fine-tuning algorithm with gradient optimization, the problems of professional terms addition and association of structured feature extraction of temporal bone image report, the performance improvement of the model and the consumption of video memory resources are solved, and efficient naming entity recognition and relationship extraction are achieved.

CN120011564APending Publication Date: 2025-05-16BEIJING UNIV OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510082344.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing natural language processing methods face problems such as the addition and association of professional terms, the performance improvement of model auxiliary to less labeled data, and the consumption of memory resources for fine-tuning of large models, resulting in unsatisfactory accuracy and recall.

Method used

The semantic tag modeling and extraction of temporal bone image reports is used to construct the ontology of temporal bone image reports, and naming entity recognition and relationship extraction are performed, and the low-rank fine-tuning algorithm with gradient optimization is combined to improve the performance and efficiency of the model.

Benefits of technology

It realizes the extraction of structured data of multi-level relationships from unstructured temporal bone image reports, improves the accuracy and recall of named entity recognition and relationship extraction, and reduces the memory usage and time-consuming of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011564A_ABST
    Figure CN120011564A_ABST
Patent Text Reader

Abstract

The invention discloses an unstructured temporal bone image report semantic tag modeling and extracting method based on large model instruction fine tuning, and relates to the technical field of natural language processing. According to the method, construction of a temporal bone image report mode layer and a data layer is completed. A low-rank fine tuning algorithm based on gradient optimization is provided, the algorithm can improve the effect of structured information extraction, improve the accuracy and the recall rate, reduce video memory occupation and training time consumption of model training in the entity recognition process, finish training of the temporal bone entity recognition model based on the optimization algorithm, and improve the recognition efficiency of the temporal bone entity recognition model. And completing relation extraction by utilizing a hierarchical relation defined by the temporal bone mode layer, and outputting final structured information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a natural language processing technology, in particular to a method for modeling and extracting semantic labels of unstructured temporal bone imaging reports based on large model instruction fine-tuning. Background Art

[0002] Ear diseases such as tinnitus, deafness, and vertigo can easily cause depression and seriously affect people's physical and mental health. Ear CT examination is one of the important technical means to find out the cause of the disease.

[0003] Temporal bone CT can show the fine structures of the middle ear and inner ear, such as mastoid pneumatization, auditory ossicles, facial nerve, Eustachian tube, cochlea and semicircular canals, jugular bulb, sigmoid sinus, etc. It can also be used to understand whether there are soft tissue masses and their location and range, and whether there are congenital abnormalities, such as congenital malformations such as external auditory canal atresia, auditory ossicular malformation, middle ear cavity dysplasia, inner ear malformation, etc. Temporal bone CT examination has certain guiding significance for the classification of chronic otitis media, the choice of surgical approach and procedure.

[0004] The main job of an ear radiologist is to read medical images (CT), find lesions in them and describe them in full and detail in the radiology image report. In reality, there are a large number of medical images to be examined. The interpretation pressure of radiologists has increased dramatically, which has also increased the probability of potential missed diagnosis and misdiagnosis. Based on the routine diagnostic process of radiologists, the study of intelligent auxiliary analysis technology based on temporal bone CT images for full disease diagnosis is expected to effectively reduce the workload of doctors.

[0005] With the development of science and technology, my country's medical field has entered the information age. Under the wave of big data research, medical data is growing explosively, and the informatization of imaging reports and text data mining are attracting more and more attention and research. Temporal bone imaging reports provide a basis and basis for medical staff to make diagnoses and take reasonable measures in clinical treatment. Converting medical information into structured data processing is an effective way to realize the reuse of clinical research data, which has important theoretical significance and practical application value.

[0006] Existing machine learning models in natural language processing usually face the following problems when applied to structured feature extraction of medical imaging reports of temporal bones:

[0007] (1) Different from traditional machine learning models, the terminology in medical reports is highly professional, and many medical terms are not universal concepts. How to add such terms to the existing knowledge base and link them together is a problem that needs to be solved.

[0008] (2) Traditional machine learning models are trained based on large amounts of labeled data sets. However, the workload of labeling medical image reports is large, and doctors from different departments have different understandings and descriptions of diseases. Therefore, how to use less labeled data to assist the model in improving performance is a difficult problem that needs to be solved urgently.

[0009] (3) The existing named entity recognition methods based on large model fine-tuning require more video memory resources when using full fine-tuning. Using low-rank fine-tuning (LoRA) can save video memory, but the accuracy and recall rate indicators are still somewhat lower than those of full fine-tuning.

[0010] The limitations of existing machine learning methods in natural language processing: lack of construction of medical terminology and pattern layers, requiring a large number of medical experts to complete labeling, and the time-consuming clinical data collection process may not fully cover all required entity and relationship types. In addition, due to scale and technical limitations, current models cannot fully understand the complex and diverse medical terms and relationship types, and the structured extraction effect and accuracy do not meet the requirements.

[0011] This paper proposes a method for semantic label modeling and extraction of unstructured temporal bone imaging reports based on large model instruction fine-tuning. We constructed an ontology for temporal bone imaging reports, used a large model fine-tuning method to extract named entities, and extracted relationships between imaging report entities according to hierarchical relationships based on the model layer constructed by the ontology, and output the final structured information. Summary of the invention

[0012] The purpose of the present invention is to adopt a semantic label modeling and extraction method of unstructured temporal bone image report based on large model instruction fine-tuning to solve the problem of temporal bone image report structuring, and complete the construction of temporal bone image report pattern layer and data layer. An algorithm of low-rank fine-tuning based on gradient optimization is proposed, which can improve the effect of structured information extraction, improve accuracy and recall rate, and reduce the memory usage and training time of model training in the entity recognition process. Based on this optimization algorithm, the training of temporal bone entity recognition model is completed, and then the hierarchical relationship defined by the temporal bone pattern layer is used to complete the relationship extraction and output the final structured information.

[0013] The present invention is achieved by adopting the following technical means:

[0014] 1. A method for constructing an ontology for temporal bone imaging reports. The goal of constructing an ontology is to acquire, describe and represent knowledge in related fields, provide a common understanding of knowledge in the field, determine commonly recognized vocabulary in the field, provide field-specific concept definitions and relationships between concepts, provide activities occurring in the field, and the main theories and basic principles in the field, so as to achieve the effect of human-computer communication. Its main uses include information exchange, sharing, interoperability, and reuse, which are used to guide us to conduct cognitive modeling of things that exist in the real world and terms and concepts in the field within a specific field, and define the schema of graph knowledge.

[0015] The ontology consists of five elements, including: ① Classes; ② Relations; ③ Functions; ④ Axioms; ⑤ Instances.

[0016] Level of detail and domain dependency: top-level Ontologies, domain Ontologies, task Ontologies and application Ontologies. The present invention aims to construct domain Ontologies, using a top-down approach.

[0017] The temporal bone image report ontology construction method is divided into 7 steps in total:

[0018] (1) Determine the professional field and scope of the ontology: The temporal bone disease ontology describes factual knowledge (such as the symptoms of temporal bone diseases), empirical knowledge (such as clearly expressing the causal relationship between symptoms, diseases, and causes), and pathological knowledge that describes the pathological process. The temporal bone imaging report ontology is the knowledge about the diagnosis process of temporal bone diseases and is a control and application strategy for the diagnosis of temporal bone diseases.

[0019] (2) Examining the possibility of reusing existing ontologies: Since the field of temporal bone imaging reporting is relatively niche, there is currently no publicly available ontology dataset.

[0020] (3) List the important terms in the ontology: temporal squamosal, petrous part, mastoid process, tympanic cavity, internal auditory canal, auditory ossicles (malleus, incus), outer wall of the cochlea, inner cavity of the cochlea, lateral semicircular canal, posterior semicircular canal, anterior semicircular canal, vestibule and internal auditory canal.

[0021] (4) Define classes and class hierarchies: Use a top-down approach to build a class hierarchy.

[0022] (5) Define the attributes of the class: including side, negation, lesion range, abnormal density, lesion location, etc.

[0023] (6) Define facets of attributes: make more detailed divisions of attributes.

[0024] (7) Create instance: Create a specific instance based on the above definition.

[0025] 2. A Named Entity Recognition Model Based on Large Model Gradient Optimization and Fine-tuning

[0026] This paper proposes a fine-tuning algorithm based on gradient optimization, which improves the LoRA fine-tuning algorithm. Specifically, in the initialization stage, the gradients of a batch of samples are subjected to singular value decomposition (SVD); the matrices A and B are initialized, and the gradient calculation of each step in the training process is aligned with the gradient of the full fine-tuning, aiming to improve the performance of efficient parameter fine-tuning and make it closer to the effect of full fine-tuning. The main process of the algorithm is as follows:

[0027] 1) Model parameter initialization:

[0028] -Select a batch of samples and calculate the initial gradient Where W0 is the initial gradient under full fine-tuning, is the loss function.

[0029] -Perform SVD (singular value decomposition) on the gradient G0 to obtain G0=UΣV.

[0030] -The original weight matrix W is decomposed into two smaller matrices A and B, where the size of A is n×r and the size of B is r×m, where n is the input dimension, m is the output dimension, and r is a number smaller than n and m, called rank, which is set to 8 in the present invention.

[0031] - Initialize the matrix A by taking the first r columns of U and initialize the matrix B by taking the r+1th to 2rth rows of V.

[0032] 2) Perform estimated gradient calculation

[0033]

[0034] in is the loss function. The present invention adopts the cross entropy loss function. η is the learning rate, and its value range is (0,1). A The optimizer that calculates the gradient of matrix A, G B is the optimizer for gradient calculation of matrix B, t is the number of gradient calculations, A t+1 is the result of matrix A after t+1 gradient calculations, B t+1 is the result of matrix B after t+1 gradient calculations

[0035] 3) Modify the optimizer

[0036] -When updating A and B each time, instead of using G directlyA and G B , but calculate the new H A and H B to replace.

[0037]

[0038] Among them, H A H B is the optimized modified by the present invention, X and Y are matrices obtained by solving a specific optimization problem, X=AC, Y=CB, C is a parameter matrix of r×r dimensions, and r is 8.

[0039] -Calculate the optimal solution of matrix C:

[0040]

[0041] 4) Update rule changes

[0042] -Use H A and H B Alternative G A and G B Update A and B:

[0043] A t+1 =A t -ηH A,t (7)

[0044] B t+1 =B t -ηH B,t (8)

[0045] Among them, H A,t H B,t They are the results of t gradient calculations on matrices A and B using the H optimizer.

[0046] 5) Fine-tune the model

[0047] - During fine-tuning, use the modified optimizer to update A and B, but use H when updating A and B A and H B .

[0048] Through the above optimization process, we can obtain a LoRA fine-tuning implementation based on gradient optimization, which uses gradient SVD to initialize A and B in the initialization stage, and modifies the optimizer's update rule during the fine-tuning process to ensure that each update is close to the effect of full fine-tuning. This method can improve the performance of efficient parameter fine-tuning in named entity recognition model training, making it closer to the effect of full fine-tuning, while maintaining low computing and storage costs.

[0049] 3. A multi-level unstructured data relationship extraction method

[0050] A common approach is to use natural language processing techniques, such as dependency parsing, to extract relationships. These techniques can help us identify entities in text and the relationships between them. Another approach is to use machine learning techniques, such as relation extraction models, to automatically learn patterns for extracting relationships. This approach usually requires a large amount of manually annotated data to train the model, but it can quickly extract relationships from a large amount of data after training. In addition, a template-based approach can be used, which uses predefined templates to extract specific types of relationships. This approach is usually simpler, but may not extract all relationships. Finally, a manual extraction method can be used, which manually reviews the data and annotates the relationships. This method is the most accurate, but it takes a long time.

[0051] Since the temporal bone image report relationship is relatively fixed, the present invention adopts a predefined template method to extract and locate the relationship. According to the hierarchical relationship defined in the temporal bone image ontology, the hierarchical relationship is extracted according to side → part → attribute (normal / abnormal), and the triples are extracted to output the final multi-level structured report data.

[0052] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:

[0053] This paper proposes a method for semantic label modeling and extraction of unstructured temporal bone imaging reports based on large model instruction fine-tuning. We constructed an ontology for temporal bone imaging reports, used a large model fine-tuning method to extract named entities, and extracted relationships between imaging report entities according to hierarchical relationships based on the model layer constructed by the ontology, and output the final structured information.

[0054] Features of the present invention:

[0055] 1. An end-to-end complete temporal bone imaging report structured extraction method is proposed. This method can extract structured data containing multi-level relationships from desensitized unstructured temporal bone imaging reports;

[0056] 2. Based on the prior knowledge of temporal bones, the construction of temporal bone imaging report mode layer and data layer was realized;

[0057] 3. An improved named entity recognition method based on large model gradient optimization fine-tuning is proposed. For the temporal bone image report entity naming task, the performance of efficient parameter fine-tuning can be improved in model training, making it closer to the effect of full fine-tuning, while maintaining low computing and GPU memory costs;

[0058] 4. According to the hierarchical relationship defined in the temporal bone image ontology, a simple and efficient method for hierarchical relationship extraction based on template is proposed.

[0059] The following is a detailed description with reference to the accompanying drawings in combination with examples, so as to gain a deeper understanding of the objects, features and advantages of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 , overall flow chart;

[0061] Figure 2 , Example diagram of temporal bone imaging report mode layer;

[0062] Figure 3 , LoRA training principle diagram;

[0063] Figure 4 ,LoRA reasoning process diagram;

[0064] Figure 5 , fine-tuning algorithm flow chart;

[0065] Figure 6 ,BIO annotation example diagram;

[0066] Figure 7 , training process;

[0067] Figure 8 , recognition confusion matrix for different entities. DETAILED DESCRIPTION

[0068] The following is a description of the implementation examples of the present invention in conjunction with the accompanying drawings:

[0069] The overall process of the present invention is as follows Figure 1 As shown:

[0070] 1) Based on the temporal bone imaging report, prior structured templates and temporal bone professional vocabulary, knowledge summary is performed, important features are extracted, and the construction of the temporal bone imaging report ontology library is completed.

[0071] 2) Preprocess and annotate unstructured reports.

[0072] a. Conduct data collection

[0073] b. Clean the content of the imaging report and desensitize the data

[0074] c. Segment and sentence the single data

[0075] d. Preprocess according to the template

[0076] e. Sequence annotation using BIO method

[0077] 3) Using the labeled data, train the named entity recognition (NER) model using a large model gradient optimization fine-tuning method to extract the lesion entities and attributes

[0078] 4) Perform hierarchical relationship extraction, extract triples, and output the final multi-level structured report data.

[0079] Among them, the detailed steps of temporal bone body construction are as follows.

[0080] 1) Determine the professional field and scope of the ontology;

[0081] The temporal bone disease ontology describes: factual knowledge (such as the symptoms of temporal bone diseases), empirical knowledge (such as the clear expression of the causal relationship between symptoms, diseases, and causes), and pathological knowledge that describes the pathological process. The temporal bone imaging report ontology is the knowledge about the diagnosis process of temporal bone diseases, and is a control and application strategy for the diagnosis knowledge of temporal bone diseases.

[0082] 2) Examine the possibility of reusing existing ontologies;

[0083] Since the field of temporal bone imaging reporting is relatively niche, there is currently no public ontology dataset

[0084] 3) List the important terms in the ontology;

[0085] Temporal squamosal, petrous part, mastoid process, tympanic cavity, internal auditory canal, auditory ossicles, malleus, incus, outer wall of cochlea, inner cavity of cochlea, lateral semicircular canal, posterior semicircular canal, anterior semicircular canal, vestibule and internal auditory canal

[0086] 4) Define classes and class hierarchies

[0087] The present invention adopts a top-down approach.

[0088] Top Level:Temporal Bone

[0089] o Sub Level 1: Side

[0090] Left

[0091] Right

[0092] Bilateral

[0093] o Sub Level 2: Localization

[0094] External Auditory Canal

[0095] ■Soft Tissue in EAC

[0096] ■Bony Wall of EAC

[0097] Temporal Bone

[0098] ■Tympanic Portion

[0099] ■Squamous Portion

[0100] ■Petrous Portion

[0101] ■Mastoid Portion

[0102] Tympanic Cavity

[0103] ■Tympanic Membrane

[0104] ■Roof

[0105] ■Tympanic Segment

[0106] ■Conus Tuberculum

[0107] ■Facial Recess

[0108] ■Eustachian Tube

[0109] Ossicular Chain

[0110] Malleus

[0111] ■Incus

[0112] ■Stapes

[0113] ■Malleo-Incudal Joint

[0114] ■Incudo-Stapedial Joint

[0115] Inner Ear

[0116] ■Cochlea

[0117] Vestibule

[0118] ■Semicircular Canals

[0119] Vestibular Aqueduct

[0120] Neuroanatomical Foramina

[0121] ■Facial Nerve Canal

[0122] ■Labyrinthine Segment

[0123] ■Tympanic Segment

[0124] Mastoid Segment

[0125] ■Foramen Ovale of Cochlear Nerve

[0126] Vestibular Nerve Canal

[0127] ■Foramen of Vesalius Minor

[0128] Vessels

[0129] ■Sigmoid Sinus

[0130] ■Internal Jugular Vein

[0131] ■High Position of the Internal Carotid Artery

[0132] Postoperative Changes

[0133] ■Prosthetic Ossicles

[0134] ■Cochlear Implant

[0135] ■Operative Site

[0136] οSub Level 3: Abnormal Density

[0137] Hyperdense

[0138] Hypodense

[0139] Moderate Density

[0140] Bone density (Osteodense)

[0141] o Sub Level 3: Negation Terms

[0142] Not Present

[0143] Clear

[0144] Intact

[0145] 5) Define the attributes of the class;

[0146] Laterality, negation, lesion extent, abnormal density, lesion localization;

[0147] 6) Define facets of attributes;

[0148] Side:

[0149] Facets: left side, right side, bilateral;

[0150] Negation Terms:

[0151] Facets: None, clear, complete, uniform, symmetrical, and acceptable in shape;

[0152] Lesion extent:

[0153] Surface analysis: small amount, large amount, local, limited, diffuse, multiple, strip-shaped, uneven, inward displacement under pressure, low position, widened;

[0154] Abnormal Density:

[0155] Facets: high, soft tissue, low density, medium density, bone density, bone defect (decreased density), thin structure, slender, not shown, poor gasification, (mastoid) mixed type, sclerotic type, widened, thickened, no shape seen, narrow, blunt, blurred, continuity interruption, local defect, enlargement;

[0156] Lesion localization:

[0157] Surfaces: soft tissue of external auditory canal, bones of the walls of the external auditory canal, tympanic part, squamous part, petrous part, mastoid part, prefenestral fissure bone, peri-cochlear bone, mastoid sinus, condyle, tympanic membrane, tympanic roof, tympanic shield, pyramidal eminence, facial recess, Eustachian tube, malleus, incus, stapes, hammer-incus joint, incus-stapedial joint, stapedial tendon, cochlea, vestibule, semicircular canals, vestibular aqueduct, cochlear duct, facial nerve canal (labyrinthine segment, tympanic segment, mastoid segment), cochlear nerve foramen, vestibular nerve canal, monopore nerve canal, sigmoid sinus, jugular vein, high internal carotid artery, artificial ossicles, cochlear implant, surgical area.

[0158] 7) Create an instance

[0159] Create an instance such as Figure 2 Shown

[0160] The present invention adopts the improved LoRA method to perform fine-tuning training of the Qwen2-7B-instruct model. The basic principle of LoRA is to use low-rank decomposition to simulate the change of parameters, so as to achieve indirect training of large models with relatively small parameters. The improved LoRA fine-tuning algorithm performs singular value decomposition (SVD) on the gradients of a batch of samples in the initialization stage; initializes matrices A and B, where A is a matrix of size n×r and B is a matrix of size r×m, and aligns the gradient calculation of each subsequent step and the alignment of full fine-tuning, aiming to improve the performance of efficient parameter fine-tuning and make it closer to the effect of full fine-tuning.

[0161] Although the pre-trained model has a large number of parameters, the intrinsic dimension (Intrinsic Dimension) corresponding to each downstream task is not large. By fine-tuning a very small number of parameters, good results can be achieved in downstream tasks.

[0162] like Figure 3 As shown, LoRA draws on the above conclusions and proposes a pre-trained parameter matrix W0∈R n×m , instead of directly fine-tuning W0, we make a low-rank decomposition assumption on the increment:

[0163] The formula is as follows:

[0164] W=W0+AB,A∈R n×r ,B∈R r×m (9)

[0165] In order to make the initial state of LoRA consistent with the pre-trained model, the present invention initializes one of A and B to all zeros, that is, A0 and B0, so that the initial W can be W0. Set W to

[0166] W=(W0-A0B0)+AB (10) A new path is added next to the original pre-trained language model PLM (Pretrained Language Model), and the amount of calculation parameters is reduced by multiplying the A matrix and the B matrix. The first matrix A is responsible for dimensionality reduction, and the second matrix B is responsible for dimensionality increase. The intermediate layer dimension is r, so as to simulate the so-called intrinsic rank. The first A matrix is ​​used to reduce the dimension, and the second matrix B is used to increase the dimension. The intermediate layer dimension is r. The dimension d is reduced to r through fc, and then mapped from r to d (where r is less than d). The present invention is set to 8, so that the matrix calculation process is changed from d×k to d×r+r×d, and the amount of calculation parameters of matrix multiplication is significantly reduced.

[0167] The model usually has N weight matrices, and LoRA is applied to the projection matrices (such as Q, K, V, O) in the self-attention layer, while the MLP module and the structure outside the self-attention layer remain unchanged.

[0168] When training downstream tasks, other parameters of the model are fixed, and only the weight parameters of the two newly added matrices are optimized. The results of the PLM and the newly added pathways are added together as the final result (the input and output dimensions of the pathways on both sides are consistent), h = WX + BAX.

[0169] The weight parameters of the first matrix A will be initialized by the Gaussian function, while the weight parameters of the second matrix B will be initialized to the zero matrix, which can ensure that the newly added pathway BA=0 at the beginning of training has no effect on the model results.

[0170] like Figure 4 As shown, during inference, the results of the left and right parts are added together, h = WX + BAX = (W + BA)X, so just add the trained matrix product BA and the original weight matrix W together as the new weight parameter to replace the original PLM W. For inference, no additional computing resources will be added.

[0171] In order to balance the effects of total fine-tuning, this paper proposes a low-rank fine-tuning method LoRA-GO (Low-Rank Adaptation with Gradient Optimized) based on gradient optimization. The training process is as follows Figure 5 shown.

[0172] In order to quantitatively describe this, we write the optimization formulas for full fine-tuning and LoRA fine-tuning under the SGD optimizer, and the results are

[0173] W t+1 =Wt -ηG t (11)

[0174] and

[0175]

[0176] in is the loss function, η is the learning rate, and as well as W t is the parameter matrix after t gradient calculations

[0177] y=W′x=(W0+ηBA)x, where y is the outermost output, x is the original input, W′ is the parameter matrix calculated by the gradient after optimization, and the gradients of matrices A and B are linear mappings of the gradient of W′:

[0178]

[0179] It is worth noting that at the beginning of training, the and full amount of fine-tuning are equal.

[0180] For the gradient in LoRA:

[0181]

[0182] At the beginning of training, both LoRA and full fine-tuning have y′=y and the same x, so:

[0183]

[0184] in, is the change in the gradient calculation parameter matrix after optimization of the present invention, is the parameter matrix change of the full gradient calculation, and They are the changes in the gradient calculation after the A and B parameter matrices are optimized.

[0185] At the beginning of training, the gradients of the low-rank matrices A and B in LoRA can be expressed as a linear mapping of the gradient of W′, and at this time the gradients of LoRA and full fine-tuning are the same, so that the fine-tuning method proposed in the present invention can better approximate the effect of full fine-tuning.

[0186] The core idea of ​​improving the algorithm is to make full fine-tuning and LoRA's W t Close, so minimize the objective:

[0187]

[0188] in is the square of the Frobenius norm of the matrix, that is, the sum of the squares of each element of the matrix.

[0189] The optimal solution can be obtained by performing SVD optimizer calculation on G0, so that we can find the optimal A0 and B0 as the initialization of A and B.

[0190] Since the optimizer is based on A t-1 ,B t-1 and their gradients, are not freely adjustable parameters, so it is not possible to directly compare the total fine-tuning with each W of LoRA. t , but the optimizer can be modified. Specifically, A t , B t The update rule is changed to:

[0191] A t+1 =A t -ηH A,t (20)

[0192] B t+1 =B t -ηH B,t (twenty one)

[0193] Among them, H A,t ,H B,t To be determined, but their shapes are consistent with A and B, so we can write:

[0194] W t+1 =W t -A t B t +A t+1 B t+1 ≈W t -η(H A,t B t +A t H B,t ) (twenty two)

[0195] The imaging reports used in the present invention were collected from the imaging department of a hospital, totaling 650 cases, of which 600 were used for training and 50 were used for testing. The labeled data adopts the BIO labeling method, and each element is labeled as "BX", "IX" or "O". Among them, "BX" means that the segment where this element is located belongs to type X and this element is at the beginning of this segment, "IX" means that the segment where this element is located belongs to type X and this element is in the middle of this segment, and "O" means that it does not belong to any type.

[0196] Select Label Studio as the annotation tool. Label Studio is an open source data annotation tool that allows users to annotate various types of data such as audio, text, images, videos, and time series through an intuitive and concise interface, and can export them to the formats required by various models. This tool is suitable for processing raw data or enhancing existing training data to obtain more accurate machine learning models. The annotation results are as follows: Figure 7 shown.

[0197] According to the task data size and resource status, the present invention uses the Qwen2-7B-instruct model as the basic model. The model has 7 billion parameters ("7B" stands for 7 billion), which makes it quite capable in processing natural language understanding and generation tasks. It has been trained to perform a variety of natural language processing tasks, such as text generation, question and answer, dialogue, text summarization, etc. The model is marked with "instruct", indicating that it has been specially fine-tuned to better follow instructions and produce expected answers. The experimental environment is shown in Table 1.

[0198] Table 1 Experimental hardware environment

[0199] model Intel(R)Xeon(R)Platinum 8358P CPU@2.60GHz Number of cores 64 System Memory 1008GB GPU driver version 535.54.03 CUDA Version 12.2 Number of GPUs 1 GPU 0 NVIDIA A800-SXM4-80GB

[0200] This task was trained for 70 steps in total. The training process is as follows: Figure 7 shown.

[0201] The experimental results are shown in Explanatory Table 2, Explanatory Table 3, Explanatory Table 4, and Explanatory Table 5, which compare the named entity recognition results using different algorithms. Figure 8 Confusion matrix diagram for the model's recognition of different entities.

[0202] For the named entity recognition model, we tested the precision, recall, and F1 value reported in the test dataset, and the results are shown in Table 2. The proposed model LoRA-GO achieved a 16% reduction in error rate compared to the baseline model rule-based entity naming recognition model (Dict-Rule). The dictionary rule-based method relies on a pre-defined vocabulary and has weak recognition capabilities for new words or variant words that are not included in the dictionary. At the same time, Bert-LSTM-CRF, as a well-known entity naming recognition model of the previous generation, is also compared in Table 2. The results show that the improved fine-tuning method based on the large model (LLM) LoRA-GO has a 4.6% improvement over the previous generation model, mainly because LLM is pre-trained on a large text corpus and has rich language knowledge and context understanding capabilities.

[0203]

[0204]

[0205] Table 2 Accuracy of our method and baseline method in the task of named entity recognition in image reports

[0206] The comparative test tasks evaluated the basic LoRA fine-tuning method (LoRA-base), the full fine-tuning method (Full-SFT) and the low-rank fine-tuning method based on gradient optimization LoRA-GO (Low-Rank Adaptation with GradientOptimized).

[0207] The accuracy test results are shown in Table 3: The LoRA-GO fine-tuning method proposed in this paper has achieved a 2.05% improvement in F1 compared with the basic LoRA fine-tuning method for the entity naming recognition model (Dict-Rule), and a 1.15% improvement in F1 compared with the LoRA-GA method, and has achieved an effect close to full-scale fine-tuning (Full-SFT).

[0208] Table 3 Performance of different fine-tuning methods

[0209]

[0210] The basic model of the experiment adopts the Qwen2-7B-instruct model. Table 3 compares the video memory usage, training time, and initialization time of the three fine-tuning methods. The LoRA fine-tuning method has greatly reduced the video memory usage compared with the full fine-tuning, and the LoRA fine-tuning method after gradient optimization (LoRA-GO) has reduced the video memory usage by 4.4G compared with the basic LoRA fine-tuning method; in terms of training time, LoRA-GO has greatly reduced the training time compared with the full fine-tuning and basic LoRA fine-tuning methods; in terms of initialization time, due to the change in the initialization process of the basic LoRA, LoRA-GO initialization takes up more time, but since the initialization is only loaded once during the training process, the overall time consumption is very small compared to the training time and can be ignored.

[0211] Table 4 Training performance of different fine-tuning methods

[0212]

[0213] Table 5 shows the recognition accuracy of the 10 entity categories. It can be seen that the number of different types of entities varies, but the accuracy of the 10 types of entities is less affected by the imbalanced data distribution. The accuracy of the three types of entities, NEUROPORE (neural foramen), OSSICULAR (ossicular chain), and SCLEROTIN (temporal bone), is lower because these entities have more diverse and complex expressions, which makes them more difficult for the model to learn.

[0214] In order to analyze the errors of entity recognition in depth, this paper calculates the classification confusion matrix of 11 categories (10 entities and 1 non-entity) as follows Figure 8 As shown in the figure. The Y axis represents the correct category, and the X axis represents the prediction result of the model. Through the confusion matrix, we can understand the misclassification of each category. From the confusion matrix, we can see that the three entities NEUROPORE (neural pore), OSSICULAR (ossicular chain) and ABNORM (abnormal) are the easiest for the model to confuse with the non-entity O. Analyzing the test data, the reason is that the appearance of these three entities in the real report is more complicated.

[0215] Among them, SIDE means side, BLOODVESSEL means blood vessel, INNEREAR means inner ear, MEATUSES means external auditory canal, NEUROPORE means neural foramen, OSSICULAR means ossicular chain, SCLEROTIN means temporal bone, TYMPANICUS means tympanic wall, NORM means normal, and ABNORM means abnormal density.

[0216] Table 5 The performance measure by entitytype

[0217]

Claims

1. A method for modeling and extracting semantic labels of unstructured temporal bone image reports based on large model instruction fine-tuning, characterized by: 1). Construction of an ontology for temporal bone imaging reports The ontology consists of five elements, including: ① Classes; ② Relations; ③ Functions; ④ Axioms; ⑤ Instances; Level of detail and domain dependency: top-level Ontologies, domain Ontologies, task Ontologies, and application Ontologies; What needs to be constructed is the domain ontology, which adopts a top-down approach; The temporal bone image report ontology construction method is divided into 7 steps in total: (1) Determine the professional field and scope of the ontology: the temporal bone disease ontology describes factual knowledge including the symptoms of temporal bone diseases, empirical knowledge including the pathological knowledge that clearly expresses the causal relationship between symptoms, diseases, and causes, and describes the pathological process; the temporal bone imaging report ontology is the knowledge about the diagnostic process of temporal bone diseases; (2) Examine the possibility of reusing existing ontologies: (3) List the important terms in the ontology: temporal squamosal, petrous part, mastoid process, tympanic cavity, internal auditory canal, auditory ossicles (malleus, incus), outer wall of cochlea, inner cavity of cochlea, lateral semicircular canal, posterior semicircular canal, anterior semicircular canal, vestibule and internal auditory canal; (4) Define classes and class hierarchies: Use a top-down approach to build a class hierarchy; (5) Define the attributes of the class: including side, negation, lesion extent, abnormal density, and lesion location; (6) Define facets of attributes: make finer divisions of attributes; (7) Create instance: Create a specific instance based on the above definition; A fine-tuning algorithm based on gradient optimization is proposed, and the process is as follows: 1) Model parameter initialization: -Select a batch of samples and calculate the initial gradient - Perform SVD singular value decomposition on the gradient G0 to obtain G0 = UΣV; -The original weight matrix W is decomposed into two matrices A and B, where the size of A is n×r and the size of B is r×m, where n is the input dimension, m is the output dimension, and r is a number smaller than n and m. It is called rank and is set to 8; - Initialize the matrix A by taking the first r columns of U, and initialize the matrix B by taking the r+1th to 2rth rows of V; 2) Perform estimated gradient calculation in is the loss function, using the cross entropy loss function, η is the learning rate, the value range is (0,1), G A The optimizer that calculates the gradient of matrix A, G B is the optimizer for gradient calculation of matrix B, t is the number of gradient calculations, A t+1 is the result of matrix A after t+1 gradient calculations, B t+1 is the result of matrix B after t+1 gradient calculations 3) Modify the optimizer - Every time A and B are updated, a new H is calculated A and H B to replace; Among them, H A H B is the modified optimizer, X and Y are matrices obtained by solving a specific optimization problem, X = AC, Y = CB, C is an r × r dimension parameter matrix, r is 8; - Calculate the matrix C in H A and H B The optimal solution under the optimizer: 4) Update rule changes -Use H A and H B Alternative G A and G B Update A and B: A t+1 =A t -ηH A,t (7) B t+1 =B t -ηH B,t (8) Among them, H A,t H B,t They are the results of t gradient calculations on matrices A and B using the H optimizer. 5) Fine-tune the model - During fine-tuning, use the modified optimizer to update A and B, but use H when updating A and B A and H B ; A multi-level unstructured data relation extraction The predefined template method is used for relationship extraction and positioning; according to the hierarchical relationship defined in the temporal bone image ontology, the hierarchical relationship is extracted according to side → location → normal / abnormal attributes, and the triples are extracted to output the final multi-level structured report data.

2. The method according to claim 1, characterized in that: The improved LoRA fine-tuning algorithm performs singular value decomposition (SVD) on the gradients of a batch of samples in the initialization stage; initializes matrices A and B, where A is a matrix of size n×r and B is a matrix of size r×m Propose a pre-trained parameter matrix W0∈R n×m , is to make a low-rank decomposition assumption on the increment, and the formula is as follows: W=W0+AB,A∈R n×r ,B∈R r×m (9) In order to make the initial state of LoRA consistent with the pre-trained model, one of A and B is initialized to zero, that is, A0 and B0. The initial W is W0; set W to W=(W0-A0B0)+AB (10) A new path is added next to the original pre-trained language model PLM. The number of calculation parameters is reduced by multiplying the A matrix and the B matrix. The first matrix A is responsible for dimensionality reduction, and the second matrix B is responsible for dimensionality increase. The intermediate layer dimension is r. The first matrix A is used to reduce the dimension, and the second matrix B is used to increase the dimension. The intermediate layer dimension is r. The dimension d is reduced to r through fc, and then mapped from r to d. Therefore, the matrix calculation process is changed from d×k to d×r+r×d. The model has N weight matrices, and LoRA is applied to the projection matrices in the self-attention layer, including Q, K, V,), while the MLP module and the structure outside the self-attention layer remain unchanged; The weight parameters of the first matrix A will be initialized by the Gaussian function, while the weight parameters of the second matrix B will be initialized to the zero matrix. During inference, the results of the left and right parts are added together, h = WX + BAX = (W + BA) X, so just add the trained matrix product BA and the original weight matrix W together as the new weight parameter to replace the original PLM W. A low-rank fine-tuning method LoRA-GO based on gradient optimization is proposed as follows: Write down the optimization formulas for full fine-tuning and LoRA fine-tuning under the SGD optimizer respectively, and the results are IN t+1 =In t -ηG t (11) and in is the loss function, η is the learning rate, and as well as W t is the parameter matrix after t gradient calculations; y=W′x=(W0+ηBA)x, where y is the outermost output, x is the original input, W′ is the parameter matrix calculated by the gradient after optimization, and the gradients of matrices A and B are linear mappings of the gradient of W′: It is worth noting that at the beginning of training, the and full amount of fine-tuning are equal; For the gradient in LoRA: At the beginning of training, both LoRA and full fine-tuning have y′=y and the same x, so: in, is the change in the optimized gradient calculation parameter matrix, is the parameter matrix change of the full gradient calculation, and They are the changes in the gradient calculation after the optimization of the A and B parameter matrices; At the beginning of training, the gradients of the low-rank matrices A and B in LoRA are expressed as linear mappings of the gradients of W′, and the gradients of LoRA and full fine-tuning are the same at this time; Minimize the objective: in is the square of the Frobenius norm of the matrix, that is, the sum of the squares of each element of the matrix; The optimal solution is obtained by calculating the SVD optimizer on G0, so that we can find the optimal A0 and B0 as the initialization of A and B; A t , B t The update rule is changed to: A t+1 =A t -ηH A,t (20) B t+1 =B t -ηH B,t (21) Among them, H A,t ,H B,t To be determined, but their shapes are consistent with A and B, write: W t+1 =W t -A t B t +A t+1 B t+1 ≈W t -η(H A,t B t +A t H B,t ) (22)。

Citation Information

Patent Citations

  • Temporal bone key anatomical structure small target segmentation method based on 3D deep supervision mechanism

    CN110544264A

  • Temporal bone key anatomical structure automatic positioning method based on spatial relative position prior

    CN112419330A

  • Man-machine combined coronary artery blood vessel image reporting system based on knowledge graph

    CN117409929A

  • Physical examination report iconography examination auxiliary interpretation method and device

    CN118170892A

  • Medical knowledge relation extraction method and system based on large language model fine tuning and retrieval enhancement generation

    CN118569263A