Artificial Intelligence-Based Endoscopic Report Quality Control Analysis Method and Device
Through the endoscopic report quality control analysis method based on artificial intelligence, events in the endoscopic text report are detected and extracted, formatted output and edited distance calculation are performed, and the problems of inefficiency and subjective artificial quality control in the existing technology are solved, and efficient and accurate quality control scores are achieved.
Patent Information
- Application Number
- CN202510127509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In the prior art, the method of manually implementing endoscopic reporting quality control is inefficient and subjective, and is prone to human error.
Using the endoscopic report quality control analysis method based on artificial intelligence, the preset endoscopic report event detection model and event extraction large model is used to detect and extract disease description events and site description events on the endoscopic text report, build a prompt word template, perform event extraction and format output, and finally calculate the quality control score based on the editing distance.
Accurate and quantifiable quality control scores for endoscopic reports are achieved, the efficiency and reliability of quality control analysis of endoscopic reports are improved, and the occurrence of human errors is reduced.
Smart Images

Figure CN119558315B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an endoscopic report quality control analysis method and device based on artificial intelligence. Background Art
[0002] The endoscopic examination report, also known as the endoscopic report, is an important part of the output in endoscopic examinations. The integrity of the endoscopic report and the quality control of the output endoscopic report require medical staff to manually review and judge. In the entire endoscopic examination, the review and ensuring the completeness and accuracy of the endoscopic report may account for 50% of the entire endoscopic examination process. This not only consumes a large amount of the time and energy of endoscopic physicians, but also the final quality control score is subjective and non-quantifiable, and there are often errors caused by human factors.
[0003] It can be seen that the method of manually implementing endoscopic report quality control in the prior art is both inefficient and subjective, and human errors are inevitably introduced during the implementation process. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed to provide an endoscopic report quality control analysis method and device based on artificial intelligence to overcome the above problems.
[0005] In one aspect of the present invention, an endoscopic report quality control analysis method based on artificial intelligence is provided. The method includes:
[0006] Using a preset endoscopic report event detection model to detect disease description events and / or location description events in the target endoscopic text report to be quality controlled, and extracting target text fragments corresponding to the detected disease description events and / or location description events from the target endoscopic text report;
[0007] Constructing a prompt template in the way of a chain of thought. The prompt template includes a task description, input format constraints and input text examples, as well as output format constraints and output text examples. The task description is used to guide the perspective of the large model in task execution and the thinking structure for describing the task analysis process;
[0008] Inputting the extracted target text fragments and the prompt template into a preset endoscopic report event extraction large model, so that the endoscopic report event extraction large model analyzes each target text fragment based on the prompt template, and extracts an event trigger word for the corresponding event, an argument role for identifying the quality control index corresponding to the event, and an argument entity for describing each argument role from each target text fragment, and formatting and outputting the extraction results of each event to obtain a structured text report;
[0009] Calculate the edit distance between the formatted output result of each event in the target endoscopic text report and the standard structured report text corresponding to the same event type preset, and determine the quality control score of the target endoscopic text report according to the edit distance.
[0010] Further, the prompt word template further includes a description of the text standardization rule for the output result;
[0011] The formatting output of the extraction results of each event includes: performing text standardization operations on the extraction results of each event according to the text standardization rule, and formatting the extraction results after the text standardization operation in a preset json format.
[0012] Further, determining the quality control score of the target endoscopic text report according to the edit distance includes:
[0013] Use the following scoring model to determine the quality control score score of the target endoscopic text report, and the scoring model includes:
[0014] score = r×100
[0015]
[0016] where sum is the total character length of the formatted output result of the current event and the corresponding standard structured report text, represents the edit distance between the formatted output result of the current event and the corresponding standard structured report text.
[0017] Further, the method further includes:
[0018] Perform a set operation on the formatted output result of each event in the target endoscopic text report and the corresponding standard structured report text to detect missing fields in the target endoscopic text report. The set operation model is as follows:
[0019]
[0020] where the fields corresponding to set C are the missing fields in the target endoscopic text report, the fields corresponding to set B are the argument role fields in the standard structured report text, and the fields corresponding to set A are the argument role fields in the formatted output result of the event;
[0021] Generate a quality control analysis report according to the quality control score and missing fields of the target endoscopic text report.
[0022] Further, the method further includes:
[0023] The quality control analysis report is displayed in real time through the AI quality control assistant plug-in program pre-installed in each digestive endoscopy image-text system workstation to remind to supplement the missing fields in the quality control analysis report;
[0024] After receiving the supplementary data for the missing fields, upload the supplementary data so that the cloud recalculates the quality control score and generates a quality control analysis report, and displays it again through the AI quality control assistant plug-in program.
[0025] Furthermore, the task description in the prompt template includes the following thinking tasks executed in sequence:
[0026] Identify the event trigger words in the text, and the event trigger words include: verbs, verb phrases or preset specific nouns;
[0027] Extract the argument roles used to identify the corresponding quality control indicators for each event, and the argument roles include: disease name, disease occurrence site, number of diseases, disease size, disease classification and mucosal status;
[0028] Label the event type for each event;
[0029] Structured output includes the trigger words, argument roles and their relationships of each event.
[0030] Furthermore, the endoscopic report event detection model is obtained through the following methods:
[0031] Construct an endoscopic text report dataset;
[0032] According to the preset event types and the labels corresponding to each event type, perform event annotation on the disease description events and / or site description events in each endoscopic text report in the endoscopic text report dataset, and construct an event detection training dataset according to the event annotation results;
[0033] Fine-tune the pre-trained RoBERTa model using the event detection training dataset to obtain an endoscopic report event detection model.
[0034] Furthermore, the endoscopic report event extraction large model is obtained through the following methods:
[0035] Use the endoscopic report event detection model to detect the disease description events and / or site description events in each endoscopic text report in the endoscopic text report dataset, and extract the corresponding sample text fragments from each endoscopic text report for the detected disease description events and / or site description events;
[0036] Annotate the sample text segments in each endoscopic text report according to the event trigger words corresponding to events of different event types, the argument roles used to identify the quality control indicators corresponding to the events, and the argument entities used to describe each argument role, and construct an event extraction training data set according to the information annotation results;
[0037] Fine-tune the pre-trained Qwen2.5-72B large language model using the event extraction training data set to obtain a large model for endoscopic report event extraction.
[0038] Another aspect of the present invention provides an endoscopic report quality control analysis device based on artificial intelligence, including a memory, a processor, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, it implements the steps of the artificial intelligence-based endoscopic report quality control analysis method described in any one of the above.
[0039] Another aspect of the present invention provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based endoscopic report quality control analysis method described in any one of the above.
[0040] The artificial intelligence-based endoscopic report quality control analysis method and device provided by the embodiments of the present invention use the sequence annotation model RoBERTa in natural language processing to detect disease description events and / or location description events in the text report, and output the event extraction results of each event based on the large language model Qwen2.5-72B and perform formatted output, so as to obtain a reverse-structured text report. Moreover, the present invention calculates the edit distance between the formatted output result of each event in the target endoscopic text report and the standard structured report text corresponding to the same event type preset, and generates a final quality control score based on the edit distance, providing an accurate and quantifiable text report quality control score, effectively improving the efficiency and reliability of endoscopic report quality control analysis, and having very significant advantages in improving the work efficiency and quality control of endoscopic physicians.
[0041] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:
[0043] Figure 1 Flow chart of an artificial intelligence-based endoscopic report quality control analysis method according to an embodiment of the present invention;
[0044] Figure 2 Schematic diagram of the annotation format for sequence annotation of endoscopic text report data in an embodiment of the present invention;
[0045] Figure 3 Schematic diagram of the annotation format for information annotation data of text fragments corresponding to events in an embodiment of the present invention;
[0046] Figure 4 Data in JSON format finally obtained after information annotation of text fragments in an embodiment of the present invention. Detailed implementation manners
[0047] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0048] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined.
[0049] Embodiment 1
[0050] An embodiment of the present invention provides an artificial intelligence-based endoscopic report quality control analysis method. As Figure 1 shown, the artificial intelligence-based endoscopic report quality control analysis method proposed by the present invention includes the following steps:
[0051] S11. Use a preset endoscopic report event detection model to detect disease description events and / or location description events in the target endoscopic text report to be quality controlled, and extract target text fragments corresponding to the detected disease description events and / or location description events from the target endoscopic text report;
[0052] S12. Construct a prompt template in the way of chain of thought. The prompt template includes a task description, input format constraints and input text examples, as well as output format constraints and output text examples. The task description is used to guide the perspective of the large model in task execution and describe the thinking structure of the task analysis process.
[0053] Among them, the task description in the prompt template includes the following sequential thinking tasks:
[0054] Identify the event trigger words in the text. The event trigger words include: verbs, verb phrases or preset specific nouns;
[0055] Extract the argument roles for identifying the quality control indicators corresponding to each event. The argument roles include: disease name, disease occurrence site, disease quantity, disease size, disease classification and mucosal status;
[0056] Annotate the event type for each event;
[0057] The structured output includes the trigger words, argument roles and their relationships of each event.
[0058] S13. Input the extracted target text fragments and the prompt template into a preset large model for endoscopic report event extraction, so that the large model for endoscopic report event extraction analyzes each target text fragment based on the prompt template, and extracts the event trigger words, argument roles for identifying the quality control indicators corresponding to each event, and argument entities for describing each argument role from each target text fragment, and formats the extraction results of each event for output to obtain a structured text report. Specifically, relevant quality control indicators can be set according to the type of examination performed, including 18 quality control indicators required by the National Digestive Medicine Quality Control Center.
[0059] S14. Statistically calculate the edit distance between the formatted output result of each event in the target endoscopic text report and the standard structured report text corresponding to the same event type preset, and determine the quality control score of the target endoscopic text report according to the edit distance. Specifically, after the analysis is completed, a complete quality control analysis report for a single examination will be formed.
[0060] The endoscopic report quality control analysis method provided by the embodiments of the present invention is an artificial intelligence-assisted quality control report generation method based on natural language processing and large language models. Through the system integration module, the endoscopic text report in the digestive endoscopy image-text system is uploaded to the AI quality control server for parsing and analysis. The AI quality control server consists of three parts. The first part is the endoscopic report event detection model, which is implemented based on the sequence labeling model RoBERTa in natural language processing. The embodiments of the present invention propose training for event detection based on RoBERTa to detect events in the text report. The second part is the large model for endoscopic report event extraction, which is implemented based on the large language model Qwen2.5-72B, outputs the event extraction results for each event, and formats the event extraction results. The third part is the quality control scoring module, which calculates the edit distance based on the event extraction results and the preset standard structured report text, and generates the final quality control score based on the edit distance. This deep learning technology through natural language processing and large language models can accurately perform quality control scoring on the text report. The embodiments of the present invention can replace part of the report review work of endoscopic physicians, provide accurate and quantifiable quality control scores for text reports, and obtain reverse-structured text reports, which have very significant advantages in improving the work efficiency and quality control of endoscopic physicians.
[0061] The first part of the AI quality control service is the event detection of endoscopic reports based on natural language understanding. In the embodiments of the present invention, the endoscopic report event detection model based on natural language processing is obtained through the following methods:
[0062] The first step is to construct a data set.
[0063] Construct an endoscopic text report data set. The endoscopic text report data set includes endoscopic text reports corresponding to several historical endoscopic examinations.
[0064] The second step is the annotation of training data.
[0065] According to the preset event types and the corresponding labels for each event type, event annotation is performed on the disease description events and / or site description events in each endoscopic text report in the endoscopic text report data set, and an event detection training data set is constructed according to the event annotation results.
[0066] In this embodiment, first, data annotation is performed. The data annotation format is as Figure 2 shown. Sequence annotation is performed on the endoscopic text report data of each endoscopic examination as Figure 2 shown. For example, in this data, "The mucosa of the descending colon is smooth, the submucosal vascular texture is clear, and the peristalsis is regular" is marked as the "site description" category, and finally it is converted into the format of training data as shown in Table 1:
[0067] Table 1 is the annotation table of the training data for the endoscopic report event detection model
[0068] T descend knot intestine adhere membrane smooth adhere , membrane beneath blood vessel vein texture clear distinct peristalsis , rule Figure 3 Figure 3 Figure 4 L 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
[0069] Here, it is assumed that there are only 2 events. The event categories are defined as: "disease description" and "location description". The corresponding labels for each event type are 0 and 1 respectively. In Table 1, the first row T represents the original text of the text report, and the second row L represents the labels that are finally converted into training data by the annotation tool. Since the label of "location description" is encoded as "1", this text segment in the event detection is marked as "1".
[0070] The reason for constructing the training data in this way is to facilitate the training of the endoscopic report event detection model. The "event detection" task described in the embodiments of the present invention is described as: based on the original endoscopic text report, training a model to extract all the events contained in the text report (the events here are defined as: "disease description" and "location description"), as well as the specific positions of the text segments where these events occur, and the position is accurate to characters. The embodiments of the present invention have annotated approximately 30,000 complete endoscopic examination reports according to this annotation method.
[0071] Step 3: Model structure and training
[0072] Fine-tune the pre-trained RoBERTa model using the event detection training data set to obtain the endoscopic report event detection model.
[0073] The embodiments of the present invention adopt the paradigm of pre-training plus fine-tuning for training in the endoscopic report event detection model. Its basic network structure adopts the Transformer Encoder structure, which is composed of spliced Transformer Encoders.
[0074] Pre-training stage: The structure of RoBERTa is exactly the same as that of the BERT model. During the pre-training stage, dynamic MLM is adopted. Original static mask: When preparing training data in BERT, each sample will only undergo random masking once (so each epoch is repeated), and the same mask is used in each subsequent training step. This is the original static mask, that is, a single static mask, which is the practice of the original BERT. RoBERTa adopts dynamic masking: Instead of performing masking during preprocessing, masks are dynamically generated each time the input is provided to the model, so it changes constantly. In addition, the NSP training strategy is removed during pre-training. The NSP training strategy in BERT is SEGMENT-PAIR + NSP: This is the practice of the original BERT. The input consists of two parts, and each part is a segment (a segment is a continuous sequence of multiple sentences) from the same document or different documents. The total number of tokens in these two segments is less than 512. RoBERTa does not adopt the NSP training strategy. In addition, during the pre-training stage, RoBERTa uses a larger Batch Size and more pre-training datasets. In the embodiments of the present invention, pre-training is not required, and only the pre-trained model file needs to be loaded.
[0075] Fine-tuning stage: Based on the labeled event detection training dataset, the pre-trained RoBERTa model is loaded for fine-tuning. The input during the fine-tuning process includes three parts. One part is the encoding of the original Tokens, one part is the SegmentEmbedding to distinguish between sentences, and one part is the Posititon Embedding to encode the positional relationship between Tokens in a sentence. The output of the fine-tuning process is the event prediction corresponding to each Token. During the fine-tuning process, the embodiments of the present invention adopt Adapter Tuning. Specifically, the Adapter module is generally added after two fully connected layers in the Transformer module. The structure of each adapter module includes an input layer, an output layer, a lower projection feed-forward layer, an upper projection feed-forward layer, a non-linear layer, and a skip connection from the input to the output. During the training process, generally only the parameters of the lower projection feed-forward layer, the upper projection feed-forward layer, the non-linear layer, and the two normalization layers in the Transformer module of the adapter are adjusted. The working principle of the adapter module is to first project the input d-dimensional feature vector into an r-dimensional vector (r << d) through the lower projection feed-forward layer (a d×r-dimensional matrix), apply the non-linear layer, and then project it back to a d-dimensional vector through the upper projection feed-forward layer (an r×d-dimensional matrix). The essence of the fine-tuning process is a text classification problem. The loss function adopted by the embodiments of the present invention is the cross-entropy loss.
[0076] In the embodiment of the present invention, the second part of the AI quality control service is the event extraction task, which can extract events, trigger words of events, argument roles of events, and arguments from text reports. The endoscopic report event extraction large model for implementing the event extraction task is obtained through the following methods:
[0077] First step, extract sample text fragments.
[0078] Use the endoscopic report event detection model to detect disease description events and / or location description events in each endoscopic text report in the endoscopic text report dataset, and extract sample text fragments corresponding to the detected disease description events and / or location description events from each endoscopic text report.
[0079] Second step, annotation of training data.
[0080] Perform information annotation on the sample text fragments in each endoscopic text report according to the event trigger words corresponding to events of different event types, the argument roles used to identify the quality control indicators corresponding to the events, and the argument entities used to describe each argument role, and construct an event extraction training dataset according to the information annotation results.
[0081] In this embodiment, the annotation format of the data is as Figure 1 shown, as Figure 1 shown. For example, in the text "2 polyps with a maximum size of about 0.001x0.2 cm were found in the ileocecal region, Paris classification", the word "found" is the trigger word of the event "disease description". Among them, "disease occurrence location (role), number of diseases (role), disease size (role), disease classification (role)" are the 4 argument roles of the event "disease description", and the corresponding arguments are: "ileocecal region, 2, 0.01X0.2 cm, Paris classification". These 4 arguments are 4 entities, and the categories of their entities correspond to: "location name, number of diseases, disease size, disease classification". The description of the final annotated data format is as shown in the json format shown. Organize the above annotation data into the data format of supervised fine-tuning of Qwen2.5-72B including prompt templates, input texts, annotation results, and output results: construct an event extraction training dataset for instruction fine-tuning according to the above format.
[0082] Third step, model structure and fine-tuning.
[0083] Use the event extraction training dataset to fine-tune the pre-trained Qwen2.5-72B large language model to obtain an endoscopic report event extraction large model.
[0084] The large model for endoscopic report event extraction in the embodiments of the present invention is implemented based on the large language model Qwen2.5-72B. The large language model is efficiently fine-tuned with parameter instructions by LoRA (Low-Rank Adaptation) for the labeled event extraction training dataset. The basic structure of the large language model depends on the Transformer Decoder. Qwen2.5-72B contains 80 Decoder layers, uses the SwiGLU activation function, adopts GQA to replace MHA, adds learnable bias terms to Q, K, and V in Attention, and uses Byte Pair Encoding (BPE) as the Tokenizer. Based on this model structure, the embodiments of the present invention perform LoRA efficient parameter instruction supervised fine-tuning training. The LoRA technology decomposes the learnable weight matrix into the product of low-rank matrices, reducing the number of parameters, and thus achieving the purpose of reducing hardware resources and accelerating the fine-tuning process.
[0085] Through the above operations, the embodiments of the present invention obtain the fine-tuned Qwen2.5-72B model. This fine-tuned large model is aligned with the event extraction task, especially for the event extraction of endoscopic text reports. Its inference accuracy has a qualitative improvement compared with the original model without fine-tuning, and it also helps to eliminate model hallucinations.
[0086] The method for endoscopic report quality control analysis based on artificial intelligence provided by the embodiments of the present invention uses the sequence labeling model RoBERTa in natural language processing to detect disease description events and / or location description events in the text report, and outputs the event extraction results of each event based on the large language model Qwen2.5-72B and formats the output, so as to obtain a reverse-structured text report. Moreover, the present invention calculates the edit distance between the formatted output result of each event in the target endoscopic text report and the standard structured report text corresponding to the same event type preset, and generates a final quality control score based on the edit distance, providing an accurate and quantifiable quality control score for the text report, effectively improving the efficiency and reliability of endoscopic report quality control analysis, and having very significant advantages in improving the work efficiency and quality control of endoscopic physicians.
[0087] The prompt template in the embodiments of the present invention also includes a description of the text standardization rules for the output results. Further, the specific steps of formatting the extraction results of each event include: performing text standardization operations on the extraction results of each event according to the text standardization rules, and formatting the extraction results after the text standardization operations in a preset json format.
[0088] In a specific embodiment of the present invention, the process of model inference is as follows:
[0089] First, the endoscopic text report is processed by the endoscopic report event detection model RoBERTa to extract the text fragments corresponding to all events in the endoscopic examination text report. These text fragments serve as the input data (Input Text) for the large endoscopic report event extraction model.
[0090] Secondly, the Input Text passes through the Prompt Engine module (prompt engineering) to obtain the Prompt, that is, the prompt template. The structure of the Prompt is as follows:
[0091] Prompt = """
[0092] You are an event extraction model, and your task is to identify and extract information related to events from the following text. Please follow the following requirements:
[0093] 1. Identify the event trigger words in the text (e.g., verbs, phrases, or specific nouns).
[0094] 2. Extract the roles related to each event (including: disease name, disease location, disease quantity, disease size, disease classification, mucosal status, etc.).
[0095] 3. Label the type of each event (there are two categories: disease description, location description).
[0096] 4. Structured output containing the trigger word, role, and their relationships for each event.
[0097] **Input text:**
[0098] {{ Input Text}}
[0099] Input Text
[0100] **Output format:**
[0101] {
[0102] 'Event 1': {
[0103] 'Trigger word': 'Event trigger word',
[0104] 'Event type': 'Event type',
[0105] 'Role': {
[0106] 'Disease name': 'Role content',
[0107] 'Disease location': 'Role content',
[0108] 'Disease quantity': 'Role content',
[0109] 'Disease size': 'Role content',
[0110] 'Disease classification': 'Role content',
[0111] 'Mucosal status': 'Role content',
[0112] }
[0113] },
[0114] 'Event 2': {
[0115] 'Trigger word': 'Event trigger word',
[0116] 'Event type': 'Event type',
[0117] 'Role': {
[0118] 'Disease name': 'Role content',
[0119] 'Disease location': 'Role content',
[0120] 'Disease quantity': 'Role content',
[0121] 'Disease size': 'Role content',
[0122] 'Disease classification': 'Role content',
[0123] 'Mucosal status': 'Role content',
[0124] }
[0125] }
[0126] ···
[0127] }
[0128] Note: Please standardize the disease size in the role to xx cm * xx cm, for example, 0.1x0.8 cm should be standardized to 0.1cm*0.8cm
[0129] **Output strictly in the above JSON format**
[0130] """
[0131] As shown above, the core of the Prompt structure includes role-playing, Input Text, and explanatory text for generating instructions and tasks. Note that in the writing structure of the entire Prompt, the COT (Chain of Thought) method is used in the embodiments of the present invention for prompting. In text standardization, the ICL (In-context Learning) method is used in the embodiments of the present invention for prompting. In the above example of text standardization rules, only the standardization of disease size is used as an example for illustration. Again, the generated Prompt is fine-tuned by Qwen2.5-72B to generate results, and its generation process will call the built-in function call module of Qwen2.5-72B to trigger the Json Output Parser function call. Finally, the built-in function call triggers the Json Output Parser function to parse the final generation result of Qwen2.5-72B into a dictionary object in Python. So far, the inference of the large model for the entire event extraction ends. The embodiments of the invention can accurately perform event extraction and perform text standardization operations on the extraction results after extraction.
[0132] The embodiments of the present invention can also implement text report quality control scoring and missing field acquisition based on the edit distance.
[0133] The edit distance in the embodiments of the present invention refers to the Levenshtein Distance. The Levenshtein Distance refers to the minimum number of edit operations required to convert one string into another between two strings. The allowed edit operations include:
[0134] 1. Replace one character with another (Substitutions).
[0135] 2. Insert a character (Insertions).
[0136] 3. Delete a character (Deletions).
[0137] The embodiments of the present invention calculate the edit distance between the result of the text report after event extraction and the preset standard structured report based on the calculation rules of the edit distance. Theoretically, the smaller the edit distance, the closer the text report is to the standard structured report, and the higher the quality control score of the text report should be at this time. The embodiments of the present invention quantify the quality control score of the text report through this method of measuring the similarity of structured texts. Specifically, determining the quality control score of the target endoscopic text report according to the edit distance specifically includes: determining the quality control score score of the target endoscopic text report by using the following scoring model. The scoring model includes:
[0138] score = r × 100
[0139]
[0140] Among them, sum is the total character length of the formatted output result of the current event and the corresponding standard structured report text, and l_dist represents the edit distance between the formatted output result of the current event and the corresponding standard structured report text. Since r is a value between 0 and 1, and the larger this value is and the closer it is to 1, it means the smaller the edit distance, and at this time it means the more similar the strings are. In the embodiment of the present invention, r is directly transformed into the quality control score of the text report. The value range of the quality control score score of the text report is from 0 to 100, and the larger the score value is, the more consistent the formatted output result and the standard structured report text are.
[0141] Due to factors such as the description order and the doctor's personalized description, the original text report and the standard structured report text may have the same semantics but completely different differences in text character descriptions. If the quality control score is directly calculated by comparing the original text report with the standard structured report text, this calculation method is definitely inaccurate and has serious deviations. To eliminate this influence, in the embodiment of the present invention, event detection and event extraction are performed in the first two stages of the AI quality control service, and the structured text obtained by event extraction and the standard structured report text are used for text quality control scoring. This method eliminates the deviation of directly using the text report for edit distance calculation and ensures the accurate quantification of the quality control score.
[0142] In addition to being able to achieve accurate and quantifiable quality control scores, the embodiment of the present invention can also quickly obtain the missing fields of the endoscopic text report and simultaneously obtain the reverse structured data of the text report through event extraction and text standardization. Among them, the specific steps for obtaining the missing fields are as follows: performing set operations on the formatted output result of each event in the target endoscopic text report and the corresponding standard structured report text to detect the missing fields in the target endoscopic text report. The set operation model is as follows: C = B - (A ∩ B), where the fields corresponding to set C are the missing fields in the target endoscopic text report, the fields corresponding to set B are the argument role fields in the standard structured report text, and the fields corresponding to set A are the argument role fields in the formatted output result of the event; generating a quality control analysis report based on the quality control score and the missing fields of the target endoscopic text report.
[0143] In this embodiment, by comparing the formatted output result obtained by event extraction and text standardization of the text report with the fields of the standard structured report text, the missing quality control indicators in the original text report can be accurately found.
[0144] An example of the standard structured report text is as follows:
[0145] {
[0146] "Standardized fields for disease description":{
[0147] "Event type":"Disease description",
[0148] "Roles":{"Disease name", "Disease location", "Number of diseases", "Size of diseases", "Disease classification", "Mucosal status"}
[0149] "Argument value range": # Omitted here
[0150] }
[0151] "Standardized fields for location description":{
[0152] "Event type":"Location description",
[0153] "Roles":{"Disease name", "Disease location", "Mucosal status"}
[0154] "Argument value range": # Omitted here
[0155] }
[0156] }}。
[0157] In a specific example, event extraction and standardization obtained the structured information of the target endoscopic text report. For example, the structured information of "Event 1" here, and the values of the corresponding roles are: A = ["Disease name", "Disease location", "Number of diseases", "Size of diseases", "Disease classification", "Mucosal status"], and the standardized fields for disease description are B = ["Disease name", "Disease location", "Number of diseases", "Size of diseases", "Disease classification", "Mucosal status"].
[0158] Through set operation: C = B - (A ∩ B), the fields corresponding to the obtained set C are the fields missing from the quality control score. In this example, C = It means that there are no missing fields in this text report in the acquisition of the quality control missing fields in the first step.
[0159] Event extraction and standardization obtained the structured information of this text report. For example, the structured information of "Event 1" here. By checking the values of the corresponding arguments, it is found that the value of the "Mucosal status" role is "", which indicates that in the original text description "2 polyps were found in the ileocecal region, the larger one is about 0.01x0.2 cm, Paris classification.", the description of the quality control index "Mucosal status" is missing.
[0160] Furthermore, the present invention can also display the quality control analysis report in real time through the AI quality control assistant plug-in program pre-installed in each digestive endoscopy graphic and text system workstation to remind the supplement of missing fields in the quality control analysis report; when the supplementary data of the missing fields is received, the supplementary data will be uploaded to enable the cloud to recalculate the quality control score and generate a quality control analysis report, and then display it again through the AI quality control assistant plug-in program. In this embodiment, the report will be displayed in real time through the AI quality control assistant plug-in program installed in each digestive endoscopy graphic and text system workstation for reminder, and doctors can view the quality control score results of each report through this plug-in.
[0161] In the embodiment of the present invention, the output result obtained through event extraction and text annotation by the endoscopy report event extraction large model is already structured text, that is, in the whole process of quality control scoring in the embodiment of the present invention, the unstructured endoscopy text report "2 polyps were found in the ileocecal region, the largest being about 0.01x0.2 cm, Paris classification." has been converted into a structured text report, specifically as follows:
[0162] ''Event 1'': {
[0163] ''Trigger word'': ''found polyps'',
[0164] ''Event type'': ''disease description'',
[0165] ''Role'': {
[0166] ''Disease name'': ''polyp'',
[0167] ''Disease occurrence site'': ''ileocecal region'',
[0168] ''Disease quantity'': ''2 pieces'',
[0169] ''Disease size'': ''0.01 cm x 0.2 cm'',
[0170] ''Disease classification'': ''Paris classification'',
[0171] ''Mucosal state'': '',
[0172] }
[0173] The endoscopic report quality control analysis method based on artificial intelligence provided by the embodiment of the present invention mainly realizes the following contents: A. Detect events in the endoscopic report through natural language processing, and extract and standardize the text report through a large language model. B. Quantify the quality control score of the text report based on the edit distance and obtain the missing fields in the endoscopic report. C. At the same time, a retrograde structured text report is obtained in the process.
[0174] The present invention accurately realizes the quality control scoring of text reports and provides the missing fields in the reports through the deep learning technology of natural language processing and large language models. Moreover, in this process, reverse-structured text is obtained as a by-product of the entire task. The present invention can assist in replacing part of the report review work of endoscopists, provide accurate and quantifiable quality control scores for text reports, and obtain reverse-structured text reports, which have very significant advantages in improving the work efficiency and quality control of endoscopists.
[0175] For the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0176] Embodiment 2
[0177] The embodiment of the present invention provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps in the above-mentioned embodiment of the endoscopic report quality control analysis method based on artificial intelligence, such as the steps S11 - S14 shown.
[0178] In the specific implementation process of Embodiment 2, reference can be made to Embodiment 1, and it has corresponding technical effects.
[0179] Embodiment 3
[0180] The embodiment of the present invention provides an endoscopic report quality control analysis device based on artificial intelligence, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it realizes the steps in the above-mentioned embodiment of the endoscopic report quality control analysis method based on artificial intelligence, such as the steps S11 - S14 shown.
[0181] In the specific implementation process of Embodiment 3, reference can be made to Embodiment 1, and it has corresponding technical effects
[0182] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any one of the claimed embodiments can be used in any combination.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence-based endoscopy report quality control analysis method, characterized in that: The method comprises: Using a preset endoscopy report event detection model to detect disease description events and / or site description events in the target endoscopy text report to be quality controlled, and extracting target text segments corresponding to the detected disease description events and / or site description events from the target endoscopy text report; A prompt word template is constructed in a thinking chain manner, wherein the prompt word template includes a task description, input format constraints and input text examples, and output format constraints and output text examples. The task description is used to guide the perspective of the large model in task execution and describe the thinking structure of the task analysis process; Input each target text segment corresponding to the disease description event and / or the site description event extracted from the target endoscopic text report and the prompt word template into a preset endoscopic report event extraction macromodel, so that the endoscopic report event extraction macromodel analyzes each target text segment based on the prompt word template, and extracts the event trigger word of the corresponding event, the argument role for identifying the quality control indicator corresponding to the event, and the argument entity for describing each argument role from each target text segment, and presents the relevant quality control indicators according to the type of inspection performed; the argument roles of the disease description event include the name of the disease, the site of the disease, the number of diseases, the size of the disease, the disease classification, and the mucosal state; the argument roles of the site description event include the name of the disease, the site of the disease, and the mucosal state; the extraction results of each event are formatted and output to obtain a structured text report; The edit distance between the formatted output result of each event in the target endoscopy text report and the preset standard structured report text corresponding to the same event type is statistically calculated, and the quality control score of the target endoscopy text report is determined according to the edit distance to form a quality control analysis report of a single inspection based on the inspection type.
2. The method according to claim 1, characterized in that: The prompt word template also includes a text standardization rule description of the output result; The formatting and outputting of the extraction results of each event includes: The text standardization operation is performed on the extraction results of each event according to the text standardization rules, and the extraction results after the text standardization operation are formatted and output according to a preset json format.
3. The method according to claim 1 or 2, characterized in that: The quality control score of the target endoscopy text report determined according to the edit distance includes: The quality control score of the target endoscopy text report is determined using the following scoring model, which includes: score=r×100 Where sum is the total character length of the formatted output result of the current event and the corresponding standard structured report text. It represents the edit distance between the formatted output result of the current event and the corresponding standard structured report text.
4. The method according to claim 1 or 2, characterized in that: The method further comprises: The formatted output results of each event in the target endoscopy text report and the corresponding standard structured report text are set to detect missing fields in the target endoscopy text report. The set operation model is as follows: C=B-(A∩B) The fields corresponding to set C are the missing fields in the target endoscopy text report, the fields corresponding to set B are the argument role fields in the standard structured report text, and the fields corresponding to set A are the argument role fields in the formatted output results of the event; Generate a QC analysis report based on the QC scores and missing fields of the target endoscopy text report.
5. The method according to claim 4, characterized in that The method further comprises: The quality control analysis report is displayed in real time through the AI quality control assistant plug-in program pre-installed in each digestive endoscopy graphic system workstation to remind the missing fields in the quality control analysis report to be supplemented; When the supplementary data for the missing fields is received, the supplementary data is uploaded so that the cloud can recalculate the quality control score and generate a quality control analysis report, and then display it again through the AI quality control assistant plug-in.
6. The method according to claim 1, characterized in that The task description in the prompt word template includes the following thinking tasks to be performed in sequence: Identify event trigger words in the text. Event trigger words include: verbs, verb phrases or pre-set specific nouns; Extract the argument roles related to each event to identify the quality control indicators corresponding to the event, including: disease name, disease site, disease quantity, disease size, disease classification and mucosal status; Label each event with its event type; The structured output contains the trigger words, argument roles and their relationships for each event.
7. The method according to claim 1, characterized in that The endoscopy report event detection model is obtained by: Construct an endoscopy text report dataset; Annotate the disease description events and / or site description events in each endoscopic text report in the endoscopic text report dataset according to the preset event types and the labels corresponding to each event type, and construct an event detection training dataset according to the event annotation results; The pre-trained RoBERTa model was fine-tuned using the event detection training dataset to obtain an endoscopy report event detection model.
8. The method according to claim 7, characterized in that The endoscopy report event extraction model is obtained in the following way: The endoscopy report event detection model is used to detect disease description events and / or site description events for each endoscopy text report in the endoscopy text report data set, and sample text segments corresponding to the detected disease description events and / or site description events are extracted from each endoscopy text report; According to the event trigger words corresponding to the events of different event types, the argument roles used to identify the quality control indicators corresponding to the events, and the argument entities used to describe each argument role, the sample text fragments in each endoscopy text report are annotated, and an event extraction training data set is constructed based on the information annotation results; The event extraction training data set is used to fine-tune the pre-trained Qwen2.5-72B large language model to obtain an endoscopy report event extraction large model.
9. An artificial intelligence-based endoscopy report quality control analysis device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Image diagnosis report writing quality evaluation method and system
CN113593683A
Trigger word and argument extraction method, system and device and medium
CN116205220A
Medical image report structuring method and related device
CN119132497A
Diagnostic reading report writing support device
JP2009093568A