Endoscope report generation method and device based on multi-modal large model and storage medium
Through the digestive endoscopic report generation method based on multimodal large model, the problem of low quality of digestive endoscopic report is solved, the standardized generation and quality improvement of reports are achieved, and the work burden of doctors is reduced.
Patent Information
- Application Number
- CN202510383320.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
The quality of existing digestive endoscopic reports is not high, mainly due to insufficient human resources, uneven levels and large reporting workloads, only one-third of the reports have good quality.
The endoscopic report generation method based on a multimodal large model is adopted to collect digestive endoscopy images, detect suspicious lesions, form image reports, and generate text reports for description and diagnosis, and finally generate a complete, accurate and standardized digestive endoscopy report.
It improves the efficiency and quality of endoscopic report generation, overcomes the problems of differences in different doctor levels, realizes standardized generation of reports, provides patients with reliable diagnostic and treatment references, and reduces the work burden of doctors.
Smart Images

Figure CN120220949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical technology assistance, and more particularly to an endoscopic report generation method, device, computing device, and storage medium based on a multimodal large model. Background Art
[0002] Digestive endoscopy examinations (mainly including gastroscopy and colonoscopy here) are important means for diagnosing and treating digestive tract diseases. Digestive endoscopy reports are important records provided by endoscopic examinations. A complete, accurate, and standardized digestive endoscopy report is the basis for determining treatment strategies, providing clinical suggestions, and evaluating treatment efficacy. However, previous studies have shown that in clinical practice, due to problems such as a lack of human resources for endoscopic physicians, uneven levels, and a large workload of reports, taking gastroscopy as an example, only one-third of the reports have good quality. Therefore, there is an urgent need to develop a method and system that can automatically generate complete, accurate, and standardized digestive endoscopy reports. In recent years, artificial intelligence technology has developed rapidly and has been proven to be able to assist doctors in tasks such as diagnosing lesions and assisting patients in understanding doctor reports in the field of medical imaging. A multimodal large model is an artificial intelligence model that contains billions of adjustable parameters in its structural framework and can process and understand multiple types of data. Traditional models often can only focus on a single type of data, while multimodal large models can accept inputs from multiple modalities and perform association, understanding, and joint learning in multiple modalities, thus having the potential to handle cross-modal tasks. Digestive endoscopy reports have videos, images, and text at the same time, and their automatic generation belongs to a typical cross-modal task. Therefore, the present invention proposes a digestive endoscopy report system based on a multimodal large model, which can automatically analyze and form a complete, accurate, and standardized digestive endoscopy examination report during the digestive endoscopy examination process, thereby overcoming the level differences between different doctors, realizing the standardized generation of digestive endoscopy reports, providing a reliable reference for the subsequent diagnosis and treatment of patients, and reducing the workload of doctors. Summary of the Invention
[0003] The embodiments of the present application provide an endoscopic report generation method, device, computing device and storage medium based on a multimodal large model. According to the collected digestive endoscopy images, the suspicious lesions existing in the digestive endoscopy images are detected and representative images are collected to form an image report, and a text report describing and diagnosing is generated. Finally, a complete, accurate and standardized digestive endoscopy report is generated to improve the efficiency and quality of endoscopic report generation. On the one hand, the embodiments of the present application provide an endoscopic report generation method based on a multimodal large model, including: collecting digestive endoscopy images, and preprocessing the digestive endoscopy images to obtain preprocessed digestive endoscopy pictures; classifying and labeling the preprocessed digestive endoscopy pictures to construct a multimodal graphic and text report database for digestive endoscopy examinations; training a preset initial report model using the multimodal graphic and text report database for digestive endoscopy examinations to obtain a trained digestive endoscopy report model; inputting the digestive endoscopy images of the report to be output into the trained digestive endoscopy report model to generate an output digestive endoscopy report.
[0004] In one embodiment, the preprocessing the digestive endoscopy images to obtain preprocessed digestive endoscopy pictures includes:
[0005] Extracting the digestive endoscopy images according to a preset frame rate to obtain pictures to be processed;
[0006] Removing the parts that do not meet the preset resolution from the pictures to be processed to obtain filtered pictures;
[0007] Performing de-similarity processing on the filtered pictures according to a preset algorithm to obtain de-similarized pictures;
[0008] Identifying the de-similarized pictures according to a picture cleanliness model to obtain pictures with qualified cleanliness;
[0009] Performing cropping processing on the pictures with qualified cleanliness to obtain preprocessed digestive endoscopy pictures.
[0010] In one embodiment, the classifying and labeling the preprocessed digestive endoscopy pictures to construct a multimodal graphic and text report database for digestive endoscopy examinations includes:
[0011] Invoking a preset digestive endoscopy lesion monitoring model to identify pictures with visible lesions from the preprocessed digestive endoscopy pictures;
[0012] Constructing a lesion picture library according to the pictures with visible lesions;
[0013] Sending the preprocessed digestive endoscopy pictures to a target terminal for report annotation;
[0014] After the preprocessed digestive endoscopy images are reported and annotated on the target terminal, obtain the report annotations corresponding to the preprocessed digestive endoscopy images;
[0015] Construct a multi-modal graphic report database for digestive endoscopy examinations based on the lesion picture library and the report annotations corresponding to the preprocessed digestive endoscopy images.
[0016] In one embodiment, the training of the preset initial report model using the multi-modal graphic report database for digestive endoscopy examinations to obtain the trained digestive endoscopy report model includes:
[0017] Attach a low-rank adaptation layer to the preset large language model to obtain the initial report model;
[0018] Perform graphic and text encoding processing on the multi-modal graphic report database for digestive endoscopy examinations to obtain training data;
[0019] Iteratively train the initial report model until a preset target is reached to obtain the trained digestive endoscopy report model.
[0020] In one embodiment, the iterative training of the initial report model until a preset target is reached to obtain the trained digestive endoscopy report model includes:
[0021] Input the training data into the initial report model to obtain the response text of the initial report model;
[0022] Define the loss function of the initial report model based on the training data and the response text;
[0023] Iteratively train the initial report model until the value of the loss function reaches the preset target, update the parameters of the low-rank adaptation layer, and obtain the trained digestive endoscopy report model.
[0024] In one embodiment, the method further includes:
[0025] Send the digestive endoscopy images of the report to be output to the target terminal for report annotation;
[0026] After the digestive endoscopy images of the report to be output are reported and annotated on the target terminal, obtain the digestive endoscopy report annotated on the target terminal;
[0027] Compare the digestive endoscopy report annotated on the target terminal with the output digestive endoscopy report to obtain a comparison result;
[0028] Score the output digestive endoscopy report according to the comparison result to obtain a scoring result;
[0029] Retrain the trained digestive endoscopy report model according to the scoring result.
[0030] In one embodiment, the retraining of the trained digestive endoscopy report model according to the scoring result includes:
[0031] Construct triple training data, where the triple training data includes digestive endoscopy examination images for which a report is to be output, the output digestive endoscopy report, and the scoring result;
[0032] Retrain the trained digestive endoscopy report model according to the triple training data.
[0033] In a second aspect, an endoscopic report generation device based on a multimodal large model provided by an embodiment of the present application has a function of implementing the endoscopic report generation method based on a multimodal large model provided in the above first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. In one embodiment, the endoscopic report generation device based on a multimodal large model includes:
[0034] A preprocessing module: used to collect digestive endoscopy examination images and preprocess the digestive endoscopy examination images to obtain preprocessed digestive endoscopy pictures;
[0035] A labeling module: used to classify and label the preprocessed digestive endoscopy pictures to construct a multimodal graphic and text report database for digestive endoscopy examinations;
[0036] A training module: used to train a preset initial report model using the multimodal graphic and text report database for digestive endoscopy examinations to obtain a trained digestive endoscopy report model;
[0037] An output module: used to input digestive endoscopy examination images for which a report is to be output into the trained digestive endoscopy report model to generate an output digestive endoscopy report;
[0038] In one embodiment, the preprocessing module is specifically used for:
[0039] Extract the digestive endoscopy examination images according to a preset frame rate to obtain pictures to be processed;
[0040] Remove parts that do not meet the preset resolution from the pictures to be processed to obtain filtered pictures;
[0041] Perform de-similarity processing on the filtered pictures according to a preset algorithm to obtain de-similarity pictures;
[0042] Identify and process the de - similarized images according to the image cleanliness model to obtain images with qualified cleanliness;
[0043] Crop the images with qualified cleanliness to obtain pre - processed digestive endoscopy images.
[0044] In one embodiment, the annotation module is specifically used for:
[0045] Call a preset digestive endoscopy lesion monitoring model to identify images with visible lesions from the pre - processed digestive endoscopy images;
[0046] Construct a lesion image library based on the images with visible lesions;
[0047] Send the pre - processed digestive endoscopy images to a target terminal for report annotation;
[0048] After the target terminal annotates the pre - processed digestive endoscopy images, obtain the report annotation corresponding to the pre - processed digestive endoscopy images;
[0049] Construct a multi - modal graphic - text report database for digestive endoscopy examinations based on the lesion image library and the report annotation corresponding to the pre - processed digestive endoscopy images.
[0050] In one embodiment, the training module is specifically used for:
[0051] Attach a low - rank adaptation layer to a preset large - language model to obtain an initial report model;
[0052] Perform graphic - text encoding processing on the multi - modal graphic - text report database for digestive endoscopy examinations to obtain training data;
[0053] Iteratively train the initial report model until a preset target is reached, and obtain a trained digestive endoscopy report model.
[0054] In one embodiment, the training module is specifically used for:
[0055] Input the training data into the initial report model to obtain the response text of the initial report model;
[0056] Define a loss function for the initial report model according to the training data and the response text;
[0057] Iteratively train the initial report model until the value of the loss function reaches a preset target, update the parameters of the low - rank adaptation layer, and obtain a trained digestive endoscopy report model.
[0058] In one embodiment, the device is also specifically used for:
[0059] Send the endoscopic image of the report to be output to the target terminal for report annotation;
[0060] After the target terminal annotates the endoscopic image of the report to be output, obtain the endoscopic report annotated on the target terminal;
[0061] Compare the endoscopic report annotated on the target terminal with the output endoscopic report to obtain a comparison result;
[0062] Score the output endoscopic report according to the comparison result to obtain a scoring result;
[0063] Retrain the trained endoscopic report model according to the scoring result.
[0064] In one embodiment, the device is further specifically configured to:
[0065] Construct triple training data, where the triple training data includes the endoscopic examination image of the report to be output, the output endoscopic report, and the scoring result;
[0066] Retrain the trained endoscopic report model according to the triple training data.
[0067] In a third aspect, an embodiment of the present application provides a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the endoscopic report generation method based on the multimodal large model described in the first aspect is implemented.
[0068] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions. When it runs on a computer, it causes the computer to execute the endoscopic report generation method based on the multimodal large model as described in the first aspect.
[0069] This application detects suspicious lesions in the collected endoscopic examination images, collects representative images, forms an image report, generates a text report of descriptions and diagnoses, and finally generates a complete, accurate, and standardized endoscopic examination report, improving the efficiency and quality of endoscopic report generation. Description of the Drawings
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0071] Figure 1 It is a schematic diagram of the scenario of the endoscopic report generation device based on the multimodal large model provided by the embodiment of the present application;
[0072] Figure 2 It is a schematic diagram of the scenario of an embodiment of the endoscopic report generation method based on the multimodal large model in the embodiment of the present application;
[0073] Figure 3 It is a schematic diagram of the process of an embodiment of the endoscopic report generation method based on the multimodal large model provided by the embodiment of the present application;
[0074] Figure 4 It is a schematic diagram of the scenario of an embodiment of preprocessing the digestive endoscopy examination images in the embodiment of the present application;
[0075] Figure 5 It is a schematic diagram of the scenario of an embodiment of an expert annotating the standard template of the endoscopic report in the embodiment of the present application;
[0076] Figure 6 It is a schematic diagram of the scenario of an embodiment of training a preset large language model in the embodiment of the present application;
[0077] Figure 7 It is a schematic diagram of the scenario of an embodiment of subjectively evaluating and calibrating the trained digestive endoscopy report model in the embodiment of the present application;
[0078] Figure 8 It is a schematic diagram of the structure of the endoscopic report generation method device based on the multimodal large model in the embodiment of the present application;
[0079] Figure 9 It is a schematic diagram of a structure of the endoscopic report generation device based on the multimodal large model in the embodiment of the present application;
[0080] Figure 10 It is a schematic diagram of a structure of a mobile phone in the embodiment of the present application;
[0081] Figure 11 It is a schematic diagram of a structure of a server in the embodiment of the present application.
[0082] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners
[0083] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0084] In the following description, specific embodiments of the present application will be described with reference to steps and symbols executed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be referred to as being executed by a computer several times. The computer execution referred to herein includes the operation of a computer processing unit that represents data in a structured form. This operation transforms the data or maintains it in a position in the computer's memory system, which can reconfigure or otherwise change the operation of the computer in a manner well known to those skilled in the art. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present application are described in the above text, which does not represent a limitation. Those skilled in the art will understand that the various steps and operations described below can also be implemented in hardware.
[0085] As used herein, the terms "module" or "unit" can be regarded as software objects executed on the computing system. The different components, modules, engines, and services described herein can be regarded as implementation objects on the computing system. The devices and methods described herein are preferably implemented in software, but of course can also be implemented in hardware, all within the protection scope of the present application.
[0086] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0087] Embodiments of the present application provide an endoscopic report generation method, device, and storage medium based on a multimodal large model.
[0088] Please refer to Figure 1 , Figure 1A scenario diagram of the endoscopy report generation device based on a multimodal large model provided by an embodiment of the present application. The endoscopy report generation device based on a multimodal large model may include an endoscopy report generation system 100 based on a multimodal large model and a user terminal 200. The endoscopy report generation system 100 based on a multimodal large model is connected through a network. An endoscopy report generation device based on a multimodal large model is integrated in the endoscopy report generation system 100 based on a multimodal large model. In an embodiment of the present application, the endoscopy report generation system 100 based on a multimodal large model may be a terminal device or a server. The endoscopy report generation system 100 based on a multimodal large model may send a digestive endoscopy report, etc., to the user terminal 200.
[0089] In an embodiment of the present application, when the endoscopy report generation system 100 based on a multimodal large model is a server, the server may be an independent server or a server network or server cluster composed of servers. For example, the servers described in the embodiments of the present application include, but are not limited to, computers, network hosts, single network servers, multiple network server sets, or cloud servers composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing (Cloud Computing). In an embodiment of the present application, communication between the server and the client may be achieved through any communication method, including but not limited to mobile communication based on the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMAX), or computer network communication based on the TCP / IP protocol suite (TCP / IP Protocol suite, TCP / IP), User Datagram Protocol (UDP) protocol, etc.
[0090] It can be understood that when the endoscopy report generation system 100 based on a multimodal large model used in an embodiment of the present application is a terminal device, the terminal device may be a device that includes both receiving hardware and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such terminal devices may include: cellular or other communication devices, which have a single-line display or a multi-line display or cellular or other communication devices without a multi-line display. Specifically, the endoscopy report generation system 100 based on a multimodal large model may specifically be a desktop terminal or a mobile terminal. The endoscopy report generation system 100 based on a multimodal large model may specifically be one of a mobile phone, a tablet computer, a laptop computer, etc.
[0091] The terminal device involved in the embodiments of this application can also be a device that provides voice and / or data connectivity to users, such as a handheld device with wireless connection capabilities or other processing devices connected to a wireless modem. For example, a mobile phone (or a "cellular" phone) and a computer with a mobile terminal. For instance, it can be a portable, pocket-sized, handheld, computer-integrated, or vehicle-mounted mobile device that exchanges voice and / or data with a wireless access network. For example, devices such as Personal Communication service (PCs) phones, cordless phones, Session Initiation Protocol (sIP) phones, Wireless Local Loop (WLL) stations, and Personal Digital Assistant (PDA).
[0092] Those skilled in the art can understand that Figure 1 the application environment shown is merely one application scenario of the solution of this application and does not constitute a limitation on the application scenario of the solution of this application. Other application environments may also include more or fewer computing devices than Figure 1 shown, or the network connection relationship of computing devices. For example Figure 1 only 1 computing device is shown in. It can be understood that the endoscopic report generation device based on the multi-modal large model may also include one or more other computing devices, or / and one or more other computing devices network-connected to the endoscopic report generation system 100 based on the multi-modal large model. The specific details are not limited here.
[0093] In addition, as Figure 1 shown, the endoscopic report generation device based on the multi-modal large model may also include a memory 300 for storing data, such as storing digestive endoscopy examination images, digestive endoscopy examination multi-modal graphic report databases, etc.
[0094] It should be noted that Figure 1 the scenario schematic diagram of the endoscopic report generation device based on the multi-modal large model shown is merely an example. The endoscopic report generation device and scenario described in the embodiments of this application are for more clearly explaining the technical solution of the embodiments of this application and do not constitute a limitation on the technical solution provided by the embodiments of this application. Those skilled in the art know that with the evolution of the endoscopic report generation device based on the multi-modal large model and the emergence of new business scenarios, the technical solution provided by the embodiments of this application is equally applicable to similar technical problems.
[0095] Based on the collected digestive endoscopy images, this application detects suspicious lesions in the digestive endoscopy images, collects representative images, forms an image report, generates a text report of description and diagnosis, and finally generates a complete, accurate, and standardized digestive endoscopy report, improving the efficiency and quality of endoscopy report generation.
[0096] In this embodiment, a description will be made from the perspective of an endoscopy report generation method based on a multi-modal large model. This endoscopy report generation method based on a multi-modal large model can be specifically integrated in an endoscopy report generation system 100 based on a multi-modal large model.
[0097] Please refer to Figure 2 , Figure 2 , which is a schematic diagram of an embodiment scenario of the endoscopy report generation method based on a multi-modal large model in an embodiment of this application. This application provides an endoscopy report generation method based on a multi-modal large model. This endoscopy report generation method based on a multi-modal large model includes: collecting digestive endoscopy examination images, and preprocessing the digestive endoscopy examination images to obtain preprocessed digestive endoscopy pictures; classifying and labeling the preprocessed digestive endoscopy pictures to construct a multi-modal graphic and text report database for digestive endoscopy examinations; using the multi-modal graphic and text report database for digestive endoscopy examinations to train a preset initial report model to obtain a trained digestive endoscopy report model; inputting the digestive endoscopy examination images to be outputted into the trained digestive endoscopy report model to generate an outputted digestive endoscopy report.
[0098] Please refer to Figure 3 , Figure 3 , which is a schematic diagram of an embodiment process of the endoscopy report generation method based on a multi-modal large model in an embodiment of this application. This endoscopy report generation method based on a multi-modal large model includes the following steps 301 to 304:
[0099] 301. Collect digestive endoscopy examination images, and preprocess the digestive endoscopy examination images to obtain preprocessed digestive endoscopy pictures.
[0100] Specifically, before collecting digestive endoscopy examination images, a data collection port can also be constructed. The data collection port can automatically collect digestive endoscopy examination images. Optionally, the data collection port automatically collects the examination requisition, and uses a preset text recognition algorithm to obtain information such as the examination category and instrument model that the patient is about to undergo in the examination requisition. Correspondingly, when it is recognized that the examination category that the patient is about to undergo is digestive endoscopy examination, the examination requisition is marked, and the corresponding digestive endoscopy image of the examination requisition is obtained, completing the collection of the digestive endoscopy examination image, and inputting the digestive endoscopy examination image of the corresponding patient into the corresponding preprocessing module. Optionally, please refer to Figure 4 , Figure 4This is a schematic diagram of an embodiment scenario for preprocessing endoscopic examination images in the embodiments of the present application. After the preprocessing module obtains the endoscopic examination images, optionally, the preprocessing operations include: (1) Extracting picture frames from the endoscopic examination images according to an adjustable specific frame rate. Among them, the collected endoscopic examination images can be pictures or videos, and the specific frame rate can be preset and used to extract relatively clear and complete endoscopic pictures from the video stream. (2) Removing the parts with substandard resolution according to an adjustable specific resolution. Optionally, when there are multiple extracted picture frames, the resolution of multiple picture frames can be screened again, and some picture frames with insufficient clarity can be removed. (3) Using a picture similarity evaluation module that combines the average hashing algorithm, perceptual hashing algorithm, differential hashing algorithm, and SIFT\GIST algorithms to evaluate the picture similarity, and removing the similar pictures obtained in the previous step according to an adjustable specific threshold. Optionally, using a preset algorithm to further desimilarize the picture frames obtained in the previous step, evaluating the similarity between the picture frames, and removing the similar pictures to facilitate the subsequent construction of a database using endoscopic pictures without generating redundant data. (4) Identifying the pictures with substandard gastrointestinal cavity cleanliness in the pictures obtained in the previous step according to a previously trained picture cleanliness model. Optionally, the picture cleanliness model can be preset and already trained, and the picture cleanliness model can be used to screen and remove the pictures with substandard gastrointestinal cavity cleanliness. (5) Cropping the edges and irrelevant information of the pictures obtained in the previous step. Optionally, when there is information irrelevant to the endoscopic image at the edge of the picture frame, the irrelevant information can be cropped. Optionally, an edge detection algorithm can be used to identify the endoscopic picture and remove the irrelevant edge information.
[0101] 202. Classify and label the preprocessed endoscopic pictures to construct a multi-modal graphic and text report database for endoscopic examinations.
[0102] Furthermore, the endoscopic images can include gastroscopy and colonoscopy examination images. The examination images can be videos or pictures. After continuously collecting gastroscopy and colonoscopy examination videos through the aforementioned data collection process and completing the preprocessing by the preprocessing module, an endoscopic examination picture library corresponding to gastroscopy and colonoscopy can be obtained. Optionally, there are pictures with obvious visible lesions in the endoscopic examination picture libraries of gastroscopy and colonoscopy. A digestive endoscopic lesion monitoring model can be used to monitor the pictures in the picture libraries to screen out the pictures with visible lesions, and all the pictures with visible lesions are used to construct a lesion picture library. Furthermore, after obtaining the lesion picture library, experts are invited to formulate a digestive endoscopic report specification template according to the digestive endoscopic diagnosis and treatment guidelines, which specifically includes an image report specification template, a text report specification template, and a lesion diagnosis specification template. Please refer to Figure 5 , Figure 5This is a schematic diagram of an example scenario of the expert-annotated endoscopic report specification template in the embodiments of this application. Optionally, as Figure 5 shown, the annotation descriptions of the same visible lesion picture can be completed by multiple doctors / experts. Figure 5 In Figure 5 , it is completed by experts A, B, and C, thus forming a complete and comprehensive endoscopic report specification template. After confirming the endoscopic report specification template, invite endoscopic doctors to perform report specification template annotation based on the preprocessed digestive endoscopy pictures obtained in the previous step. Specifically, for the selected pictures in the preprocessed digestive endoscopy picture library, describe the lesion pictures and match them to generate a text report, combine the picture report and the text report to generate a diagnosis, and finally integrate them to form a multi-modal graphic and text report database.
[0103] 203. Use the multi-modal graphic and text report database of the digestive endoscopy examination to train a preset initial report model to obtain a trained digestive endoscopy report model.
[0104] Specifically, the initial report model can be a preset large language model. A large language model (LLMs) is an artificial intelligence system trained on a large amount of text data using a neural network architecture. General large language models such as PaLM, LLaMA, and ChatGLM have developed rapidly. Medical researchers have developed medical large language models by training and fine-tuning general large language models to be applicable to the medical field. Optionally, the initial report model can be MedPaLM-2, a large language model covering 193k medical Q&A. Specifically, select one from a group of selectable large language models and freeze most of its parameters as a base model, and attach a low-rank adaptation layer as the part that needs to be fine-tuned. Use a visual diffusion module with open parameters and a cross-modal connection layer with open parameters to match and encode the graphic and text pairs to obtain training data, and input the training data into the base large language model to obtain the answer text of the base large language model. Define the contrast loss function between the semantic embedding of the answer text of the large model obtained in the previous step and the training data as the target to train the large model, and update the obtained parameters in the low-rank adaptation layer mentioned in the previous step, that is, update the parameters of the low-rank adaptation layer. Please refer to Figure 6 , Figure 6This is a schematic diagram of an example scenario for training a preset large language model in an embodiment of this application. Optionally, after training is completed and the trained digestive endoscopy report model is obtained, the large model can also be subjectively evaluated and calibrated. A subjective evaluation and calibration module is constructed that can identify the picture report, text report, and diagnosis in the model results, and the doctor fills in the corresponding partition scores or grades. After training, test data is randomly selected from the multi-modal graphic and text reports, and endoscopic doctors are invited to label a specific proportion of the data, and other endoscopic doctors are invited to evaluate the results until the score of the result generated by the model is not inferior to that of the doctor, which is used as the adoptable version. Please refer to Figure 7 , Figure 7 This is a schematic diagram of an example scenario for subjectively evaluating and calibrating the trained digestive endoscopy report model in an embodiment of this application. Specifically, cross-scoring is used, and the reports output by the model are respectively labeled by AI, junior doctors, and senior doctors to construct a test set. In each round of evaluation by experts, one-third is randomly selected from the labels from three sources to form a new test set, and through three rounds of traversal, the labeling of all pictures and output results is completed.
[0105] 204. Input the digestive endoscopy examination images of the report to be output into the trained digestive endoscopy report model to generate the output digestive endoscopy report.
[0106] Optionally, after obtaining the trained digestive endoscopy report model, modules such as data collection, preprocessing, multi-modal digestive endoscopy graphic and text report large model, and report output can also be integrated to form a system for the digestive endoscopy graphic and text report model based on the multi-modal large model. Specifically, the modules and workflows required for the above-mentioned sorting, labeling, and database construction processes are integrated to form the input module of the graphic and text report system; the finally obtained multi-modal digestive endoscopy report model after training is integrated to form the report module of the graphic and text report system; the output module of the graphic and text report system is constructed according to the expert report specification template obtained in the previous step; a proofreading and review module of the graphic and text report system is constructed after the output module, which can modify the labeling reminder and text content of the pictures in the input module except for the preprocessing step, and resubmit them to the report module for reporting, and finally enter the review port for review by doctors. Optionally, after inputting the digestive endoscopy examination images of the report to be output into the trained digestive endoscopy report model to generate the output digestive endoscopy report, an automatic printing module can also be constructed, and a standard printing template is designed according to the expert report specification template obtained in the previous step, and the output result of the digestive endoscopy graphic and text report system based on the multi-modal large model obtained after the patient's examination, that is, the output digestive endoscopy report, is typeset and printed.
[0107] Based on the collected endoscopic examination images, this application detects suspicious lesions in the endoscopic examination images, collects representative images, forms an image report, generates a text report with descriptions and diagnoses, and finally generates a complete, accurate, and standardized endoscopic examination report, improving the efficiency and quality of endoscopic report generation.
[0108] In one implementation of this application, preprocessing the endoscopic examination images to obtain preprocessed endoscopic pictures includes:
[0109] Extracting the endoscopic examination images according to a preset frame rate to obtain pictures to be processed; removing parts that do not meet the preset resolution in the pictures to be processed to obtain filtered pictures; performing de-similarity processing on the filtered pictures according to a preset algorithm to obtain de-similarized pictures; performing recognition processing on the de-similarized pictures according to a picture cleanliness model to obtain pictures with qualified cleanliness; and performing cropping processing on the pictures with qualified cleanliness to obtain preprocessed endoscopic pictures.
[0110] Specifically, optionally, the preprocessing operations include: (1) Extracting picture frames from the endoscopic examination images according to an adjustable specific frame rate. The collected endoscopic examination images can be pictures or videos, and the specific frame rate can be preset for extracting relatively clear and complete endoscopic picture frames from the video stream. (2) Removing parts with unqualified resolution according to an adjustable specific resolution. Optionally, when there are multiple extracted picture frames, the resolution of multiple picture frames can be further screened to remove some picture frames with insufficient clarity. (3) Using a picture similarity evaluation module that combines the average hash algorithm, perceptual hash algorithm, difference hash algorithm, and SIFT\GIST algorithms to evaluate the picture similarity, and removing the similar pictures obtained in the previous step according to an adjustable specific threshold. Optionally, using a preset algorithm to further perform de-similarity processing on the picture frames obtained in the previous step, evaluating the similarity between picture frames, and removing the similar pictures to prevent redundant data from being generated when constructing a database using endoscopic pictures later. (4) Identifying pictures with unqualified cleanliness in the digestive tract cavity in the pictures obtained in the previous step according to a previously trained picture cleanliness model. Optionally, the picture cleanliness model can be preset and already trained, and the picture cleanliness model can be used to screen and remove pictures with unqualified cleanliness in the digestive tract cavity. (5) Cropping the edges and irrelevant information of the pictures obtained in the previous step. Optionally, when there is information irrelevant to the endoscopic image at the edge of the picture frame, the irrelevant information can be cropped. Optionally, an edge detection algorithm can be used to identify the endoscopic picture and remove the irrelevant edge information.
[0111] In the embodiments of the present application, a preprocessing module preprocesses the digestive endoscopy images to obtain higher-quality digestive endoscopy pictures, which facilitates the subsequent construction of a database.
[0112] In an implementation manner of the present application, the classification and annotation of the preprocessed digestive endoscopy pictures to construct a multi-modal graphic report database for digestive endoscopy examinations includes:
[0113] Invoking a preset digestive endoscopy lesion monitoring model to identify pictures of visible lesions from the preprocessed digestive endoscopy pictures; constructing a lesion picture library based on the pictures of visible lesions;
[0114] Sending the preprocessed digestive endoscopy pictures to a target terminal for report annotation; after the target terminal annotates the preprocessed digestive endoscopy pictures, obtaining the report annotation corresponding to the preprocessed digestive endoscopy pictures; constructing a multi-modal graphic report database for digestive endoscopy examinations based on the lesion picture library and the report annotation corresponding to the preprocessed digestive endoscopy pictures.
[0115] Digestive endoscopy images can include gastroscopy and colonoscopy examination images, which can be videos or pictures. After continuously collecting gastroscopy and colonoscopy examination videos through the aforementioned data collection process and completing preprocessing by the preprocessing module, an endoscopic examination picture library corresponding to gastroscopy and colonoscopy can be obtained. Optionally, there are pictures with obvious visible lesions in the endoscopic examination picture libraries of gastroscopy and colonoscopy. The digestive endoscopy lesion monitoring model can be used to monitor the pictures in the picture libraries to screen out pictures of visible lesions, and all pictures of visible lesions are used to construct a lesion picture library. Further, after obtaining the lesion picture library, experts are invited to formulate a digestive endoscopy report specification template according to the digestive endoscopy diagnosis and treatment guidelines, which specifically includes an image report specification template, a text report specification template, and a lesion diagnosis specification template. Optionally, the annotation description of the same picture of visible lesions can be completed by multiple doctors / experts, so as to form a complete and comprehensive digestive endoscopy report specification template. After confirming the endoscopic report specification template, endoscopic doctors are invited to perform report specification template annotation based on the preprocessed digestive endoscopy pictures obtained in the previous step. Specifically, for the selected pictures in the preprocessed digestive endoscopy picture library, the lesions in the pictures are described and a text report is generated by matching, and a diagnosis is generated by integrating the picture report and the text report, and finally a multi-modal graphic report database is formed.
[0116] In the embodiments of the present application, a multi-modal graphic report database is constructed by performing report annotation on digestive endoscopy pictures, which facilitates the subsequent training of an initial report model using the database.
[0117] In an implementation manner of the present application, the training of a preset initial report model using the multi-modal graphic report database for digestive endoscopy examinations to obtain a trained digestive endoscopy report model includes:
[0118] Attach a low-rank adaptation layer to a preset large language model to obtain an initial report model; perform graphic and text encoding processing on the multi-modal graphic and text report database of the digestive endoscopy examination to obtain training data; perform iterative training on the initial report model until a preset target is reached to obtain a trained digestive endoscopy report model.
[0119] Optionally, the initial report model can be a preset large language model. Specifically, select one from a group of selectable large language models and freeze most of its parameters as a base model, and attach a low-rank adaptation layer as the required fine-tuning part. Use a vision diffusion module with open parameters and a cross-modal connection layer with open parameters to perform matching encoding on the graphic and text pairs to obtain training data, and input the training data into the base large language model to obtain the response text of the base large language model. Define a loss function, and perform iterative training on the initial report model until a preset target of model training is reached to obtain a trained digestive endoscopy report model.
[0120] In the embodiment of the present application, training data is obtained through graphic and text encoding, and the initial report model is trained using the training data to obtain a trained digestive endoscopy report model, improving the accuracy of report output.
[0121] In an implementation manner of the present application, the iterative training of the initial report model until a preset target is reached to obtain a trained digestive endoscopy report model includes:
[0122] Input the training data into the initial report model to obtain the response text of the initial report model; define a loss function of the initial report model according to the training data and the response text; perform iterative training on the initial report model until the value of the loss function reaches a preset target, update the parameters of the low-rank adaptation layer, and obtain a trained digestive endoscopy report model.
[0123] Specifically, use a vision diffusion module with open parameters and a cross-modal connection layer with open parameters to perform matching encoding on the graphic and text pairs to obtain training data, and input the training data into the base large language model to obtain the response text of the base large language model. Optionally, the training data is a binary group data of digestive endoscopy pictures - report annotations. Input the digestive endoscopy pictures in the training data into the initial report model, define a contrast loss function between the semantic embedding of the response text of the large model obtained in the previous step and the report annotations in the training data as the target to train the large model, and update the obtained parameters in the low-rank adaptation layer in the previous step to obtain a trained digestive endoscopy report model.
[0124] In the embodiment of the present application, the initial report model is iteratively trained by defining a loss function to obtain a trained digestive endoscopy report model, thereby improving the accuracy of the output of the digestive endoscopy report model.
[0125] In an implementation manner of the present application, the method further includes:
[0126] Sending the digestive endoscopy image of the report to be output to a target terminal for report annotation;
[0127] After the target terminal annotates the digestive endoscopy image of the report to be output, obtaining the digestive endoscopy report annotated on the target terminal; comparing the digestive endoscopy report annotated on the target terminal with the output digestive endoscopy report to obtain a comparison result; scoring the output digestive endoscopy report according to the comparison result to obtain a scoring result; and retraining the trained digestive endoscopy report model according to the scoring result.
[0128] Specifically, after obtaining the trained digestive endoscopy report model, in order to optimize the accuracy of the report output by the digestive endoscopy report model, the digestive endoscopy report model can be evaluated and optimized. Optionally, reports are randomly selected from the output reports of the digestive endoscopy report model to form test data. Endoscopy doctors are invited to annotate a specific proportion of the report data to obtain a manually annotated digestive endoscopy report, and the manually annotated digestive endoscopy report is compared with the digestive endoscopy report output by the model. Optionally, other endoscopy doctors / experts can be invited to evaluate the comparison result, and manual scores are given to the manually annotated digestive endoscopy report and the digestive endoscopy report output by the model respectively. The trained digestive endoscopy report model is retrained according to the scoring result until the result score obtained by the model in generating the report is not inferior to that of the doctor, then it is considered that the retraining is completed.
[0129] The embodiment of the present application improves the accuracy of the report output by the digestive endoscopy report model by optimizing the output result of the digestive endoscopy report model.
[0130] In an implementation manner of the present application, the retraining of the trained digestive endoscopy report model according to the scoring result includes:
[0131] Constructing triple training data, where the triple training data includes the digestive endoscopy examination image of the report to be output, the output digestive endoscopy report, and the scoring result; and retraining the trained digestive endoscopy report model according to the triple training data.
[0132] Optionally, construct triple data. After obtaining the scoring results of the test data, by constructing triple data, the triple data can be the endoscopic examination images to be output in the report - the endoscopic report output by the model - the scoring results. Embed the scoring results into the training data to facilitate model optimization, and use the triple data to retrain the endoscopic report model to further improve the accuracy of the model output.
[0133] In the embodiment of the present application, the triple data is used to retrain the endoscopic report model to further improve the accuracy of the model output.
[0134] To facilitate better implementation of the endoscopic report generation method based on the multi - modal large model provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above - mentioned endoscopic report generation method based on the multi - modal large model. The meanings of the nouns are the same as those in the above - mentioned endoscopic report generation method based on the multi - modal large model, and the specific implementation details can refer to the description in the embodiment of the endoscopic report generation method based on the multi - modal large model.
[0135] The endoscopic report generation device based on the multi - modal large model in the embodiment of the present application has the function of implementing the endoscopic report generation method based on the multi - modal large model provided in the above - mentioned embodiment. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above - mentioned function, and the modules can be software and / or hardware.
[0136] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the endoscopic report generation device based on the multi - modal large model provided in the embodiment of the present application. The endoscopic report generation device based on the multi - modal large model can be applied to a computing device in a content push scenario. Specifically, the endoscopic report generation method device 800 based on the multi - modal large model may include a pre - processing module 801, a labeling module 802, a training module 803, and an output module 804, as follows:
[0137] Pre - processing module: used to collect endoscopic examination images and pre - process the endoscopic examination images to obtain pre - processed endoscopic pictures;
[0138] Labeling module: used to classify and label the pre - processed endoscopic pictures to construct a multi - modal graphic - text report database for endoscopic examinations;
[0139] Training module: used to train a preset initial report model using the multi - modal graphic - text report database for endoscopic examinations to obtain a trained endoscopic report model;
[0140] Output module: used to input the endoscopic examination images of the report to be output into the trained endoscopic report model to generate the output endoscopic report;
[0141] In one embodiment, the preprocessing module is specifically configured to:
[0142] Extract the endoscopic examination images according to a preset frame rate to obtain the pictures to be processed;
[0143] Remove the parts that do not meet the preset resolution from the pictures to be processed to obtain the filtered pictures;
[0144] Perform de-similarity processing on the filtered pictures according to a preset algorithm to obtain the de-similarity pictures;
[0145] Perform recognition processing on the de-similarity pictures according to the picture cleanliness model to obtain the pictures with qualified cleanliness;
[0146] Perform cropping processing on the pictures with qualified cleanliness to obtain the preprocessed endoscopic pictures.
[0147] In one embodiment, the annotation module is specifically configured to:
[0148] Call a preset endoscopic lesion monitoring model to identify the pictures with visible lesions from the preprocessed endoscopic pictures;
[0149] Construct a lesion picture library according to the pictures with visible lesions;
[0150] Send the preprocessed endoscopic pictures to the target terminal for report annotation;
[0151] After the target terminal performs report annotation on the preprocessed endoscopic pictures, obtain the report annotation corresponding to the preprocessed endoscopic pictures;
[0152] Construct an endoscopic examination multi-modal graphic report database according to the lesion picture library and the report annotation corresponding to the preprocessed endoscopic pictures.
[0153] In one embodiment, the training module is specifically configured to:
[0154] Attach a low-rank adaptation layer to a preset large language model to obtain an initial report model;
[0155] Perform graphic and text encoding processing on the endoscopic examination multi-modal graphic report database to obtain training data;
[0156] Iteratively train the initial report model until a preset target is reached to obtain the trained endoscopic report model.
[0157] In one embodiment, the training module is specifically configured to:
[0158] Input the training data into the initial report model to obtain the response text of the initial report model;
[0159] Define the loss function of the initial report model according to the training data and the response text;
[0160] Iteratively train the initial report model until the value of the loss function reaches a preset target, update the parameters of the low-rank adaptation layer, and obtain the trained digestive endoscopy report model.
[0161] In one embodiment, the device is further specifically configured to:
[0162] Send the digestive endoscopy image of the report to be output to the target terminal for report annotation;
[0163] After the target terminal annotates the digestive endoscopy image of the report to be output, obtain the digestive endoscopy report annotated on the target terminal;
[0164] Compare the digestive endoscopy report annotated on the target terminal with the output digestive endoscopy report to obtain a comparison result;
[0165] Score the output digestive endoscopy report according to the comparison result to obtain a scoring result;
[0166] Retrain the trained digestive endoscopy report model according to the scoring result.
[0167] In one embodiment, the device is further specifically configured to:
[0168] Construct triple training data, where the triple training data includes the digestive endoscopy examination image of the report to be output, the output digestive endoscopy report, and the scoring result;
[0169] Retrain the trained digestive endoscopy report model according to the triple training data.
[0170] According to the collected digestive endoscopy examination images in the embodiments of the present application, detect the suspicious lesions existing in the digestive endoscopy examination images, collect representative images, form an image report, generate a text report of description and diagnosis, and finally generate a complete, accurate, and standardized digestive endoscopy examination report, improving the generation efficiency and quality of the endoscopy report.
[0171] The above describes the endoscopic report generation device based on a multimodal large model in the embodiments of the present application from the perspective of modular functional entities. The following describes the endoscopic report generation device based on a multimodal large model in the embodiments of the present application from the perspective of hardware processing.
[0172] It should be noted that Figure 8 The entity device corresponding to the output module 804 shown may be a transceiver, a radio frequency circuit, a communication module, an input / output (I / O) interface, etc., and the entity device corresponding to the training module 803 may be a processor.
[0173] Figure 9 The devices shown may all have a structure as Figure 8 shown. When Figure 9 the endoscopic report generation method device based on a multimodal large model shown has a structure as Figure 8 shown, Figure 9 the processor and transceiver in it can implement the same or similar functions of the training module 803 and the output module 804 provided in the device embodiment corresponding to the device. Figure 8 The memory in it stores the computer program that the processor needs to call when executing the above endoscopic report generation method based on a multimodal large model.
[0174] When the computing device in the embodiments of the present application is a terminal device, the embodiments of the present application also provide a terminal device. As Figure 10 shown, for ease of illustration, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0175] Figure 10 Shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided in the embodiments of the present application. Referring to Figure 10 , the mobile phone includes: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090, etc. Those skilled in the art can understand that Figure 10 the structure of the mobile phone shown in
[0176] The following will specifically introduce each component of the mobile phone in conjunction with Figure 10 :
[0177] The RF circuit 1010 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1080 for processing; in addition, the uplink data is sent to the base station. Generally, the RF circuit 1010 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, short messaging service (SMS), etc.
[0178] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 1020 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0179] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 1031), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1031. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0180] The display unit 1040 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation thereon or nearby, it transmits it to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 10 the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.
[0181] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0182] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and then converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent through the RF circuit 1010 to, for example, another mobile phone, or the audio data is output to the memory 1020 for further processing.
[0183] Wi-Fi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the Wi-Fi module 1070, which provides users with wireless broadband Internet access. Although Figure 10 the Wi-Fi module 1070 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.
[0184] The processor 1080 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.
[0185] The mobile phone further includes a power supply 1090 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 1080 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0186] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0187] In the embodiment of the present application, the processor 1080 included in the mobile phone also has the function of controlling the execution of the multi-modal large model-based endoscopic report generation method process performed by the above-mentioned multi-modal large model-based endoscopic report generation method device.
[0188] The embodiment of the present application also provides a server. Please refer to Figure 11 , Figure 11 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (English full name: central processing units, English abbreviation: CPU) 1122 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 for storing application programs 1142 or data 1144 (for example, one or more mass storage devices). Among them, the memory 1132 and the storage media 1130 may be transient storage or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1122 may be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the server 1100.
[0189] The server 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows server, Mac Os X, Unix, Linux, FreeBsD, etc.
[0190] The steps in the multi-modal large model-based endoscopic report generation method in the above embodiment may be based on the Figure 11 structure of the server 1100 shown.
[0191] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0192] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0193] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0194] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0195] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software function modules. If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0196] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0197] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired means (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless means (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0198] The technical solutions provided in the embodiments of the present application have been introduced in detail above. Specific examples are used in the embodiments of the present application to illustrate the principles and implementation manners of the embodiments of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the embodiments of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the embodiments of the present application.
Claims
1. A method for generating an endoscopic report based on a multimodal large model, characterized in that: The method for generating an endoscopic report based on a multimodal large model comprises: Collecting digestive endoscopic examination images, and preprocessing the digestive endoscopic examination images to obtain preprocessed digestive endoscopic pictures; Classifying and annotating the preprocessed digestive endoscopy images to construct a multimodal graphic report database for digestive endoscopy examinations; Using the digestive endoscopy multimodal graphic report database to train a preset initial report model to obtain a trained digestive endoscopy report model; The digestive endoscopy examination images of the to-be-output report are input into the trained digestive endoscopy report model to generate an output digestive endoscopy report.
2. The method for generating an endoscopic report based on a multimodal large model according to claim 1, characterized in that: The method of preprocessing the digestive endoscopy images to obtain preprocessed digestive endoscopy images comprises: Extracting the digestive endoscopy images according to a preset frame rate to obtain images to be processed; Removing a portion that does not meet a preset resolution from the image to be processed to obtain a filtered image; Performing de-similarization processing on the screened images according to a preset algorithm to obtain de-similarized images; Performing recognition processing on the removed similarity images according to the image cleanliness model to obtain images that meet the cleanliness standards; The images that meet the cleanliness standards are cropped to obtain pre-processed digestive endoscopy images.
3. The method for generating an endoscopic report based on a multimodal large model according to claim 1, characterized in that: The method of classifying and labeling the pre-processed digestive endoscopy images and constructing a digestive endoscopy multimodal graphic report database includes: Calling a preset digestive endoscopy lesion monitoring model to identify images of visible lesions from the preprocessed digestive endoscopy images; Building a lesion picture library according to the pictures of the visible lesions; Sending the preprocessed digestive endoscopy image to a target terminal for report annotation; After the target terminal performs report annotation on the preprocessed digestive endoscopy image, obtaining the report annotation corresponding to the preprocessed digestive endoscopy image; A multimodal graphic report database for digestive endoscopy is constructed based on the lesion image library and the report annotations corresponding to the preprocessed digestive endoscopy images.
4. The method for generating an endoscopic report based on a multimodal large model according to claim 1, characterized in that: The method of using the digestive endoscopy multimodal graphic report database to train a preset initial report model to obtain a trained digestive endoscopy report model includes: Add a low-rank adaptation layer to the preset large language model to obtain the initial report model; Performing image-text coding processing on the digestive endoscopy multimodal image-text report database to obtain training data; The initial report model is iteratively trained until a preset goal is reached to obtain a trained digestive endoscopy report model.
5. The method for generating an endoscopic report based on a multimodal large model according to claim 4, characterized in that: The iterative training of the initial report model until a preset goal is reached to obtain a trained digestive endoscopy report model includes: Inputting the training data into an initial report model to obtain an answer text of the initial report model; Defining a loss function of an initial report model based on the training data and the answer text; The initial report model is iteratively trained until the value of the loss function reaches a preset target, the parameters of the low-rank adaptation layer are updated, and the trained digestive endoscopy report model is obtained.
6. The method for generating an endoscopic report based on a multimodal large model according to claim 1, characterized in that: The method further comprises: Sending the digestive endoscopic image of the report to be output to the target terminal for report annotation; After the target terminal marks the digestive endoscopy image to be output as a report, obtaining the digestive endoscopy report marked on the target terminal; Comparing the digestive endoscopy report marked on the target terminal with the output digestive endoscopy report to obtain a comparison result; Scoring the output digestive endoscopy report according to the comparison result to obtain a scoring result; The trained digestive endoscopy report model is retrained according to the scoring results.
7. The method for generating an endoscopic report based on a multimodal large model according to claim 6, characterized in that: The step of retraining the trained digestive endoscopy report model according to the scoring result comprises: Constructing triplet training data, wherein the triplet training data includes digestive endoscopy images to be output as reports, output digestive endoscopy reports, and scoring results; The trained digestive endoscopy report model is retrained according to the triplet training data.
8. An endoscopic report generation device based on a multimodal large model, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating an endoscopic report based on a multimodal large model as described in any one of claims 1 to 7 is implemented.
9. A computing device, characterized in that It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating an endoscopic report based on a multimodal large model as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, enable the computer to execute the method for generating an endoscopic report based on a multimodal large model as claimed in any one of claims 1 to 7.