Construction method of ophthalmology language large model and ophthalmology intelligent inquiry method
By automating the screening and cleaning of data using a large language model, combining it with professional knowledge training, and optimizing computing resources through multi-machine, multi-card distributed training, the problem of data scarcity and high computing resource requirements in ophthalmology large models has been solved, realizing an efficient and accurate intelligent auxiliary diagnostic tool.
Patent Information
- Application Number
- CN202510854599.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing large-scale models lack high-quality, large-scale datasets for application in ophthalmology. The inconsistent data formats and high computational resource requirements result in poor model performance in professional tasks, making it difficult to meet clinical needs.
We use a large language model for automated data filtering and cleaning to build a high-quality ophthalmology dataset. We then train the model using ophthalmology expertise, optimize computing resources with multi-machine, multi-GPU distributed training, and fine-tune the model using the DeepSpeed ZeRO-3 strategy.
It significantly improves the professionalism and accuracy of the model in ophthalmological tasks, increases data processing efficiency and resource utilization, provides efficient and reliable intelligent auxiliary diagnostic tools, reduces the burden on doctors, and enhances the patient's medical experience.
Smart Images

Figure CN120913804A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular, the application of a large language model in the field of ophthalmology, more particularly, a method for constructing an ophthalmic language large model and an intelligent ophthalmic diagnosis method. BACKGROUND
[0002] Global vision health problems are increasingly prominent, with the continuous progress of information technology, digital ophthalmology technology has become an important means to improve ophthalmic medical services, making remote medical treatment possible, especially in areas where medical resources are scarce, it can provide timely and efficient ophthalmic services for patients, while improving the efficiency of diagnosis and treatment, it also promotes the fair distribution of medical resources. At present, with the rapid development of deep learning technology, especially large language models (Large Language Model, LLM), medical intelligent diagnosis has become an important research direction of medical artificial intelligence. In ophthalmic clinical practice, doctors often face heavy tasks, inconsistent standards, and strong subjectivity during diagnosis, especially in primary medical institutions, the shortage of ophthalmic specialists makes standardized and efficient ophthalmic medical services a realistic demand. Therefore, how to effectively use ophthalmic clinical data to build an ophthalmic large model with diagnostic ability is a hot issue in current research and application.
[0003] With the rapid development of intelligent medical technology, clinical data in the field of ophthalmology is growing at an alarming rate, and these data often come from various sources and have complex structures, covering information in different formats. How to efficiently integrate and deeply utilize clinical ophthalmic data has become a key problem that needs to be broken through in the current digital ophthalmology construction. However, the progress of data set construction in the field of ophthalmology is relatively lagging behind, and the existing publicly available data resources are mostly limited to single-modal image information, generally lacking corresponding structured text descriptions, making it difficult to meet the training needs of intelligent diagnosis and decision support systems. In addition, high-quality ophthalmic question-and-answer text data is also extremely scarce, further restricting the application of large models as represented by generative artificial intelligence in ophthalmic scenarios.
[0004] In the field of medical large models, HuatuoGPT and ChatMed have released Chinese medical dialogue data and large model structures, promoting the development of general medical question and answer systems. However, these works mainly focus on general scenarios and lack deep modeling and vertical optimization of ophthalmic professional knowledge. The clinical ophthalmic tasks that require high professional and accurate performance are still insufficient. In the field of ophthalmic large models, RetFound and VisionFM use massive amounts of unlabeled fundus image data for self-supervised learning pre-training, and have strong ophthalmic visual feature extraction capabilities. However, relying solely on image information, even if it covers multiple modalities, sometimes it is difficult to fully depict the overall picture of the disease. In clinical practice, ophthalmic diagnosis not only relies on imaging examination, but also needs to consider multiple sources of text information such as patient history, chief complaint, family history of genetic diseases, and laboratory test results. DeepDR-LLM fine-tunes the pre-trained large language model with diabetes care recommendation data, but the model's function focuses on the screening and care recommendations for diabetic retinopathy, making it difficult to cover other ophthalmic disease types. FFA-GPT and Slit Lamp GPT integrate pre-trained large language models to support interactive open question and answer, but their training data mainly comes from paired structured medical record documents, which are limited in quantity and lack open question and answer corpus in real-world scenarios, resulting in insufficient understanding and generation capabilities of the model for complex consultations, multi-turn conversations, and unstructured information.
[0005] Overall, current ophthalmic-related large model systems rely on specific disease types or specific tasks for construction, lack a unified data construction process and general training methods for all disease types and multiple tasks, and cannot meet the actual needs of clinical ophthalmic large models.
[0006] In summary, current large models have made significant progress in general fields, but their application in professional medical fields still faces many challenges. The existing technology mainly has the following problems when applying general large models to the ophthalmic field:
[0007] First, the scarcity of high-quality and large-scale data sets in the ophthalmic field is a key factor that hinders the improvement of model performance. Although there are some medical data sets, the text question and answer data is insufficient in quantity and professional depth. This makes it difficult for general models to achieve high precision and clinical practicality in ophthalmic professional tasks without sufficient domain-specific knowledge.
[0008] Secondly, the existing medical data sources are complex and have inconsistent formats, and the data cleaning and integration efficiency is low. The data collected from the Internet consultation platform, electronic medical record system, ophthalmology teaching materials, etc. often have different formats and contain a large amount of unstructured, redundant or noisy information, such as containing patients' casual language, emotional expression, non-medical description, etc. Traditional data preprocessing methods usually rely on a large amount of manual intervention, which is inefficient and easy to introduce subjective bias, seriously hindering the rapid construction and standardization of large-scale data sets.
[0009] Furthermore, in the existing training framework, the training task of the ophthalmic large model often faces challenges in computing resources and training efficiency. Due to the large number of model parameters, the demand for computing resources is high, and if there is no reasonable resource allocation scheme and multi-machine multi-card scheduling mechanism, it may lead to problems such as decreased training efficiency and prolonged training period. How to reasonably organize data flow and optimize distributed training configuration based on existing training methods to improve the controllability and stability of model training is still a key practical difficulty in the process of model application landing.
[0010] Therefore, it is urgent to build a large model training scheme for general ophthalmic tasks, which can systematically complete data screening, cleaning, enhancement and standard format conversion based on multi-source heterogeneous clinical data, and cooperate with instruction fine-tuning and training configuration optimization, etc. Technical path to effectively improve the professional expression ability, task generalization ability and clinical practicability of the model in the field of ophthalmology.
[0011] It should be noted that: the background technology is only used to introduce the relevant information of the present application, in order to help understand the technical solutions of the present application, but it does not mean that the relevant information must be prior art. In the absence of evidence that the relevant information has been disclosed before the filing date of the present application, the relevant information should not be regarded as prior art. SUMMARY
[0012] Therefore, the purpose of the present application is to overcome the defects of the prior art, and to provide a new method for constructing an ophthalmic language large model and an ophthalmic intelligent consultation method.
[0013] According to a first aspect of the present application, the present application provides a method for constructing an ophthalmology language large model, the ophthalmology language large model being used to output answers to questions of a user according to a prompt word and a question of the user, the method comprising: S1, obtaining a plurality of ophthalmology field original data sets; S2, performing data screening on the original data sets obtained in step S1 using a pre-configured first language large model to screen out data related to ophthalmology; S3, performing data cleaning and standardization processing on the data screened out in step S2 using the first language large model to uniformly standardize different data into data samples of a {prompt word, user question, question answer} structure; S4, collecting ophthalmology professional knowledge and constructing ophthalmology professional medical textbook fragments based thereon, generating ophthalmology question and answer data related to the content of the fragments using the first large language model according to the ophthalmology professional medical textbook fragments, and uniformly standardizing the question and answer data into data samples of a {prompt word, user question, question answer} structure; S5, using the first large language model to produce a plurality of groups of new data that are semantically equivalent but have different expression manners according to the data in step S4, and uniformly standardizing the new data into data samples of a {prompt word, user question, question answer} structure; S6, grouping the data samples obtained in steps S3, S4, and S5 into a training set, taking the prompt word and the user question as input and the question answer as label, and performing supervised training on a base large language model to obtain an ophthalmology field large language model.
[0014] Preferably, the plurality of ophthalmology field original data sets comprise: Medical, Chinese-medical-dialogue-data, cMedQA2, ChatMed_Consult_Dataset, Huatuo-Llama-Med-Chinese, HuatuoGPT-sft-data-v1, huatuo_medical_qa_sharegpt, DISC-Med-SFT, huatuo_encyclopedia_qa, Huatuo26M-Lite, and MedDialog.
[0015] Preferably, in the step S2, the first large language model is configured as a Qwen2.5-14B-Instruct large language model or a DeepSeek, and is deployed in a local manner; wherein the first large language model is configured to judge whether the data belongs to basic ophthalmology consultation or is ophthalmology knowledge, and to screen out data belonging to basic ophthalmology consultation and ophthalmology knowledge.
[0016] Preferably, in the step S3, the first large language model is positioned as an ophthalmology consultation data information assistant using a prompt word, and the data screened out in step S2 is executed to perform data cleaning and standardization processing: S31, extracting valid conversation content from the data, clearing irrelevant small talk and information unrelated to the disease; S32, arranging multi-round conversations, labeling each round of speech as "user" or "doctor" and arranging in sequence; S33, analyzing the patient's complaint information, integrating to form a new instruction field, including standard instructions and complaint summaries; S34, according to the multi-round conversation after arrangement, the user's last question is used to construct the doctor's standard answer; S35, the final remaining data is standardized, and the data sample of the {prompt word, user question, question answer} structure is output, wherein only the question and answer content related to the disease, the disease, and the treatment are retained in the data sample.
[0017] Preferably, in the step S4, the first large language model is positioned as an ophthalmology field data assistant using a prompt word, and a plurality of ophthalmology question and answer data related to the ophthalmology medical textbook content are generated according to the ophthalmology professional medical textbook fragments.
[0018] Preferably, the base large model is Qwen2.5-14B-Instruct large language model.
[0019] Preferably, in the step S6, a full-parameter fine-tuning strategy is used when the base large language model is supervised trained, so that all weights of the model are supervised adjusted under the premise of keeping the overall structure of the model unchanged.
[0020] Preferably, in the step S6, the base model is trained in a distributed manner based on multiple machines and multiple cards, and DeepSpeedZeRO-3 optimization strategy is used.
[0021] According to the second aspect of the present application, an ophthalmology intelligent consultation method based on a large language model is provided. The method comprises: T1, obtaining an ophthalmology field large language model constructed by the method of the first aspect of the present application; T2, inputting the symptoms and problems complained by the consultation person as user questions into the model obtained in step T1, and obtaining the question answers.
[0022] Compared with the prior art, the advantages of the present application are: (1) significantly improving the professional ability of the ophthalmic large model: by constructing a large-scale high-quality ophthalmic data set and combining intelligent data processing and advanced fine-tuning training methods, the large model trained by the present application performs far better than the existing general model in terms of professional knowledge and clinical inquiry assistance. The experimental results show that the fine-tuned model has a large improvement in various evaluation indicators, and can provide more reliable intelligent auxiliary diagnosis and more accurate medical information services. (2) greatly improving the data processing efficiency and quality: the introduction of large language models for data screening, question and answer pair generation and high-noise data cleaning realizes the automation and intelligentization of the data preparation process, significantly reduces the labor cost and time investment, and at the same time ensures the purity, professionalism and uniformity of the training data format, making the originally time-consuming and laborious data engineering link become efficient and feasible, speeding up the entire life cycle of the model from data to deployment. (3) optimizing the resource utilization and scalability of large model training: using DeepSpeed ZeRO-3 strategy optimization and parameter efficient fine-tuning technology, effectively reducing the memory consumption and computing demand in the training process. The implementation of multi-machine multi-card distributed training further breaks through the bottleneck of single-machine training, realizes the scalability and acceleration of training, so that the method of the present application can adapt to the training needs of larger scale models in the future, and reduce the deployment and maintenance cost. (4) providing efficient and reliable auxiliary tools for ophthalmic clinical practice: the final formed model can provide intelligent auxiliary diagnosis suggestions for ophthalmologists, provide professional and personalized ophthalmic consultation services for patients, and serve as a powerful knowledge base for medical education and research. This will effectively improve the efficiency of ophthalmic diagnosis and treatment, reduce the workload of doctors, improve the patient experience and satisfaction, and promote the intelligent development of ophthalmic medicine. BRIEF DESCRIPTION OF DRAWINGS
[0023] The embodiments of the present application will be further described below with reference to the accompanying drawings, in which:
[0024] Figure 1 The flowchart of the method for constructing an ophthalmic large language model according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0026] As mentioned in the background section, the existing technology mainly has the following problems when applying general large models to the ophthalmology field: First, the scarcity of high-quality and large-scale data sets in the ophthalmology field is a key factor restricting the improvement of model performance. Although there are some medical data sets, the text question and answer data are insufficient in quantity and professional depth. This makes it difficult for general models to achieve high precision and clinical practicability in ophthalmology professional tasks without sufficient domain-specific knowledge. Second, the existing medical data sources are complex and have inconsistent formats, and the data cleaning and integration efficiency is low. The data collected from internet consultation platforms, electronic medical record systems, ophthalmology textbooks, etc. often have different formats and contain a large amount of unstructured, redundant or noisy information, such as containing patients' casual language, emotional expression, non-medical description, etc. Traditional data preprocessing methods usually rely on a large amount of manual intervention, which is inefficient and easy to introduce subjective bias, seriously hindering the rapid construction and standardization of large-scale data sets. Third, in the existing training framework, the training task of the ophthalmology large model often faces challenges in computing resources and training efficiency. Due to the large number of model parameters, the demand for computing resources is high. If there is no reasonable resource allocation scheme and multi-machine multi-card scheduling mechanism, it may lead to problems such as decreased training efficiency and prolonged training period. How to reasonably organize data flow and optimize distributed training configuration based on existing training methods to improve the controllability and stability of model training is still a key practical difficulty in the application process of the model.
[0027] In addition, the inventors have found that directly training the model using a general large model and existing scattered ophthalmology data sets performs poorly in handling ophthalmology professional question and answer tasks, and the model output content is generalized, lacks professional terminology, and even has "hallucination" phenomena that do not conform to medical facts. Therefore, the transfer performance of the general large model in the professional medical scene is significantly insufficient in the absence of large-scale high-quality domain data support.
[0028] Therefore, the present application proposes a new ophthalmology large language model scheme from the following key technical difficulties:
[0029] (1) Large-scale high-quality ophthalmology field data acquisition and integration. Existing ophthalmology data sets are limited in size, and data formats, annotation standards are not unified, which makes it difficult to be directly used for large model training. Although the medical consultation data on the Internet is large, it is full of a lot of noise, unstructured information and irrelevant to ophthalmology. How to efficiently screen, clean and build a high-quality, large-scale ophthalmology field data set that meets the training needs of large models from massive, low-quality raw data is the first technical difficulty. The inventors found that manual screening and traditional script cleaning used in existing technologies are inefficient and difficult to ensure quality. In order to solve this problem, the inventors found through research that an automated screening and cleaning mechanism based on a large model can effectively solve this problem, such as using a large model to screen ophthalmology field data from open source medical data sets, and deeply cleaning high-noise Internet consultation data. This method greatly improves the efficiency and accuracy of data processing, solving the problem that traditional methods cannot scale to process heterogeneous data.
[0030] (2) Effective injection of professional knowledge and overcoming of model "illusion". Even after preliminary data processing, the model is still prone to "illusion" without the support of ophthalmology professional knowledge, outputting inaccurate or logically chaotic medical content. How to systematically integrate authoritative ophthalmology medical knowledge into training data and guide the model to learn medical logic rather than just statistical associations is another core difficulty. The inventors have tried simple keyword matching or rule injection, but the effect is not good, and it cannot achieve deep understanding of knowledge. To solve this problem, the inventors found through research that by using a large model to automatically generate a large number of ophthalmology textbooks, a high-quality professional question and answer data can be constructed. This method directly injects authoritative medical knowledge into training data, enabling the model to learn more accurate medical concepts and reasoning paths, thereby effectively reducing the occurrence of "illusion" phenomenon.
[0031] (3) Fine-tuning and performance optimization of large models for specific fields. According to the characteristics of the ophthalmology field, how to choose the appropriate base model, design an efficient fine-tuning strategy and optimize the training parameters to maximize the performance of the model on professional tasks, while taking into account the efficiency of computing resources, is the focus of technical research. The inventors have explored and verified different fine-tuning methods and parameter settings, and found that full-parameter fine-tuning works best. In addition, the present invention also uses multi-machine multi-card distributed training to support the training of larger scale models. These fine-tuning strategies and optimization measures enable the model to better adapt to the professional semantics and task requirements of the ophthalmology field, ultimately significantly improving the performance of the model on ophthalmology tasks.
[0032] In summary, the present application gradually overcomes the core technical difficulties of large-scale high-quality ophthalmic field data acquisition, effective injection of professional knowledge, and model fine-tuning training, etc. by combining automated tools (such as LLM assisted screening, cleaning and question and answer generation) and fine-tuning model training strategies, and finally forms a systematic ophthalmic field large model training method, which lays a solid foundation for realizing a high-reliable intelligent ophthalmic auxiliary system.
[0033] In order to better understand the present application, the present application will be described in detail below in conjunction with embodiments.
[0034] According to one embodiment of the present application, as shown in Figure 1 The present application provides a method for constructing an ophthalmic language large model, the ophthalmic language large model is used to output answers to user's questions according to prompt words and user's questions, and the method comprises the following steps: S1, obtaining a plurality of ophthalmic field original data sets; S2, performing data screening on the original data sets obtained in step S1 by using a pre-configured first language large model to screen out data related to ophthalmology; S3, performing data cleaning and standardization processing on the data screened out in step S2 by using the first language large model to uniformly standardize different data into data samples in the structure of {prompt word, user question, question answer}; S4, collecting ophthalmic professional knowledge and constructing ophthalmic professional medical textbook fragments based thereon, generating ophthalmic question and answer data related to the content of the ophthalmic professional medical textbook fragments by using the first large language model, and uniformly standardizing the question and answer data into data samples in the structure of {prompt word, user question, question answer}; S5, producing a plurality of groups of new data which are semantically equivalent but have different expression manners by using the first large language model according to the data in step S4, and uniformly standardizing the new data into data samples in the structure of {prompt word, user question, question answer}; S6, grouping the data samples obtained in steps S3, S4 and S5 into a training set, taking the prompt word and the user question as input and the question answer as label, and performing supervised training on the base large language model to obtain an ophthalmic field large language model.
[0035] In order to better understand the present application, the present application will be described in detail below in conjunction with embodiments.
[0036] I. Problem definition
[0037] The core problem to be solved by the present application can be formalized as: given an initial general large model and a set of original, possibly noisy ophthalmic medical data sources , how to design and implement an efficient and scalable training method so that the ophthalmic field large model Can: significantly outperform in ophthalmic professional tasks ; effectively overcome the "illusion" phenomenon, ensure the medical reliability of the generated content; have good generalization ability and be applicable to various ophthalmic diseases and clinical scenarios. Among them, indicates a high-quality, large-scale, structured ophthalmic field training data set after the method of the present application.
[0038] II. Construction of large-scale ophthalmic field data set
[0039] In the specific implementation process, the present application constructs a text question and answer data set of a total of 9,005,048 data, reaching the scale of 10 billion level corpus, aiming to provide sufficient and high-quality corpus for the training of large models, and the construction process is based on automatic data cleaning and data enhancement of large models. The specific construction process mainly includes the following five steps:
[0040] Step 1: Collection of open source medical question and answer data set
[0041] The present application first widely collects 11 open source data sets in the general medical field as shown in Table 1, laying the foundation for subsequent ophthalmic field data screening. This strategy of collecting from extensive general medical data and then screening for field-specific data is an effective way to deal with the scarcity of pure ophthalmic data. It provides a large data pool from which highly relevant medical knowledge in the ophthalmic field can be extracted.
[0042] Table 1
[0043] Step 2: Extraction of ophthalmic question and answer data
[0044] In order to accurately extract ophthalmic field related content from the above general medical data sets, the present application uses the Qwen2.5-14B-Instruct large language model deployed on a local server for data screening. The screening process uses specific system prompt words to guide the large language model to determine whether the medical description belongs to ophthalmic disease consultation or ophthalmic knowledge, and provides multiple positive and negative examples for context guidance to ensure the accuracy of the screening. For example, for questions such as "treatment of capillary hemangioma in the eye in Western medicine", the large model will judge "yes", while for questions such as "what are the symptoms of pediatric sinus tachycardia?", The large model will judge "no". After preliminary screening by the large model, a total of 611,505 ophthalmic field data were obtained. Compared with manual screening or keyword-based matching, the large model can understand the context and semantic association, thereby achieving more accurate and more scalable screening of massive heterogeneous medical data, effectively solving the problems of data noise and relevance.
[0045] Step 3: Standardization of ophthalmic question and answer data
[0046] Due to the inconsistent format of different data sets collected in step two, it is difficult to directly use them for model training. To adapt to the instruction fine-tuning requirements of large models, the present application designs a plurality of data format processing functions to convert the screened data into a unified SFT (Supervised Fine-Tuning) format, i.e. {"instruction": "", "input":"", "output": ""}. Among them, instruction tells the model what task to complete, i.e. prompt words such as question and answer, summary, diagnosis and suggestion, etc. The input is the question raised by the user or the chief complaint of the patient (closely related to the instruction), and the output is the professional answer generated by the model. Table 2 shows examples of two common types of data.
[0047] Table 2
[0048] Converting to SFT format ensures data consistency and solves the problem of data heterogeneity at the structural level, allowing heterogeneous data to be integrated into a unified training paradigm and ensuring the stability of model training.
[0049] In addition, to address the noise (such as irrelevant chatter and emotional expression) in internet consultation data, the present application designs a deep cleaning mechanism based on large models to address the inherent "noise" and "reliability" problems of internet data. The cleaning process uses specific system prompt words, using a large language model (here, the large language model can be the large language model used in the previous screening, or other large language models such as OpenAI GPT or DeepSeek) as an "ophthalmic consultation data cleaning assistant" and performing the following tasks:
[0050] 1. Extracting valid dialogue content from the original data, removing irrelevant chatter and other information unrelated to the disease.
[0051] 2. Organizing multi-round conversations, clearly labeling each round of speech as "user" or "doctor", and arranging them in order.
[0052] 3. Analyzing the patient's chief complaint information and integrating it into a new instruction field, including standard instructions and a summary of the complaint.
[0053] 4. Based on the multi-round conversation after sorting, the user's last question is used to construct the doctor's standard answer.
[0054] 5. Output follows the SFT format and emphasizes only retaining truly relevant questions and answers related to the disease, condition, and treatment.
[0055] The application significantly improves the accuracy and practicality of downstream ophthalmic large models in the diagnosis understanding and generation task by using large models for semantic cleaning, converting low-quality noise data into high-quality, clinically relevant training materials, and laying the foundation for building a high-trust, high-standard medical dialogue corpus.
[0056] Step four: ophthalmic textbook collection and question and answer pair construction
[0057] In order to inject authoritative ophthalmic medical knowledge into the training data, 212 ophthalmic field textbooks are collected in the application, and the text content is extracted through OCR technology to form a large-scale professional text corpus. On this basis, the application uses a large language model to automatically construct 207,876 question and answer data. The construction process uses specific system prompt words, uses the large model as an "ophthalmic field data assistant", generates as much ophthalmic question and answer data as possible related to the content according to the input medical textbook fragments, and requires the question to be clinically relevant, such as disease definition, pathogenesis, differential diagnosis, examination method, treatment plan, surgical indication, complication, postoperative management, prevention suggestion, etc., and returns in standard SFT format. By synthesizing question and answer pairs from authoritative textbooks through a large model, it is an efficient way to create a large-scale, authoritative knowledge base. The prompt words emphasize "clinical relevance" and "cover all important knowledge points", ensuring that the generated question and answer data directly contribute to training ophthalmic field large models.
[0058] Step five: data expansion based on large models
[0059] In order to further improve the robustness and generalization ability of the model, the application uses a large model as a "medical question and answer data enhancement expert" in the data expansion stage, and generates multiple sets of semantically equivalent but different expression methods of new data according to the input ophthalmic field SFT data. The system prompt words clearly require the large model to generate multiple sets of semantically equivalent but different expression methods of new data, including diversified rewriting of instructions / input and output, and require each set of expanded data to be high-quality, authentic and reliable medical questions and answers, and cannot introduce factual errors. After expansion, the total amount of text question and answer data set reaches 9,005,048 data. This data enhancement method based on large models, focusing on semantic equivalence rather than simple repetition or random noise addition, is crucial for improving model robustness and generalization ability. By generating diverse language expressions for medical content, the model's sensitivity to wording changes in actual queries is reduced, allowing it to better understand and respond to users' diverse expression methods, while avoiding overfitting to specific sentence patterns or vocabulary, which helps to enhance the model's practicality.
[0060] After steps three, four and five, a standardized large-scale ophthalmic field data set is obtained.
[0061] III. Eye field large model training
[0062] The present application performs instruction fine-tuning training of a large language model on a large-scale eye field dataset constructed to maximize the performance of the model on eye tasks. The present application selects Qwen2.5-14B-Instruct as a base large language model for full-parameter supervised fine-tuning to enhance its generation ability and professional adaptability in eye intelligent consultation tasks. Distributed training is performed based on multiple machines and multiple cards, the training process is performed on two high-performance computing servers equipped with 8 NVIDIA A800 80GB GPUs, and the DeepSpeed ZeRO-3 optimization strategy is combined to significantly improve the resource utilization efficiency and system scalability of large model training.
[0063] In terms of communication, the present application explicitly configures parameters such as the total number of nodes, node number, master node address and port, and simultaneously enables the RDMA network acceleration mechanism, adapts the high-performance NCCL communication interface, ensures high-bandwidth and low-latency data transmission capability between GPUs across nodes, and provides strong support for large-scale gradient synchronization and parameter broadcast.
[0064] According to one embodiment of the present application, a full-parameter fine-tuning strategy is adopted during training, and all weights are supervised and adjusted under the premise of keeping the overall structure of the model unchanged. The system integrates a mixed precision training mechanism to achieve better memory utilization and computing throughput, and uses DeepSpeed ZeRO-3 technology to divide model parameters, gradients and optimizer states into multiple partitions and allocate them to each GPU, achieving fine control of memory resources. The weight saving aggregation strategy is also enabled during the training process, effectively improving the stability and efficiency of the model persistence process, facilitating subsequent deployment and version management. In addition, the system adjusts the gradient continuity and parameter scheduling granularity to further improve the memory access efficiency and parameter life cycle management capability, enhancing the robustness and controllability of large model training.
[0065] In terms of training configuration, the present application adopts a setting of batch size 8 and gradient accumulation step number 8 on each GPU, achieving an equivalent global batch size of 1024, which helps to improve training stability and model generalization ability. The total number of training rounds is set to 3, the initial learning rate is set to 1e-5 using a cosine annealing scheduler, and a warmup ratio of 10% is configured to alleviate gradient shock in the initial stage. The multi-thread concurrent mechanism (16 preprocessing threads and 4 data loading threads) is enabled in the data loading stage to optimize the data processing throughput in the training process.
[0066] IV. Eye field large model evaluation
[0067] To verify the effectiveness of the ophthalmic large language model constructed according to the scheme, the ophthalmic large language model test set construction method in two stages is adopted: first, 5,000 data are randomly screened, and then 2,000 data are screened by a large model as the final test set. A strict system prompt word is used as the evaluation standard in the screening process to determine whether the question belongs to the ophthalmic clinical inquiry, whether the answer can be effectively inferred by professional knowledge, whether it meets the clinical professional requirements, and whether it reflects the representative clinical inquiry scene. This fine screening mechanism ensures that the model is evaluated on high-quality and highly relevant data, avoiding false high performance on atypical or low-quality data, thus truly reflecting the core capabilities of the model in the ophthalmic field.
[0068] Traditional automated evaluation indicators such as BLEU and ROUGE scores, while having a basic role in measuring language understanding, often fail to capture the inherent dynamics, interactivity, and subtleties of inquiry dialogue. To comprehensively evaluate the applicability and professionalism of large language models in the ophthalmic field in clinical scenarios, the present application constructs a multi-dimensional fine evaluation system for ophthalmic intelligent inquiry based on large models, focusing on performance measurement from four key dimensions: medical accuracy, information integrity, clinical utility, and expression professionalism. The evaluation system uses annotated answers as a reference benchmark, combined with real clinical communication scenarios, to ensure that the evaluation of model-generated content has high credibility and practical orientation. The evaluation criteria include medical accuracy, information integrity, clinical utility, and expression professionalism, each defined as follows.
[0069] Medical accuracy: examines whether the model-generated content conforms to current mainstream medical consensus, clinical guidelines, or normative diagnosis and treatment logic, focusing on the professional correctness and compliance of diagnosis and treatment recommendations.
[0070] Information integrity: examines whether the model's answer covers key clinical information points in the reference answer (such as symptom description, diagnosis conclusion, treatment measures, etc.), measuring the comprehensiveness and coverage of information transmission.
[0071] Clinical utility: evaluates whether the model-generated content has clear clinical executability, judges whether the recommendations are specific and meaningful, and avoids vague or ineffective responses.
[0072] Expression professionalism: tests whether the language style is close to the real doctor inquiry scene, focusing on whether the terminology used is accurate, the language structure is compact, and the expression is medically communicative.
[0073] Each dimension adopts a refined scoring mechanism from 0 to 5 points, and provides standard judgment basis. To further improve the objectivity of evaluation, a unified system prompt word is constructed to ensure that the large model scores under the premise of standard consistency. The prompt word is from the perspective of professional ophthalmologists, and scores the answers of the large language model in the ophthalmic consultation scene based on the strict and refined four indicators, emphasizing that the answers must be professional, accurate and compact in expression, and strictly prohibiting lengthy content and non-clinical style. The present application compares the average scores of the original baseline model and the model after supervised fine-tuning in the above four dimensions, and the results are shown in Table 3 (SFT model represents the model constructed based on the present application).
[0074] Table 3
[0075] The experimental results show that, compared with the original model, the ophthalmic field large language model proposed by the present application has achieved significant improvement in medical decision quality, information coverage, executable suggestion output and clinical language style fitting, and has shown good practical application potential.
[0076] To verify the discrimination ability and practical application guidance of the scoring system, the present application also sets up typical case comparison, covering basic ophthalmic terminology explanation (such as "aniseikonia", "disjunctive position"), structural cognition (such as "inner plexiform layer thickness") and treatment method description (such as "silicone oil filling operation"). The results are shown in Table 4, which shows that the model after fine-tuning not only can provide high information density, low redundancy and professional answers, but also can be more accurately aligned with expert consensus, avoiding excessive generalization and language style drift.
[0077] Table 4
[0078] As can be seen from the above embodiments, the present application integrates multi-source data and constructs text question and answer data through multi-stage intelligent process to ensure that the model can comprehensively understand the professional knowledge in the field of ophthalmology, laying a foundation for the model to achieve high precision and high generalization ability. The present application uses a large language model as an intelligent assistant to accurately screen a large amount of general medical data in the field of ophthalmology, and deeply cleanses high-noise Internet consultation data, improves the clinical relevance of training data, and avoids the model from learning irrelevant or inaccurate information, thereby ensuring the professionalism and reliability of the model output, which can improve the automation level and efficiency of data preprocessing. The present application fine-tunes the instructions on a large-scale high-quality ophthalmic field data set, and uses DeepSpeed ZeRO-3 for memory optimization. At the same time, multi-machine and multi-card distributed training is implemented to achieve efficient and scalable model training, solve the computing resource bottleneck of large model training, and significantly improve the performance of the model in ophthalmic report generation and professional question answering.
[0079] Compared with the prior art, the present application has the following technical effects: (1) significantly improving the professional ability of the ophthalmic large model: by constructing a large-scale high-quality ophthalmic data set and combining intelligent data processing and advanced fine-tuning training methods, the large model trained by the present application performs far better than existing general models in terms of professional knowledge and clinical inquiry assistance. Experimental results show that the fine-tuned model has a significant improvement in various evaluation indicators, and can provide more reliable intelligent auxiliary diagnosis and more accurate medical information services. (2) greatly improving data processing efficiency and quality: introducing large language models for data screening, question and answer pair generation, and high-noise data cleaning, realizing the automation and intelligentization of the data preparation process, significantly reducing the labor cost and time input, while ensuring the purity, professionalism and uniformity of the training data format, making the originally time-consuming and laborious data engineering link become efficient and feasible, accelerating the entire life cycle of the model from data to deployment. (3) optimizing the resource utilization and scalability of large model training: using DeepSpeedZeRO-3 optimization strategy and parameter efficient fine-tuning technology, effectively reducing the memory consumption and computing demand in the training process. The implementation of multi-machine multi-card distributed training further breaks through the bottleneck of single-machine training, realizes the scalability and acceleration of training, so that the method of the present application can adapt to the training needs of larger scale models in the future, and reduce the deployment and maintenance cost. (4) providing efficient and reliable auxiliary tools for ophthalmic clinical practice: the final model can provide intelligent auxiliary diagnosis suggestions for ophthalmologists, provide professional and personalized ophthalmic consultation services for patients, and serve as a powerful knowledge base for medical education and research. This will effectively improve the efficiency of ophthalmic diagnosis and treatment, reduce the workload of doctors, improve the patient experience and satisfaction, and promote the intelligent development of ophthalmic medicine.
[0080] It should be noted that although the above describes each step in a specific order, it does not mean that each step must be performed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the desired function can be achieved.
[0081] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for causing a processor to implement various aspects of the present application.
[0082] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves.
[0083] Embodiments of the application have been described above, with examples of the description being exemplary, but not exhaustive, and are not limited to the disclosed embodiments. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The scope of the embodiments described herein is not to be limited by the specific illustrative examples contained herein. The selection of the terms to be used in this document were chosen to best explain the principles of the embodiments, practical application, or technical improvement over the technology found in the marketplace, or to enable other of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for constructing a large-scale language model in the field of ophthalmology, wherein the large-scale language model in the ophthalmology field is used to output answers to user questions based on prompt words and user questions, characterized in that, The method comprises: S1, obtaining a plurality of ophthalmology field original data sets; S2, using a pre-configured first language large model to perform data screening on the original data sets obtained in step S1 to screen out data related to ophthalmology; S3, using the first language large model to perform data cleaning and standardization processing on the data screened out in step S2 to uniformly standardize different data into data samples in the structure of {prompt word, user question, question answer}; S4, collecting ophthalmology professional knowledge and constructing ophthalmology professional medical textbook fragments based thereon, using the first large language model to generate ophthalmology question and answer data related to the content of the fragments according to the ophthalmology professional medical textbook fragments, and uniformly standardizing the question and answer data into data samples in the structure of {prompt word, user question, question answer}; S5, using the first large language model to produce a plurality of groups of new data which are semantically equivalent but have different expression manners according to the data in step S4, and uniformly standardizing the new data into data samples in the structure of {prompt word, user question, question answer}; S6, grouping the data samples obtained in steps S3, S4 and S5 into a training set, taking the prompt word and the user question as input and the question answer as label, and performing supervised training on the base large language model to obtain an ophthalmology field large language model. 2.The method of claim 1, wherein, The plurality of ophthalmology field original data sets comprise: Medical, Chinese-medical-dialogue-data, cMedQA2, ChatMed_Consult_Dataset, Huatuo-Llama-Med-Chinese, HuatuoGPT-sft-data-v1, huatuo_medical_qa_sharegpt, DISC-Med-SFT, huatuo_encyclopedia_qa, Huatuo26M-Lite, MedDialog. 3.The method of claim 2, wherein, In the step S2, the first large language model is configured as a Qwen2.5-14B-Instruct large language model or DeepSeek, and is deployed in a local manner; The first large language model is configured to judge whether the data belongs to basic ophthalmology consultation or is ophthalmology knowledge, and to screen out data belonging to basic ophthalmology consultation and ophthalmology knowledge. 4.The method of claim 3, wherein, In the step S3, the first large language model is positioned as an ophthalmology consultation data information assistant by using the prompt word, and performs the following tasks on the data screened out in step S2 to perform data cleaning and standardization processing: S31, extracting valid conversation content from the data, and clearing irrelevant chitchat and information unrelated to diseases; S32, arranging multi-round conversations, labeling each round of speech as "user" or "doctor" and arranging in sequence; S33, analyzing the patient's complaint information, integrating to form a new instruction field including standard instructions and complaint summaries; S34, constructing the doctor's standard answer according to the last question of the user in the arranged multi-round conversation; S35, the final remaining data is standardized, and the data sample of the {prompt word, user question, question answer} structure is output, wherein only the question and answer content related to the disease, condition and treatment is retained in the data sample. 5.The method of claim 4, wherein, In the step S4, the first language large model is positioned as an ophthalmology field data assistant by using the prompt word, and a plurality of ophthalmology question and answer data related to the ophthalmology medical textbook fragment content is generated according to the ophthalmology professional medical textbook fragment. 6.The method of claim 5, wherein, The base large model is a Qwen2.5-14B-Instruct large language model. 7.The method of claim 6, wherein, In the step S6, a full parameter fine-tuning strategy is used when the base large language model is supervised trained, so as to supervise and adjust all the weights of the model while keeping the overall structure of the model unchanged. 8.The method of claim 7, wherein, In the step S6, the base model is trained in a distributed manner based on multiple machines and multiple cards, and a DeepSpeed ZeRO-3 optimization strategy is used.
9. An ophthalmic intelligent inquiry method based on a large language model, characterized in that, The method comprises: T1, obtaining the ophthalmology field large language model constructed by the method in any one of claims 1-8; T2, inputting the symptoms and questions complained by the inquiring person into the model obtained in the step T1 as the user question, and obtaining the question answer.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method in any one of claims 1-8.
11. An electronic device, comprising: Comprise: One or more processors; And A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method in any one of claims 1-8 by executing the executable instructions.