First-person film reading experience training method constructed based on mixed large model and eye movement tracker

By constructing a first-person video reading experience training method based on hybrid large models and eye trackers, the problem that new doctors find it difficult to learn specific video reading experience through AI-assisted diagnostic software is solved, and the effect of improving the new doctors' video reading ability and learning efficiency is achieved.

CN120103976APending Publication Date: 2025-06-06PEOPLES HOSPITAL OF HENAN PROV

Patent Information

Application Number
CN202510177663.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When new doctors conduct film reading training through existing AI-assisted diagnostic software, it is difficult to learn specific film reading experience and cannot learn specific film reading skills through the AI ​​diagnosis process.

Method used

A first-person video reading experience training method is adopted based on a hybrid big model and eye movement tracker. By obtaining EEG data and eye movement data from senior imaging physicians, a visual attention model and a generative model of imaging to diagnostic reports is constructed to form a hybrid big model to provide personalized video reading training.

Benefits of technology

Provide new doctors with a highly simulated video reading training environment, quantify the video reading process through eye movement technology, EEG measurement technology and deep learning technology, improve the objectivity of video reading evaluation, help new doctors learn and master the video reading habits and methods of senior imaging doctors, and improve their video reading ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103976A_ABST
    Figure CN120103976A_ABST
Patent Text Reader

Abstract

The invention discloses a first-person film reading experience training method constructed based on a mixed large model and an eye movement tracker, and relates to the technical field of medical technology.A first-person view angle film reading experience training system is constructed, a high-simulation film reading training environment is provided for a new doctor, and when the new doctor carries out film reading training, a first-person view angle film reading experience training system is provided for the new doctor; a new doctor is helped to know the reading level through reading evaluation, the new doctor is helped to learn and grasp the reading habit and method of a doctor of a qualification imaging department by guiding the new doctor to pay attention to the important attention area of the medical image, the reading ability of the new doctor is improved, the new doctor is helped to learn specific reading experience quickly and effectively, and the medical image reading efficiency is improved. An innovative method is provided for film reading experience training, and the method is high in practicability and suitable for popularization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical technology, and in particular to a first-person film reading experience training method. Background Art

[0002] The main contradiction currently faced by domestic radiologists is the conflict between insufficient number of doctors and excessive clinical demand. The distribution of radiologists between large tertiary hospitals and primary hospitals is seriously unbalanced. Doctors in large hospitals are busy dealing with a large number of imaging examinations every day, while radiologists in primary hospitals often lack professional guidance and training, which has led to a widening gap in medical technology levels.

[0003] This demand requires newly graduated imaging doctors to independently undertake the task of reading films as soon as possible. However, their film reading experience is inevitably insufficient, and it is often difficult to quickly transform theoretical knowledge into practical experience. Because the pathology in actual clinical practice is complex and changeable, and there are differences from the typical cases in textbooks, they will feel uncomfortable when reading films. Secondly, the accumulation of film reading experience requires a lot of practice, and the cultivation of diagnostic thinking also needs to be exercised through real cases. The equipment in different hospitals may be different, and the images produced by the examination may have subtle differences, which requires time to run in. In standardized training, the teaching teacher not only needs to train new doctors, but also has to undertake corresponding clinical tasks, making it difficult to conduct detailed and targeted explanations.

[0004] At present, most newly graduated imaging doctors improve their film reading experience through traditional methods. By participating in standardized residency training, new doctors can gradually master the basic principles and skills of film reading in theoretical learning and clinical practice. At the same time, they can find an experienced mentor for one-on-one guidance, who can impart film reading experience, help analyze cases, and point out misconceptions in diagnosis. In addition, actively participate in case discussions within the department and analyze typical and difficult cases with colleagues, so as to learn from others' diagnostic ideas and broaden their horizons. In actual work, increasing the amount of film reading, especially the practice of reading complex and rare cases, is the most direct and effective way to improve film reading experience. At the same time, regularly attending academic conferences and seminars to understand the latest developments and technological advances in the field of radiology and exchange experiences with peers can also effectively improve film reading ability.

[0005] Relevant emerging technologies can also be used to assist in film reading training. Through a variety of AI-assisted software for film reading and diagnosis, diagnostic efficiency can be improved and the clinical workload can be reduced. The application of teleradiology technology allows new doctors to collaborate across regions, share case resources, and be exposed to more types of cases. The popularization of digitalization and cloud computing technology makes medical imaging data easier to store, retrieve and analyze. New doctors can quickly obtain a large number of case resources for learning and practice through cloud computing platforms. Virtual reality (VR) and augmented reality (AR) technologies provide a new film reading experience, presenting image data in a three-dimensional way, which helps to understand the case more deeply. In addition, using Internet resources such as online courses, forums and professional social media, new doctors can learn radiology knowledge anytime and anywhere.

[0006] In summary, the traditional method of improving film reading experience is mainly through teaching by senior doctors and some emerging technologies for auxiliary training. Teaching by senior doctors relies on the time and energy of senior doctors. At the same time, due to manpower limitations, it is impossible to provide one-on-one guidance while facing busy clinical tasks, which affects the learning progress and quality of new doctors. The application of emerging technologies also has corresponding limitations. There are many AI-assisted diagnosis software at present, but the software only provides labels, such as benign or malignant. Its own judgment process is a black box and difficult to explain. Therefore, new doctors cannot learn specific film reading experience through the AI ​​diagnosis process.

[0007] The invention patent with application number 202211198780.6 discloses a method for analyzing and evaluating the ability to read medical images. The method uses an eye tracker to obtain the gaze data of the reader during the reading process, and obtains a gaze hotspot map and various indicators based on the gaze data; and compares the gaze hotspot map and various indicators with the gaze hotspot map and various indicators of the pathology experts when reading the same pathological images, to obtain an evaluation result for evaluating the reader's medical image reading ability. The judgment process is also difficult to explain. New doctors can only recognize their own reading level, but cannot learn specific reading experience to improve their reading ability. Summary of the invention

[0008] In view of the technical problem that it is difficult for new doctors to learn specific film reading experience when they are trained to read films through existing AI-assisted diagnosis software, the present invention proposes a first-person film reading experience training method based on a hybrid large model and an eye tracker, aiming to provide an innovative method for film reading experience training, so as to help new doctors learn specific film reading experience quickly and effectively.

[0009] In order to achieve the above object, the technical solution of the present invention is achieved as follows: A first-person film reading experience training method based on a hybrid large model and an eye tracker comprises the following steps: S1. When the senior radiologist is reading the film, the EEG data and eye movement data of the senior radiologist are obtained, and the eye movement data and EEG data of the senior radiologist are pre-processed respectively; S2. Build and train a visual attention model, form a preliminary attention heat map based on the pre-processed eye movement data and EEG data of senior radiologists, input the preliminary attention heat map into the trained visual attention model to obtain the personalized attention heat map of senior radiologists, collect the personalized attention heat maps of senior radiologists to produce a personalized attention heat map dataset of senior radiologists; S3. Based on the personalized attention heat map dataset of senior radiologists, collect the medical images and text data corresponding to the personalized attention heat map in the personalized attention heat map dataset, and build a text report training database; build an image-to-diagnosis report generation model, and use the text report training database to train the image-to-diagnosis report generation model to obtain a trained image-to-diagnosis report generation model; S4. Based on the trained visual attention model and the image-to-diagnosis report generation model, a hybrid large model is constructed and trained to obtain a trained hybrid large model. When a personalized attention heat map and a diagnosis report are input, the hybrid large model outputs an organ lesion level attention heat map, a standard report, and an attention block-description text pair; S5. Based on the training of the hybrid large model, a first-person perspective film reading experience training system is constructed, and the first-person perspective film reading experience training system is used to train new doctors in film reading.

[0010] Preferably, the eye movement data is acquired using an eye tracker, and the eye movement data includes gaze point, gaze duration and scanning path; the method for preprocessing the eye movement data of senior radiologists is: for the collected eye movement data, first use median filtering or Kalman filtering to remove noise in the eye movement data, and then perform standardization processing, combined with the eye tracker reference coordinate system, and refer to the eye tracker position offset and rotation to convert all eye movement data into a unified coordinate system.

[0011] Preferably, the method for preprocessing the EEG data of senior doctors is: analyzing the frequency components of the EEG signals through signal processing techniques including Fourier transform, splitting the EEG signals into band signals according to the frequency components, combining the split band signals with the original EEG signals, and inputting them into a machine learning model built based on support vector machines or random forests to obtain the concentration of EEG signals at different time periods.

[0012] Preferably, the method for forming a preliminary attention heat map is: generating attention blocks according to the gaze points and gaze duration in the eye movement data, adjusting the weights of the attention blocks in combination with the concentration of the EEG signals, and when the concentration of the EEG signals of a senior radiologist in a certain area is high, increasing the weight of the attention block corresponding to the area, and after completing the weight adjustment, obtaining a preliminary attention heat map.

[0013] Preferably, the text data includes diagnostic content specifications in medical textbooks related to medical imaging diagnosis, medical imaging diagnostic reports of senior radiologists and writing specifications of medical imaging diagnostic reports of hospitals; the medical imaging diagnostic reports of senior radiologists include case descriptions and diagnostic conclusions; before constructing a text report training database, the collected text data is cleaned to remove irrelevant information such as doctor's signature and date; then the text is segmented and professional terms are annotated; and the medical images are standardized including scaling, cropping and denoising.

[0014] Preferably, for medical image data, the standardized medical image is sent to a convolutional neural network, the convolutional neural network outputs image features in the form of feature maps, and the image features in the form of feature maps are converted into sequence form to obtain image feature vectors; for text data, the cleaned and segmented text is encoded to obtain a text embedding vector; the text embedding vector is concatenated with the image feature vector and used as the input of a generation model from image to diagnosis report, and the generation model from image to diagnosis report outputs a diagnosis report of the medical image.

[0015] Preferably, a logical reasoning module is integrated into the image-to-diagnosis report generation model. The logical reasoning module uses logical rules defined based on medical knowledge and the experience of senior doctors and quantifiable indicators such as disease probability and image feature weights to logically verify and reason the generated diagnostic report to obtain a standard report that conforms to clinical diagnostic logic.

[0016] Preferably, the basic architecture of the hybrid large model is built based on Transformer, with a multi-modal input fusion layer, and a multi-head attention mechanism and a self-attention mechanism are introduced; when training the hybrid large model, the visual attention model and the image-to-diagnosis report generation model are trained separately first, and then the hybrid large model is jointly trained using a multi-task learning framework.

[0017] Preferably, when using the first-person perspective film reading experience training system to train new doctors, the first-person perspective film reading experience training system presents medical images for film reading training to the new doctors, and generates organ lesion level attention heat maps and standard reports corresponding to the medical images used for film reading training. When the new doctor is reading the film, the hot spots observed by the new doctor are captured in real time, and the voice descriptions of the new doctor's attention to different areas are recorded. By matching the hot spots observed by the new doctor with the organ lesion level attention heat maps corresponding to the medical images used for film reading training, it is determined whether the area observed by the new doctor is a high-hot spot attention area. If the new doctor does not pay attention to the high-hot spot attention area for a long time when reading the film, automatic guidance is performed; the voice descriptions of the new doctor's attention to different areas are converted into text and matched with the text in the standard report. The attention areas covered by the film reading and the matching degree of the voice description are comprehensively considered to perform film reading evaluation, quantify the film reading process, and conduct film reading training for the new doctor.

[0018] Preferably, the film reading evaluation includes attention Figure 1 consistency score, speech description consistency score and comprehensive score; among them, attention Figure 1 The expression of consistency score is: NMSE = (Σ(|A doctor - A model|²) / Σ(A model²)) Among them, NMSE is the normalized mean square error, Doctor A represents the attention heat map of the new doctor, and Model A represents the organ lesion level attention heat map generated by the hybrid large model; The expression of speech description consistency score is: Recall = (number of correctly identified positives / total number of positive samples) Among them, Recall is the recall rate.

[0019] Compared with the prior art, the present invention has the following beneficial effects: the present invention provides new doctors with a highly simulated film reading training environment by constructing a first-person perspective film reading experience training system. It quantifies the film reading process and improves the objectivity of film reading evaluation through eye movement technology, EEG measurement technology and deep learning technology. While helping new doctors to understand their own film reading level, it can also effectively assist new doctors to learn and master the film reading habits and methods of senior radiologists, improve the new doctors' film reading ability, and help new doctors to quickly and effectively learn specific film reading experience. It provides an innovative method for film reading experience training, which is highly practical and suitable for promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 It is a flow chart of the present invention.

[0022] Figure 2 This is a schematic diagram of training new doctors using a first-person perspective film reading experience training system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] like Figure 1 As shown, a first-person film reading experience training method based on a hybrid large model and an eye tracker includes the following steps: S1. When the senior radiologist is reading the film, the EEG data and eye movement data of the senior radiologist are obtained at the same time, and the eye movement data and EEG data of the senior radiologist are pre-processed respectively.

[0025] Furthermore, the eye movement data is obtained using an existing wearable eye tracker, such as Tobii Pro or SMIEye Tracking. When sampling, the sampling rate is set to at least 120Hz to improve the accuracy and reliability of the eye movement data and ensure that the subtle eye movement changes of senior radiologists when reading films can be captured. The eye movement data of senior doctors when reading films is detected in real time, including gaze point, gaze duration and scanning path, and the software built into the eye tracker is used for preliminary processing: the collected eye movement data is firstly removed from the noise by median filtering or Kalman filtering, and then standardized, combined with the eye tracker reference coordinate system, and with reference to the position offset and rotation of the eye tracker, all eye movement data are converted to a unified coordinate system for subsequent analysis and application.

[0026] Furthermore, the method for acquiring EEG data is as follows: when senior radiologists read films, use an acquisition device with good anti-interference ability and high temporal resolution, combine EEG (Electroencephalogram) and fNIRS (functional near-infrared spectroscopy) technology to obtain detailed brain activity information of senior radiologists, record the EEG signals of senior radiologists in real time, focus on the EEG bands α and β waves related to attention and cognitive load, analyze the frequency components of the EEG signals through signal processing technology including Fourier transform, so as to split the EEG signals into different band signals according to the frequency components, combine the split band signals with the original EEG signals, input them into a machine learning model constructed based on support vector machines or random forests, etc., identify the concentration of EEG signals in different time periods, so as to mark the high-attention areas later. Since the EEG data is collected synchronously with the eye movement data, the two are aligned in time sequence, and no other processing is required.

[0027] S2. Construct and train a visual attention model, form a preliminary attention heat map based on the preprocessed eye movement data and EEG data of senior radiologists, input the preliminary attention heat map into the trained visual attention model to obtain the personalized attention heat map of senior radiologists, collect the personalized attention heat maps of senior radiologists to produce a personalized attention heat map dataset of senior radiologists.

[0028] The visual attention model combines eye movement data and EEG data through multimodal data fusion technology. First, an attention model such as an attention mechanism network is used to combine the gaze points in the eye movement data and the concentration in the EEG data to generate a preliminary attention heat map. The method is: attention blocks are generated according to the gaze points and gaze duration in the eye movement data, and the weights of the attention blocks are adjusted in combination with the intensity of the EEG signal to obtain a preliminary attention heat map. Specifically, when a senior radiologist shows a high degree of concentration in a certain area (this concentration is obtained by identifying the EEG data of the senior radiologist through the machine learning model in step S1), the weight of the attention block corresponding to the area will be increased accordingly, thereby more accurately reflecting the distribution of the doctor's concentration.

[0029] The visual attention model takes into account individual differences, including personal information such as gender, age, and years of experience in reading films, and assigns personalized weights to the fixation points. To this end, the visual attention model takes personal information and a preliminary attention heat map as input, feeds it into a convolutional neural network (CNN) for processing, and ultimately generates a personalized attention heat map.

[0030] Specifically, the convolutional neural network is an end-to-end structure, including an encoder and a decoder. The encoder is used to extract features, and the decoder is used to restore the original input size. The preliminary attention heat map is used as the input of the convolutional neural network, and the features are extracted through the convolution layer and the pooling layer to obtain the feature map. The personal information encoding vector is added to the feature map, such as gender, age, years of reading experience, etc. after convolution encoding. After the information is fused by splicing or element addition, the feature map is gradually restored to the original size through the deconvolution layer and the upsampling layer; finally, an activation function, such as ReLU or Sigmoid, is applied to generate a personalized attention heat map.

[0031] The visual attention model is trained using a large amount of labeled data, the cross entropy loss function is used to evaluate the performance of the visual attention model, and the gradient descent algorithm and optimizer (such as Adam) are applied to update the weights. During the training process, the input data of the visual attention model is enhanced and regularization technology is introduced to enhance the robustness of model training by randomly adding noise to prevent overfitting of the visual attention model.

[0032] S3. Based on the personalized attention heat map dataset of senior radiologists, collect the medical images and text data corresponding to the personalized attention heat maps in the personalized attention heat map dataset, and build a text report training database; build an image-to-diagnosis report generation model, and use the text report training database to train the image-to-diagnosis report generation model to obtain a trained image-to-diagnosis report generation model.

[0033] The text data includes diagnostic content specifications in medical textbooks related to medical imaging diagnosis, medical imaging diagnosis reports of senior radiologists and writing specifications of medical imaging diagnosis reports of hospitals; the medical imaging diagnosis reports of senior radiologists include case descriptions and diagnostic conclusions.

[0034] First, the collected text data is cleaned to remove irrelevant information such as doctor's signature and date. Then the text is segmented and professional terms are annotated. Medical images are standardized including scaling, cropping and denoising to facilitate subsequent processing and analysis of text data and medical images.

[0035] The data collection also needs to collect medical case reports related to medical images, extract medical images and text descriptions, and expand the clinical atypical cases in the text report training database to improve the generalization ability of the generation model from images to diagnostic reports.

[0036] For medical image data, the standardized medical images are sent to the convolutional neural network (CNN), and the image features are extracted by CNN. The image features are output by CNN in the form of feature maps and converted into sequence form to obtain image feature vectors. For text data, the cleaned and segmented text is encoded to obtain a text embedding vector. After the text embedding vector and the image feature vector are concatenated, they are used as the input of the image-to-diagnosis report generation model, which outputs the medical image diagnosis report.

[0037] The image-to-diagnosis report generation model uses a Transformer-based model architecture, such as BERT (Bidirectional Encoder Representations from Transformers) or ViT (Vision Transformer), to handle the fusion of medical image and text data. The Transformer model captures the complex relationship between medical images and text through a multi-head self-attention mechanism.

[0038] In order to ensure that the generated diagnostic report conforms to medical logic, diagnostic logic rules are defined based on medical knowledge and the experience of senior doctors, and the logic rules are converted into quantifiable indicators such as disease probability and image feature weights. Integrate logical reasoning modules, such as conditional generative networks (CGNs), into the image-to-diagnosis report generation model. The logical reasoning module uses the previously defined logical rules and quantifiable indicators to perform logical verification and reasoning on the generated diagnostic report to ensure that the standard report generated by the image-to-diagnosis report generation model conforms to clinical diagnostic logic.

[0039] When training the image-to-diagnosis report generation model, data enhancement is performed on medical image data including rotation, flipping, and scaling, and on text data including synonym replacement and sentence reorganization to increase the generalization ability of the image-to-diagnosis report generation model. The cross-entropy loss function is used to evaluate the model performance of the image-to-diagnosis report generation model, and the AdamW optimizer is applied in combination with weight decay regularization. A fine-tuning strategy is adopted to ensure that the input under small sample data still has a high accuracy by directly calling the Transformer weights that have been pre-trained on a large number of public datasets and training directly on the pre-trained Transformer model. Gradient accumulation and mixed precision training are implemented to accelerate the training process.

[0040] S4. Based on the trained visual attention model and the image-to-diagnosis report generation model, a hybrid large model is constructed and trained to obtain a trained hybrid large model. The input of the hybrid large model is the personalized attention heat map generated by the visual attention model in step S2 and the diagnosis report generated by the image-to-diagnosis report generation model in step S3. The output of the hybrid large model is the organ lesion level attention heat map, the standard report, and the attention block-description text pair.

[0041] The hybrid large model supports the generation of organ lesion level attention heatmaps. The hybrid large model uses the output of the visual attention model to determine the area that the doctor focuses on when reading the film through high-precision eye tracking data. Combined with segmentation techniques in deep learning, such as U-Net, medical images are segmented into organ and lesion areas. The attention blocks are matched with the segmentation results to generate an organ lesion level attention heatmap. The hybrid large model supports the construction of attention block-description text pairs. The hybrid large model extracts the text description corresponding to the attention block in the diagnostic report generated by the generation model of the image to the diagnostic report. Use natural language processing (NLP) techniques, such as entity recognition and relationship extraction, to identify the descriptions of organs and lesions in the diagnostic report.

[0042] The basic architecture of the hybrid model is built on the basis of Transformer, and a multimodal input fusion layer is designed to fuse image features and text features. The hybrid model is enhanced to understand multimodal data by using methods such as feature concatenation or feature interaction, such as multi-head attention mechanism. The self-attention mechanism is introduced into the Transformer model, so that the hybrid model can learn the association between the attention block and the text description. The weights of the attention block and the text description are dynamically adjusted using soft attention or self-attention. A multi-task loss function is designed, including the loss of attention heatmap generation, the loss of text generation, and the loss of attention block-description text pair matching. A weighted loss function is used to assign different weights according to the importance of different tasks. A progressive training strategy is adopted to first train the visual attention model and the image-to-diagnosis report generation model separately, and then the hybrid model is jointly trained using a multi-task learning framework, and the organ lesion level attention heatmap and standard report generation are trained at the same time. An evaluation framework is developed, including quantitative evaluation (such as IoU, precision, recall) and qualitative evaluation (such as doctor review). Stress testing was performed on the hybrid large model to improve its stability and accuracy in handling complex cases.

[0043] S5. Based on the training of the hybrid large model, a first-person perspective film reading experience training system is constructed. The first-person perspective film reading experience training system presents medical images for film reading training to new doctors, and generates organ lesion level attention heat maps and standard reports corresponding to the medical images used for film reading training. When the new doctor is reading the film, the first-person perspective is used to capture the hot spots observed by the new doctor in real time with the eye tracking technology, and at the same time, the voice description of the new doctor when paying attention to different areas is recorded, and the film reading evaluation is made after matching with the organ lesion level attention heat maps and standard reports corresponding to the medical images used for film reading training; if the new doctor does not pay attention to the high hot spots in the organ lesion level attention heat maps corresponding to the medical images used for film reading training for a long time when reading the film, automatic guidance will be performed, so as to train the new doctor so that the new doctor can quickly grow into a senior new doctor in the imaging department.

[0044] Specifically, when the new doctor is reading the film, the new doctor's visual focus is captured in real time by the method described in step S1. By integrating advanced speech recognition technology, such as a deep learning speech recognition model, the new doctor's voice description is converted into text data in real time, and the voice data is preprocessed, including noise reduction, sentence segmentation, word segmentation, etc., so as to capture the new doctor's voice description when focusing on different areas in real time. The organ lesion level attention heat map generated by the hybrid large model and the standard report are used as a reference, and the attention area covered by the film reading and the matching degree of the voice description are comprehensively evaluated to quantify the film reading process. According to the new doctor's real-time film reading situation, if the high hot spot area has not been paid attention to for a long time, automatic guidance is performed.

[0045] Furthermore, when training new doctors, the organ lesion-level attention heat map generated by the hybrid large model is used as a reference to detect the new doctor's attention area in real time. When matching it with the organ lesion-level attention heat map generated by the hybrid large model, it is also necessary to dynamically adjust the score weight of the attention block-description text pair according to the new doctor's local attention and global attention level, and the score weight of the key diagnosis area is higher than that of the non-critical area.

[0046] The film reading evaluation includes attention Figure 1 consistency score, speech description consistency score and comprehensive score.

[0047] attention Figure 1 Consistency score: The normalized mean square error (NMSE) is used to calculate the difference between the new doctor's attention heat map generated by the new doctor's reading and the organ lesion-level attention heat map corresponding to the medical image generated by the hybrid large model for reading training.

[0048] NMSE = (Σ(|A doctor - A model|²) / Σ(A model²)) Among them, Doctor A represents the attention heat map of the new doctor, and Model A represents the organ lesion level attention heat map generated by the hybrid large model.

[0049] Speech description consistency score: Natural language processing techniques such as word embedding and sequence matching are used to calculate the consistency between the new doctor's speech description and the standard report corresponding to the medical image output by the hybrid large model for film reading training. The accuracy of the new doctor's speech description is evaluated by recall.

[0050] Recall = (number of correctly identified positives / total number of positive samples) Comprehensive scoring: Design a comprehensive scoring algorithm that combines attention Figure 1 Use weighted average or other composite scoring methods, such as F1 score, to comprehensively analyze the attention heat Figure 1 The consistency score and the speech description consistency score were used to obtain a comprehensive score.

[0051] The first-person perspective film reading experience training system can provide a highly simulated film reading training environment. By real-time detection and matching of new doctors' attention blocks and description text pairs, and calculating consistency scores, it can effectively assist new doctors in learning and mastering the film reading habits and methods of senior radiologists. It not only improves the objectivity of film reading evaluation, but also promotes the rapid growth of new doctors.

[0052] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A first-person film reading experience training method based on a hybrid large model and an eye tracker, characterized in that: The following steps are involved: S1. When the senior radiologist is reading the film, the EEG data and eye movement data of the senior radiologist are obtained, and the eye movement data and EEG data of the senior radiologist are pre-processed respectively; S2. Build and train a visual attention model, form a preliminary attention heat map based on the pre-processed eye movement data and EEG data of senior radiologists, input the preliminary attention heat map into the trained visual attention model to obtain the personalized attention heat map of senior radiologists, collect the personalized attention heat maps of senior radiologists to produce a personalized attention heat map dataset of senior radiologists; S3. Based on the personalized attention heat map dataset of senior radiologists, collect the medical images and text data corresponding to the personalized attention heat map in the personalized attention heat map dataset, and build a text report training database; build an image-to-diagnosis report generation model, and use the text report training database to train the image-to-diagnosis report generation model to obtain a trained image-to-diagnosis report generation model; S4. Based on the trained visual attention model and the image-to-diagnosis report generation model, a hybrid large model is constructed and trained to obtain a trained hybrid large model. When a personalized attention heat map and a diagnosis report are input, the hybrid large model outputs an organ lesion level attention heat map, a standard report, and an attention block-description text pair; S5. Based on the training of the hybrid large model, a first-person perspective film reading experience training system is constructed, and the first-person perspective film reading experience training system is used to train new doctors in film reading.

2. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 1, characterized in that: The eye movement data is obtained by using an eye tracker, and the eye movement data includes gaze point, gaze duration and scanning path; the method for preprocessing the eye movement data of senior radiologists is: for the collected eye movement data, first remove the noise in the eye movement data by median filtering or Kalman filtering, and then perform standardization processing, combined with the eye tracker reference coordinate system, and refer to the eye tracker position offset and rotation, to convert all eye movement data into a unified coordinate system.

3. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 1 or 2, characterized in that: The method for preprocessing the EEG data of senior doctors is as follows: analyze the frequency components of the EEG signals through signal processing techniques including Fourier transform, split the EEG signals into band signals according to the frequency components, combine the split band signals with the original EEG signals, and input them into a machine learning model built based on support vector machines or random forests to obtain the concentration of EEG signals at different time periods.

4. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 3 is characterized in that: The method for forming a preliminary attention heat map is: generating attention blocks according to the gaze points and gaze duration in the eye movement data, adjusting the weights of the attention blocks in combination with the concentration of the EEG signals, and when the concentration of the EEG signals of a senior radiologist in a certain area is high, increasing the weight of the attention block corresponding to the area, and after completing the weight adjustment, obtaining a preliminary attention heat map.

5. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 4 is characterized in that: The text data includes diagnostic content specifications in medical textbooks related to medical imaging diagnosis, medical imaging diagnosis reports of senior radiologists and writing specifications of medical imaging diagnosis reports of hospitals; the medical imaging diagnosis reports of senior radiologists include case descriptions and diagnostic conclusions; before constructing a text report training database, the collected text data is cleaned to remove irrelevant information such as doctor's signature and date; then the text is segmented and professional terms are annotated; and the medical images are standardized including scaling, cropping and denoising.

6. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 5, characterized in that: For medical image data, the standardized medical images are sent to the convolutional neural network, which outputs image features in the form of feature maps. The image features in the form of feature maps are converted into sequence form to obtain image feature vectors. For text data, the cleaned and segmented text is encoded to obtain a text embedding vector. The text embedding vector is concatenated with the image feature vector and used as the input of the image-to-diagnosis report generation model. The image-to-diagnosis report generation model outputs a diagnosis report of the medical image.

7. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 6, characterized in that: The image-to-diagnosis report generation model integrates a logical reasoning module, which uses logical rules defined based on medical knowledge and the experience of senior doctors and quantifiable indicators such as disease probability and image feature weights to logically verify and reason the generated diagnosis report to obtain a standard report that conforms to clinical diagnosis logic.

8. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 1 or 7, characterized in that: The basic architecture of the hybrid large model is built on the basis of Transformer, with a multi-modal input fusion layer, and also introduces a multi-head attention mechanism and a self-attention mechanism. When training the hybrid large model, the visual attention model and the image-to-diagnosis report generation model are trained separately first, and then the hybrid large model is jointly trained using a multi-task learning framework.

9. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 8, characterized in that: When using the first-person perspective film reading experience training system to train new doctors, the first-person perspective film reading experience training system presents medical images for film reading training to the new doctors, and generates organ lesion level attention heat maps and standard reports corresponding to the medical images used for film reading training. When the new doctor is reading the film, the hot spots observed by the new doctor are captured in real time, and the voice descriptions of the new doctor's attention to different areas are recorded. By matching the hot spots observed by the new doctor with the organ lesion level attention heat maps corresponding to the medical images used for film reading training, it is determined whether the area observed by the new doctor is a high-hot spot attention area. If the new doctor does not pay attention to the high-hot spot attention area for a long time when reading the film, automatic guidance is performed; the voice descriptions of the new doctor's attention to different areas are converted into text and matched with the text in the standard report. The attention areas covered by the film reading and the matching degree of the voice description are comprehensively considered to evaluate the film reading, quantify the film reading process, and conduct film reading training for the new doctor.

10. The first-person film reading experience training method based on a hybrid large model and an eye tracker according to claim 9, characterized in that: The film reading evaluation includes an attention map consistency score, a speech description consistency score and a comprehensive score; wherein the expression of the attention map consistency score is: NMSE = (Σ(|A doctor - A model|²) / Σ(A model²)) Among them, NMSE is the normalized mean square error, Doctor A represents the attention heat map of the new doctor, and Model A represents the organ lesion level attention heat map generated by the hybrid large model; The expression of speech description consistency score is: Recall = (number of correctly identified positives / total number of positive samples) Among them, Recall is the recall rate.

Citation Information

Patent Citations

  • Medical image reading ability analysis and evaluation method

    CN115564220A

Cited By

  • Data annotation method and system based on user behavior and attention tracking

    CN120929832A

  • A data labeling method and system based on user behavior and attention tracking

    CN120929832B

  • Morphological teaching system and method based on dynamic image generation

    CN121033591A