Newborn fundus image description generation method and imaging method

By constructing a neonatal fundus image description generation model combining image feature extraction network, feature mapping network and description generation module, the problem that the prior art cannot effectively interpret the fundus image of neonatal is solved, and efficient and reliable image description generation is achieved.

CN119992242AActive Publication Date: 2025-05-13CENT SOUTH UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510480474.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing general image description generation scheme cannot effectively adapt to the characteristics of neonatal fundus images, resulting in low interpretation efficiency, low reliability and consistency.

Method used

A technical solution combining image feature extraction network, feature mapping network and description generation module is adopted to construct a newborn fundus image description generation model, and through feature extraction, mapping and description generation modules, the description of the newborn fundus image is automatically generated.

Benefits of technology

The reliability and accuracy of the image description of the fundus of the neonatal baby is improved, the dependence of manual interpretation is reduced, and the efficiency and consistency of image interpretation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992242A_ABST
    Figure CN119992242A_ABST
Patent Text Reader

Abstract

The invention discloses a newborn fundus image description generation method and an imaging method. The newborn fundus image description generation method comprises the steps of obtaining an existing newborn fundus image, and performing description and preprocessing to construct a training data set; based on the feature extraction network, the feature mapping network and the large language model, constructing a newborn fundus image description generation initial model, and training to obtain a newborn fundus image description generation model; and automatically generating the neonatal fundus image description by adopting the obtained neonatal fundus image description generation model. According to the newborn eye fundus image description generation method and the imaging method provided by the invention, the technical scheme of combining feature extraction, feature mapping and a large language model is adopted, the corresponding image feature extraction network and the feature mapping network are designed for the newborn eye fundus image, and the description generation scheme is adopted, so that the imaging accuracy is improved. Generation and imaging of neonatal fundus image description are achieved in a targeted mode, the reliability is higher, and the accuracy is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of digital image processing, and in particular relates to a method for generating description of a neonatal fundus image and an imaging method. Background Art

[0002] Fundus images of newborns are of great significance in both clinical and basic medical research. However, in the current clinical and basic medical application processes, fundus images of newborns need to be interpreted by professional medical imaging personnel or ophthalmologists in advance. This manual interpretation process is not only time-consuming and laborious, but also inefficient, and has average reliability and consistency.

[0003] At present, researchers have proposed a scheme for generating descriptions of general images. However, this type of generation scheme is for general images, while the fundus images of newborns have their own characteristics, such as sparse and thin blood vessels, large optic nerve heads and blurred edges. These characteristics make the existing general image description generation scheme not suitable for the generation of fundus image descriptions of newborns. Summary of the invention

[0004] One of the purposes of the present invention is to provide a method for generating descriptions of neonatal fundus images with high reliability and good accuracy.

[0005] A second object of the present invention is to provide an imaging method including the method for generating a description of a neonatal fundus image.

[0006] The method for generating a description of a fundus image of a newborn provided by the present invention comprises the following steps: S1. Obtain existing fundus images of newborns; S2. describing and preprocessing the neonatal fundus image obtained in step S1 to construct a training data set; S3. Based on the image feature extraction network, feature mapping network and description generation module, an initial model for generating descriptions of neonatal fundus images was constructed; The initial model for generating descriptions of neonatal fundus images includes an image feature extraction network, a feature mapping network and a description generation module which are connected in series in sequence; the image feature extraction network is used to extract features from input image data; the feature mapping network is used to map the extracted features to corresponding image descriptions; the description generation module is used to generate a description of the input image according to the obtained image features and corresponding text prompt words; S4. Using the training data set obtained in step S2, based on the likelihood function, the initial model for generating the fundus image description of the neonate constructed in step S3 is trained to obtain a fundus image description generation model for the neonate; S5. Using the neonatal fundus image description generation model obtained in step S4, automatically generate neonatal fundus image descriptions.

[0007] The step S2 specifically includes the following steps: Describe the fundus images of newborn babies of various set types obtained in step S1; The described fundus image of the newborn is preprocessed; the preprocessing comprises the following steps: For each set type, the fundus images of newborns of the type with a number of images less than the first set value are repeated several times; the fundus images of newborns of the type with a number of images greater than the second set value are discarded according to the set ratio; The obtained fundus image of the newborn is adjusted to a set resolution; The neonatal fundus image with adjusted resolution is divided into nn square images; the neonatal fundus image with adjusted resolution is reduced to the same size as the divided square images; finally, an original neonatal fundus image is preprocessed into nn+1 images.

[0008] The step S3 comprises the following steps: An image feature extraction network is constructed based on the SigLip model to extract features from input image data; A feature mapping network is constructed based on a multi-layer perceptron model to map the extracted features to the corresponding image descriptions; A description generation module is constructed based on the Qwen2 model to generate a description of the input image according to the obtained image features and the corresponding text prompt words.

[0009] The image feature extraction network based on the SigLip model is constructed, and specifically comprises the following steps: The SigLip model is used to construct an image feature extraction network. The processing process of the image feature extraction network is expressed as: Where Z is the feature extraction result; X is the input image; The parameters are The processing function of the SigLip model; An original neonatal fundus image and the corresponding preprocessed n images are collectively represented as ,in is the compressed original fundus image of the newborn. is the corresponding preprocessed n images; nn+1 images , are input into the image feature extraction network to obtain the feature extraction results for ,in for The feature extraction results are for The feature extraction results.

[0010] The construction of the feature mapping network based on the multi-layer perceptron model specifically includes the following steps: The input of the feature mapping network is the feature extraction result output by the image feature extraction network ; The feature extraction results for the input , the feature mapping network adopts The convolutions are aligned and pooled in the spatial dimension to reduce the spatial dimension; for The jth feature extraction result in , and ; Map the pooled feature data to the target feature dimension and flatten the three-dimensional features into two-dimensional features to obtain the flattened image features. ; The flattened image features Input a multilayer perceptron with two linear layers for feature mapping, expressed as: In the formula The feature extraction result output by the first linear layer The hidden state of is the activation function, and ; is the projection matrix of the first linear layer; is the bias term of the first linear layer; Feature extraction result output by multi-layer perceptron Image features; is the projection matrix of the second linear layer; is the bias term of the second linear layer; The feature extraction results obtained by the image feature extraction network Both are input into the feature mapping network to get the corresponding output for ; for The output of the corresponding feature mapping network, for The output of the corresponding feature mapping network; Will get Splicing is performed to obtain the features of the neonatal fundus image in the text space , and is used as input to the description generation module in the form of a token sequence.

[0011] The Qwen2 model-based description generation module specifically includes the following steps: The Qwen2 model is used to construct the description generation module; The input text instruction is passed through the embedding layer to get the text embedding ; Embed text and Perform splicing and generate a description of the input image using an autoregressive approach; The autoregressive method comprises the following steps: During the generation process, the previously generated value is recursively used to generate the next value until the generation process ends; During the generation process, the current position prediction value is expressed as , all generated values ​​before the current position are expressed as , then the generation process is expressed as: in represents the probability of generating a description, and , L is the length of the generated description, To predict the input command before the position, The generated content before the predicted position.

[0012] The training described in step S4 specifically includes the following steps: Through the training method of efficient parameter fine-tuning, the likelihood function of the model is maximized; the likelihood function within a batch is expressed as: In the formula For the parameters The probability of ; B is the size of the batch; is the current position prediction value of the b-th sample in the corresponding batch; is the image of the bth sample in the corresponding batch ; is the input instruction before the predicted position of the bth sample in the corresponding batch, is the generated content before the predicted position of the b-th sample in the corresponding batch; In order to improve the generalization ability of the model, the gradient accumulation method is used to average the gradients obtained in K batches; K is the set batch value; The training method for efficient parameter fine-tuning specifically includes the following contents: The initial model generated by describing the fundus image of the newborn constructed in step S3 is used as the primary model, and the corresponding primary model parameters are expressed as ; During the training process, by training a parameter relative to the primary model The parameter set is less than the third set value , to encode the increments of the primary model while keeping the original weights unchanged; add a trainable low-rank matrix factorization to the linear layer of the model; during training, the input data passes through the linear layer, and the gradient update only acts on the low-rank matrix, while keeping the parameters of the initial model unchanged; in the inference phase, the incremental weights obtained from training are merged into the initial model to achieve more efficient parameter updates; The update formula of the model weight is expressed as: In the formula are the updated model parameters; The trainable parameters Encoded model weight increments; When updating the model weights, the weights of the feature mapping network and the description generation module are updated: for the linear layer, the weight matrix is ​​represented as W, the bias is represented as b, and the increment of the weight matrix is ​​represented as , assuming that the input of the linear layer is z, then after efficient parameter fine-tuning, the forward propagation of the linear layer is expressed as: Where h is the output of the linear layer; is a scaling factor used to adjust the value of the weight increment. Through this efficient parameter fine-tuning mechanism, while ensuring the efficient and adjustable model parameters, it can reduce the computing and storage costs, making the fine-tuning of large-scale models more efficient.

[0013] The present invention also provides an imaging method including the method for generating description of a neonatal fundus image, further comprising the following steps: S6. The neonatal fundus image description content generated in step S5 is marked and re-imaged on the neonatal fundus image to obtain a neonatal fundus image with the fundus image description content.

[0014] The method for generating descriptions of fundus images of newborns and the imaging method provided by the present invention combine the technical solutions of feature extraction, feature mapping and description generation modules, design corresponding image feature extraction networks and feature mapping networks for fundus images of newborns, and adopt a description generation scheme, which not only realizes the generation of descriptions of fundus images of newborns in a targeted manner, but also has higher reliability and better accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The figure is a schematic diagram of a method flow of the generation method of the present invention.

[0016] Figure 2 The figure is a schematic diagram of the method flow of the imaging method of the present invention. DETAILED DESCRIPTION

[0017] like Figure 1 The method flow chart of the generation method of the present invention is shown as follows: The method for generating the description of the fundus image of a newborn disclosed in the present invention comprises the following steps: S1. Obtain existing fundus images of newborns; In specific implementation, the preferred solution is to obtain multiple types of neonatal fundus images, including normal neonatal fundus images and various types of abnormal neonatal fundus images; S2. Describing and preprocessing the neonatal fundus image obtained in step S1 to construct a training data set; specifically comprising the following steps: Describe the fundus images of newborn babies of various set types obtained in step S1; The described fundus image of the newborn is preprocessed; the preprocessing comprises the following steps: For each set type, the fundus images of newborns of the type with a number of images less than the first set value are repeated several times; the fundus images of newborns of the type with a number of images greater than the second set value are discarded according to the set ratio; The obtained fundus image of the newborn is adjusted to a set resolution; The fundus image of the newborn after adjusting the resolution is divided into n square images; the fundus image of the newborn after adjusting the resolution is reduced to the same size as the square image after segmentation; finally, an original fundus image of the newborn is preprocessed into n+1 images; a preferred solution is to input Resolution image, resized to resolution, and then cut into square image, and correspondingly reduce the original image to , that is, processing a newborn fundus image into 13 square images.

[0018] S3. Based on the image feature extraction network, feature mapping network and large language model, an initial model for generating descriptions of neonatal fundus images was constructed; The initial model for generating descriptions of neonatal fundus images includes an image feature extraction network, a feature mapping network and a description generation module which are connected in series in sequence; the image feature extraction network is used to extract features from input image data; the feature mapping network is used to map the extracted features to corresponding image descriptions; the description generation module is used to generate a description of the input image according to the obtained image features and corresponding text prompt words; The specific implementation includes the following steps: An image feature extraction network is constructed based on the SigLip model to extract features from input image data; A feature mapping network is constructed based on a multi-layer perceptron model to map the extracted features to the corresponding image descriptions; A description generation module is constructed based on the Qwen2 model to generate a description of the input image according to the obtained image features and the corresponding text prompt words.

[0019] The construction of the image feature extraction network based on the SigLip model specifically includes the following steps: The SigLip model is used to construct an image feature extraction network, and the network processing process is expressed as: Where Z is the feature extraction result; X is the input image; The parameters are The processing function of the SigLip model; An original neonatal fundus image and the corresponding preprocessed n images are collectively represented as ,in is the compressed original fundus image of the newborn. is the corresponding preprocessed n images; nn+1 images , are input into the image feature extraction network to obtain the feature extraction results for ,in for The feature extraction results are for The feature extraction results.

[0020] The construction of the feature mapping network based on the multi-layer perceptron model specifically includes the following steps: The input of the feature mapping network is the feature extraction result output by the image feature extraction network ; The feature extraction results for the input , the feature mapping network adopts The convolutions are aligned and pooled in the spatial dimension to reduce the spatial dimension; is the jth feature extraction result in v, and ; Map the pooled feature data to the target feature dimension and flatten the three-dimensional features into two-dimensional features to obtain the flattened image features. ; The flattened image features Input a multilayer perceptron with two linear layers for feature mapping, expressed as: In the formula The feature extraction result output by the first linear layer The hidden state of is the activation function, and ; is the projection matrix of the first linear layer; is the bias term of the first linear layer; Feature extraction result output by multi-layer perceptron Image features; is the projection matrix of the second linear layer; is the bias term of the second linear layer; The activation function is used to introduce nonlinear transformations in the multi-layer perceptron to enhance the expressiveness of the model; The feature extraction results obtained by the image feature extraction network Both are input into the feature mapping network to get the corresponding output for ; for The output of the corresponding feature mapping network, for The output of the corresponding feature mapping network; Will get Splicing is performed to obtain the features of the neonatal fundus image in the text space , and is used as input to the description generation module in the form of a token sequence.

[0021] The Qwen2 model-based description generation module specifically includes the following steps: The Qwen2 model is used to construct the description generation module; The input text instruction is passed through the embedding layer to get the text embedding ; Embed text and Perform splicing and generate a description of the input image using an autoregressive approach; The autoregressive method comprises the following steps: During the generation process, the previously generated value is recursively used to generate the next value until the generation process ends; During the generation process, the current position prediction value is expressed as , all generated values ​​before the current position are expressed as , then the generation process is expressed as: in represents the probability of generating a description, and , L is the length of the generated description, To predict the input command before the position, The generated content before the predicted position.

[0022] S4. Using the training data set obtained in step S2, based on the likelihood function, the initial model for generating the fundus image description of the neonate constructed in step S3 is trained to obtain a fundus image description generation model for the neonate; In specific implementation, the training process includes: Through the training method of efficient parameter fine-tuning, the likelihood function of the model is maximized; the likelihood function within a batch is expressed as: In the formula For the parameters B is the batch size, which can be set according to the hardware configuration of the system where the model is deployed; is the current position prediction value of the b-th sample in the corresponding batch; is the image of the bth sample in the corresponding batch ; is the input instruction before the predicted position of the bth sample in the corresponding batch, is the generated content before the predicted position of the b-th sample in the corresponding batch; In order to improve the generalization ability of the model, the gradient accumulation method is used to average the gradients obtained in K batches; K is the set batch value; The training method for efficient parameter fine-tuning specifically includes the following contents: The initial model generated by describing the fundus image of the newborn constructed in step S3 is used as the primary model, and the corresponding primary model parameters are expressed as ; During the training process, by training a parameter relative to the primary model The parameter set is less than the third set value , to encode the increments of the primary model while keeping the original weights unchanged; add a trainable low-rank matrix factorization to the linear layer of the model; during training, the input data passes through the linear layer, and the gradient update only acts on the low-rank matrix, while keeping the parameters of the initial model unchanged; in the inference phase, the incremental weights obtained from training are merged into the initial model to achieve more efficient parameter updates; The update formula of the model weight is expressed as: In the formula are the updated model parameters; The trainable parameters Encoded model weight increments; When updating the model weights, the weights of the feature mapping network and the description generation module are updated: for the linear layer, the weight matrix is ​​represented as W, the bias is represented as b, and the increment of the weight matrix is ​​represented as , assuming that the input of the linear layer is z, then after efficient parameter fine-tuning, the forward propagation of the linear layer is expressed as: Where h is the output of the linear layer; is a scaling factor used to adjust the value of the weight increment. This efficient parameter fine-tuning mechanism can reduce the computational and storage costs while ensuring efficient and adjustable model parameters, making fine-tuning of large-scale models more efficient. This adjustment method can more accurately optimize the model's mapping and description generation capabilities for neonatal fundus image features without significantly increasing the amount of calculation, and is closely integrated with the entire neonatal fundus image description generation method to improve the model's performance when processing neonatal fundus images. The description generation module can also ensure that when generating image descriptions, it can better adapt to the characteristics of neonatal fundus images and improve the accuracy and reliability of the descriptions.

[0023] S5. Using the neonatal fundus image description generation model obtained in step S4, automatically generate neonatal fundus image descriptions.

[0024] The effect of the generation method of the present invention is described below in conjunction with an embodiment: Experiments were conducted on a constructed dataset of neonatal fundus image text descriptions using the pytorch2.1 framework. The model generation performance was evaluated from the following aspects: (1) Description text evaluation based on word matching; (2) Image description evaluation; (3) Disease diagnosis performance evaluation based on binary classification tasks; (4) Disease diagnosis performance evaluation based on multi-classification tasks. In addition, an ablation experiment was designed in this section to demonstrate the effectiveness of increasing the number of image slices in improving text generation performance. In text evaluation, the python jieba library was used to segment the text and use words as the smallest unit of text evaluation.

[0025] The description text evaluation based on word matching is used to evaluate the mutual inclusion of phrases in the true value and the predicted value. The selected indicators include Accuracy, Precision, Recall and F1 score.

[0026] In order to compare with the method of the present invention, two open source models with good performance in the medical field and fundus image field, LLaVA-Med-v1.5 and MM-Retinal, were selected for comparison. LLaVA-Med-v1.5 is a multimodal large language model obtained by tuning the instructions on the biomedical dataset using the LLaVA-v1.5 model; MM-Retinal is an image-text model pre-trained on a large number of ophthalmic images.

[0027] The evaluation results of description text based on phrase matching are shown in Table 1: Table 1 Schematic diagram of description text evaluation results based on phrase matching

[0028] It can be seen from Table 1 that the various indicators of the method of the present invention are far higher than those of the existing methods, which shows the effectiveness of the method of the present invention.

[0029] Then, we select the evaluation indicators widely used in image captioning tasks to evaluate the effect of the proposed method; the evaluation indicators include BLEU (Bilingual Evaluation Score), ROUGE (Recall Summary Evaluation Alternative Method), METEOR (Translation Evaluation Index Based on Display Ranking) and CIDER (Consistent Image Description Evaluation Index); the comparison results are shown in Table 2: Table 2 Schematic diagram of comparison results based on image description evaluation

[0030] It can be seen from Table 2 that compared with the existing solutions, the method of the present invention has obvious improvements in various indicators.

[0031] In terms of the evaluation of neonatal fundus image detection indicators based on binary classification, the output text is divided into "normal" and "abnormal" and its statistical indicators are calculated, including accuracy, specificity, recall and F1 score. The specific results are shown in Table 3: Table 3 Schematic diagram of evaluation indicators of neonatal fundus images based on binary classification

[0032] From Table 3, we can see that this method is significantly higher than the existing methods in terms of F1 score, recall rate and accuracy, and only lags behind the existing methods in terms of specificity. This is because the specificity index refers to the ability of the test to identify normal individuals. Since the output is almost all "no obvious abnormality", this index is too high, but the specificity index of this method is also very close, indicating that this method is superior.

[0033] In terms of the evaluation of neonatal fundus image detection indicators based on multi-classification, simple word matching is used to classify and calculate its statistical indicators, including accuracy, specificity, recall rate and F1 score. The specific results are shown in Table 4: Table 4 Schematic diagram of neonatal fundus image evaluation indicators based on multi-classification

[0034] It can be seen from Table 4 that the various indicators of the method of the present invention are far superior to other indicators.

[0035] The present invention also proposes a feature block extraction scheme; for this scheme, an ablation experiment is designed to prove that the number of image blocks (corresponding to the number of blocks in S2, that is, the number of nn, for example, the present invention divides the image into , i.e. the total number of processed images is 13) on the description generation performance; measured by multi-classification indicators; the experimental results are shown in Table 5: Table 5 Schematic diagram of ablation experiment results

[0036] It can be seen from Table 5 that the method of the present invention is much better than the result obtained by using only the original image or dividing the image into 4 blocks, which shows that improving the image resolution can significantly improve the accuracy of the description generated by the model.

[0037] It can be seen from Tables 1 to 5 that the various indicators of the method of the present invention are significantly higher than those of the existing method, which once again proves the effectiveness of the present invention.

[0038] like Figure 2 The method flow diagram of the imaging method of the present invention is shown as follows: the imaging method disclosed in the present invention, including the method for generating the description of the fundus image of a newborn, comprises the following steps: S1. Obtain existing fundus images of newborns; S2. describing and preprocessing the neonatal fundus image obtained in step S1 to construct a training data set; S3. Based on the image feature extraction network, feature mapping network and large language model, an initial model for generating descriptions of neonatal fundus images was constructed; The initial model for generating descriptions of neonatal fundus images includes an image feature extraction network, a feature mapping network and a description generation module which are connected in series in sequence; the image feature extraction network is used to extract features from input image data; the feature mapping network is used to map the extracted features to corresponding image descriptions; the description generation module is used to generate a description of the input image according to the obtained image features and corresponding text prompt words; S4. Using the training data set obtained in step S2, based on the likelihood function, the initial model for generating the fundus image description of the neonate constructed in step S3 is trained to obtain a fundus image description generation model for the neonate; S5. The neonatal fundus image description generation model obtained in step S4 is used to automatically generate a description of the neonatal fundus image; S6. The neonatal fundus image description content generated in step S5 is marked and re-imaged on the neonatal fundus image to obtain a neonatal fundus image with the fundus image description content.

[0039] The imaging method provided by the present invention can be directly applied to an existing neonatal fundus image device (such as a neonatal fundus image imaging system), or directly applied to a terminal (such as a computer); in specific application, the existing scheme is used to obtain the actual neonatal fundus image, and then the obtained data is input into the corresponding machine device or terminal. At this time, the machine device or terminal can obtain the description content of the actual neonatal fundus image according to the imaging method disclosed in the present invention, and display the description content on the original image through different types of representation (such as color), and then perform secondary imaging and output; at this time, the output image is an image with fundus image description content, which can reflect the actual neonatal fundus image and the corresponding description content, thereby greatly facilitating clinical medical staff and laboratory experimenters to carry out subsequent work.

Claims

1. A method for generating description of fundus images of newborns, characterized in that The steps include: S1. Obtain existing fundus images of newborns; S2. describing and preprocessing the neonatal fundus image obtained in step S1 to construct a training data set; S3. Based on the image feature extraction network, feature mapping network and large language model, an initial model for generating descriptions of neonatal fundus images was constructed; The initial model for generating descriptions of neonatal fundus images includes an image feature extraction network, a feature mapping network and a description generation module which are connected in series in sequence; the image feature extraction network is used to extract features from input image data; the feature mapping network is used to map the extracted features to corresponding image descriptions; the description generation module is used to generate a description of the input image according to the obtained image features and corresponding text prompt words; S4. Using the training data set obtained in step S2, based on the likelihood function, the initial model for describing the neonatal fundus image constructed in step S3 is trained to obtain a neonatal fundus image description generation model; S5. Using the neonatal fundus image description generation model obtained in step S4, automatically generate neonatal fundus image descriptions.

2. The method for generating a description of a fundus image of a newborn according to claim 1, characterized in that The step S2 specifically includes the following steps: Describe the fundus images of newborn babies of various set types obtained in step S1; The described fundus image of the newborn is preprocessed; the preprocessing comprises the following steps: For each set type, the fundus images of newborns of the type with a number of images less than the first set value are repeated several times; the fundus images of newborns of the type with a number of images greater than the second set value are discarded according to the set ratio; The obtained fundus image of the newborn is adjusted to a set resolution; The neonatal fundus image with adjusted resolution is divided into nn square images; the neonatal fundus image with adjusted resolution is reduced to the same size as the divided square images; finally, an original neonatal fundus image is preprocessed into nn+1 images.

3. The method for generating a description of a fundus image of a newborn according to claim 2, characterized in that The step S3 comprises the following steps: An image feature extraction network is constructed based on the SigLip model to extract features from input image data; A feature mapping network is constructed based on a multi-layer perceptron model to map the extracted features to the corresponding image descriptions; A description generation module is constructed based on the Qwen2 model to generate a description of the input image according to the obtained image features and the corresponding text prompt words.

4. The method for generating a description of a fundus image of a newborn according to claim 3, characterized in that The image feature extraction network based on the SigLip model is constructed, and specifically comprises the following steps: The SigLip model is used to construct an image feature extraction network. The processing process of the image feature extraction network is expressed as: Where Z is the feature extraction result; X is the input image; The parameters are The processing function of the SigLip model; An original neonatal fundus image and the corresponding preprocessed n images are collectively represented as ,in is the compressed original fundus image of the newborn. is the corresponding preprocessed n images; nn+1 images , are input into the image feature extraction network to obtain the feature extraction results for ,in for The feature extraction results are for The feature extraction results.

5. The method for generating a description of a neonatal fundus image according to claim 4, characterized in that The construction of the feature mapping network based on the multi-layer perceptron model specifically includes the following steps: The input of the feature mapping network is the feature extraction result output by the image feature extraction network ; For the input feature extraction results , the feature mapping network adopts The convolutions are aligned and pooled in the spatial dimension to reduce the spatial dimension; for The jth feature extraction result in , and ; Map the pooled feature data to the target feature dimension and flatten the three-dimensional features into two-dimensional features to obtain the flattened image features. ; The flattened image features Input a multilayer perceptron with two linear layers for feature mapping, expressed as: In the formula The feature extraction result output by the first linear layer The hidden state of is the activation function, and ; is the projection matrix of the first linear layer; is the bias term of the first linear layer; Feature extraction result output by multi-layer perceptron Image features; is the projection matrix of the second linear layer; is the bias term of the second linear layer; The feature extraction results obtained by the image feature extraction network Both are input into the feature mapping network to get the corresponding output for ; for The output of the corresponding feature mapping network, for The output of the corresponding feature mapping network; Will get Splicing is performed to obtain the features of the neonatal fundus image in the text space , and is used as input to the description generation module in the form of a token sequence.

6. The method for generating a description of a fundus image of a newborn according to claim 5, characterized in that The Qwen2 model-based description generation module specifically includes the following steps: The Qwen2 model is used to construct the description generation module; The input text instruction is passed through the embedding layer to get the text embedding ; Embed text and Perform splicing and generate a description of the input image using an autoregressive approach; The autoregressive method comprises the following steps: During the generation process, the previously generated value is recursively used to generate the next value until the generation process ends; During the generation process, the current position prediction value is expressed as , all generated values ​​before the current position are expressed as , then the generation process is expressed as: in represents the probability of generating a description, and , L is the length of the generated description, To predict the input command before the position, The generated content before the predicted position.

7. The method for generating a description of a fundus image of a newborn according to claim 6, characterized in that The training described in step S4 specifically includes the following steps: Through the training method of efficient parameter fine-tuning, the likelihood function of the maximum model is obtained; the likelihood function within a batch is expressed as: In the formula For the parameters The probability of ; B is the size of the batch; is the current position prediction value of the b-th sample in the corresponding batch; is the image of the bth sample in the corresponding batch ; is the input instruction before the predicted position of the bth sample in the corresponding batch, is the generated content before the predicted position of the b-th sample in the corresponding batch; In order to improve the generalization ability of the model, the gradient accumulation method is used to average the gradients obtained in K batches; K is the set batch value; The training method for efficient parameter fine-tuning specifically includes the following contents: The initial model generated by describing the fundus image of the newborn constructed in step S3 is used as the primary model, and the corresponding primary model parameters are expressed as ; During the training process, by training a parameter relative to the primary model The parameter set is less than the third set value , to encode the increments of the primary model while keeping the original weights unchanged; add a trainable low-rank matrix factorization to the linear layer of the model; during training, the input data passes through the linear layer, and the gradient update only acts on the low-rank matrix, while keeping the parameters of the initial model unchanged; In the inference phase, the incremental weights obtained through training are merged into the initial model to achieve more efficient parameter updates; The update formula of the model weight is expressed as: In the formula are the updated model parameters; The trainable parameters Encoded model weight increments; When updating the model weights, the weights of the feature mapping network and the description generation module are updated: for the linear layer, the weight matrix is ​​represented as W, the bias is represented as b, and the increment of the weight matrix is ​​represented as , assuming that the input of the linear layer is z, then after efficient parameter fine-tuning, the forward propagation of the linear layer is expressed as: Where h is the output of the linear layer; is a scaling factor used to adjust the value of the weight increment.

8. An imaging method comprising the method for generating a description of a fundus image of a newborn according to any one of claims 1 to 7, characterized in that The following steps are also included: S6. The neonatal fundus image description content generated in step S5 is marked and re-imaged on the neonatal fundus image to obtain a neonatal fundus image with the fundus image description content.

Citation Information

Patent Citations

  • Chinese image semantic description method combined with multilayer GRU based on residual error connection Inception network

    CN108830287A

  • Eye fundus image classification method based on convolutional neural network

    CN114913592A

  • Eye fundus image matching method and system based on deep learning, and readable medium

    CN114926892A

  • Image paragraph description text generation method based on information entropy

    CN118314573A

  • Eye fundus image basic model pre-training method based on image text multi-modality

    CN118397393A