Image recognition method and apparatus for fundus image, and computer device and storage medium

By using a federated training feature extraction and evaluation recognition model, and leveraging fundus images and clinical indicators, the complexity and cost of traditional hypertension target organ damage assessment are addressed, achieving efficient and accurate assessment.

WO2026091304A1PCT designated stage Publication Date: 2026-05-07TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-01-08
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Traditional hypertension target organ damage assessment requires examination of multiple organs, which is technically demanding and costly. Existing technologies are difficult to effectively train data models in multi-center medical collaborations.

Method used

A federated training model based on an initial mask autoencoder using historical fundus images was employed. By combining an embedding layer, encoder, and decoder, and through a feature extraction model and an evaluation recognition model, the damage to target organs in hypertension was assessed using fundus images and clinical indicators.

Benefits of technology

It simplifies the assessment process for target organ damage in hypertension, reduces assessment costs, improves the accuracy and reliability of assessments, breaks down medical data silos, and protects privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071221_07052026_PF_FP_ABST
    Figure CN2025071221_07052026_PF_FP_ABST
Patent Text Reader

Abstract

An image recognition method and apparatus for a fundus image, and a computer device and a storage medium. The method comprises: inputting into a feature extraction model a fundus image of a subject to be examined, so as to obtain a feature image corresponding to the fundus image, wherein the feature extraction model is obtained by means of performing federated training on an initial masked autoencoder model on the basis of historical fundus images; and inputting the feature image and clinical indicator information of said subject into an evaluation and recognition model, so as to obtain an evaluation result, wherein the evaluation result comprises the condition of hypertensive target organ damage, and / or the probability of occurrence of hypertensive target organ damage.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition methods and apparatus for fundus images, computer equipment and storage media

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on October 28, 2024, application number 202411514707.4, entitled "Image Recognition Method, Apparatus, Computer Equipment and Storage Medium for Fundus Images", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, and in particular to an image recognition method, apparatus, computer device, and storage medium for fundus images. Background Technology

[0004] Hypertensive target organ damage refers to structural or functional changes in the heart, arteries, kidneys, retina, and brain caused by elevated blood pressure. It serves as a marker of preclinical or asymptomatic cardiovascular disease and has a strong predictive value for subsequent cardiovascular and cerebrovascular events. Traditional assessment of hypertensive target organ damage typically involves multiple organs and requires technical support from various laboratory and imaging examinations, placing high demands on the examiner's skill level and related medical resources.

[0005] Artificial intelligence-assisted diagnosis and risk warning technology for hypertension and target organ damage based on large-scale fundus image models can identify fundus images and assess target organ damage in hypertension. This is of great significance for guiding the diagnosis and treatment of hypertension patients and reducing the occurrence of complications.

[0006] Therefore, simplifying the assessment of target organ damage in hypertension has become a pressing technical problem in this field. Summary of the Invention

[0007] Therefore, it is necessary to provide an image recognition method, apparatus, computer device, and storage medium for fundus images that can simplify the assessment of target organ damage in hypertension, addressing the aforementioned technical problems.

[0008] Firstly, this application provides an image recognition method for fundus images. The method includes:

[0009] The fundus image of the object to be detected is input into the feature extraction model to obtain the feature image corresponding to the fundus image; the feature extraction model is obtained by federated training of the initial mask autoencoder model based on historical fundus images;

[0010] The feature image and the clinical indicator information of the subject to be detected are input into the evaluation and recognition model to obtain the evaluation results; the evaluation results include: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0011] In one embodiment, the feature extraction model includes an embedding layer, an encoder, and a decoder. The step of inputting a fundus image of the object to be detected into the feature extraction model to obtain a feature image corresponding to the fundus image includes: inputting the fundus image of the object to be detected into the embedding layer for preliminary feature extraction to obtain a first feature vector; inputting the first feature vector into the encoder for feature abstraction to obtain a second feature vector of a preset length; the preset length being greater than the length of the first feature vector; and inputting the second feature vector into the decoder to obtain the feature image corresponding to the fundus image.

[0012] In one embodiment, before inputting the fundus image of the object to be detected into the feature extraction model, the method further includes: masking a first preset number of images in the historical fundus images to obtain a mask image; labeling a second preset number of images in the historical fundus images to obtain labeled images; and performing federated training on the initial mask autoencoder model based on the mask image, the labeled images, and the historical fundus images to obtain the feature extraction model.

[0013] In one embodiment, the initial mask autoencoder model is federatedly trained based on the mask image, the labeled image, and the historical fundus image to obtain the feature extraction model. This includes: training the initial mask autoencoder model based on the mask image and the historical fundus image to obtain a first intermediate mask autoencoder model; training the first intermediate mask autoencoder model based on the labeled image and the historical fundus image to obtain a second intermediate mask autoencoder model; determining the mask autoencoder model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training; the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of other participants in the federated training.

[0014] In one embodiment, determining the mask autoencoder model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training includes: obtaining shared model parameters from the second model parameters; the shared model parameters being model parameters of a preset proportion in the second model parameters; aggregating the shared model parameters and the first model parameters to obtain target model parameters; and determining the mask autoencoder model based on the target model parameters.

[0015] In one embodiment, the assessment and identification model includes a diagnostic model and / or a risk prediction model. The input of the feature image and the clinical indicator information of the subject to be tested into the assessment and identification model to obtain an assessment result includes: inputting the feature image and the clinical indicator information of the subject to be tested into the diagnostic model to obtain a diagnostic result; the diagnostic result is used to characterize the hypertension target organ damage status of the subject to be tested; and inputting the feature image and the clinical indicator information of the subject to be tested into the risk prediction model to obtain a risk prediction result; the risk prediction result is used to characterize the probability of hypertension target organ damage in the subject to be tested.

[0016] In one embodiment, the process of inputting the feature image and the clinical indicator information of the subject to be tested into a risk prediction model to obtain a risk prediction result includes: inputting the feature image and the clinical indicator information of the subject to be tested into a risk prediction model to obtain a risk prediction score; determining the risk prediction result based on the risk prediction score and a survival curve; wherein the survival curve is determined based on historical risk prediction scores.

[0017] In one embodiment, inputting the feature image and the clinical indicator information of the object to be detected into the risk prediction model to obtain the risk prediction score includes: concatenating the feature image and the clinical indicator information of the object to be detected in the form of a vector to obtain a model input vector; and inputting the model input vector into the risk function of the risk prediction model to obtain the risk prediction score.

[0018] In one embodiment, determining the risk prediction result based on the risk prediction score and the survival curve includes: using the model input vector of the study population as a training set, determining the historical risk prediction score based on each model input vector in the training set; determining the median of multiple historical risk prediction scores as a division threshold, dividing individuals with negative hypertension into low-risk and high-risk groups according to the division threshold; plotting survival curves for the two groups using Kaplan-Meier survival analysis, and then determining the risk prediction result based on the risk prediction score and the survival curve.

[0019] In one embodiment, the clinical indicators of the subject to be tested may include, for example, the subject's age, gender, body mass index (BMI), smoking status, diabetes, triglycerides (TG), low-density lipoprotein cholesterol (LDL-C), and high-density lipoprotein cholesterol (HDL-C).

[0020] In one embodiment, masking a first preset number of images in the historical fundus images to obtain a masked image includes: randomly selecting the first preset number of images from the historical fundus images, and masking a portion of the image blocks of the first preset number of images to obtain a masked image, wherein the first preset number is less than or equal to the number of historical fundus images.

[0021] In one embodiment, annotating a second preset number of images in the historical fundus images to obtain the annotated images includes: randomly selecting the second preset number of images from the historical fundus images, and annotating the hypertensive target organ damage areas in the second preset number of images to obtain the annotated images; wherein the second preset number may be less than or equal to the number of historical fundus images.

[0022] In one embodiment, aggregating the shared model parameters and the first model parameters to obtain target model parameters includes: performing a weighted summation of the shared model parameters and the first model parameters to obtain the target model parameters.

[0023] Secondly, this application also provides an image recognition device for fundus images. The device includes: a first determining module and a second determining module.

[0024] The first determining module is used to input the fundus image of the object to be detected into the feature extraction model to obtain the feature image corresponding to the fundus image; the feature extraction model is obtained by federated training of the initial mask autoencoder model based on historical fundus images.

[0025] The second determining module is used to input the feature image and the clinical indicator information of the subject to be detected into the evaluation and recognition model to obtain the evaluation result; the evaluation result includes: the damage status of the target organs of hypertension, and / or the probability of damage to the target organs of hypertension.

[0026] Thirdly, this application also provides a computer device, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the above methods.

[0027] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above methods.

[0028] The aforementioned image recognition method, apparatus, computer equipment, and storage medium for fundus images obtain a feature extraction model by inputting the fundus image of the subject to be detected into a feature extraction model obtained through federated training of an initial mask autoencoder model based on historical fundus images. This yields a feature image corresponding to the fundus image. Then, the feature image and the clinical indicator information of the subject to be detected are input into an evaluation and recognition model to obtain the extent of hypertension target organ damage and / or the probability of damage to hypertension target organs. In this embodiment, the diagnosis and prediction of hypertension for the subject to be detected can be determined based solely on the fundus image and clinical indicator information, eliminating the need to examine multiple organs involved in hypertension target organ damage. This simplifies the operational process of hypertension target organ damage assessment, reduces assessment costs, and improves the accuracy of hypertension target organ damage assessment by using the feature extraction model and the evaluation and recognition model. Attached Figure Description

[0029] Figure 1 is an internal structural diagram of a computer device provided in an embodiment of this application;

[0030] Figure 2 is a flowchart illustrating an image recognition method for fundus images provided in an embodiment of this application;

[0031] Figure 3 is a flowchart of a hypertension screening and prediction method provided in an embodiment of this application;

[0032] Figure 4 is a flowchart illustrating a feature image acquisition method provided in an embodiment of this application;

[0033] Figure 5 is a flowchart illustrating a feature extraction model acquisition method provided in an embodiment of this application;

[0034] Figure 6 is a flowchart illustrating a feature extraction model determination method provided in an embodiment of this application;

[0035] Figure 7 is a flowchart illustrating another feature extraction model determination method provided in an embodiment of this application;

[0036] Figure 8 is a schematic diagram of a federated training process provided in an embodiment of this application;

[0037] Figure 9 is a flowchart illustrating an evaluation result acquisition method provided in an embodiment of this application;

[0038] Figure 10 is a flowchart illustrating a method for obtaining risk prediction results provided in an embodiment of this application;

[0039] Figure 11 is a flowchart illustrating a diagnostic and risk warning method for hypertension and target organ damage based on fundus images provided in an embodiment of this application.

[0040] Figure 12 is a structural block diagram of an image recognition device for fundus images provided in an embodiment of this application. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0042] Hypertensive target organ damage refers to structural or functional changes in the heart, arteries, kidneys, retina, and brain caused by elevated blood pressure. It serves as a marker of preclinical or asymptomatic cardiovascular disease and has a strong predictive value for subsequent cardiovascular and cerebrovascular events. Traditional assessment of hypertensive target organ damage typically involves multiple organs and requires technical support from various laboratory and imaging examinations, placing high demands on the examiner's skill level and related medical resources.

[0043] For computer vision models, models trained on a single center are often prone to overfitting, and their performance may significantly degrade across other centers, sample distributions, and tasks. Furthermore, a larger training sample size often leads to better model performance. In real-world multi-center medical collaborations, issues such as data silos between institutions and a lack of effective solutions to privacy concerns can hinder model training on large datasets.

[0044] Artificial intelligence-assisted diagnosis and risk warning technology for hypertension and target organ damage based on large-scale fundus image models, which performs image recognition on fundus images to assess target organ damage in hypertension, is of great significance for guiding the diagnosis and treatment of hypertension patients and reducing the occurrence of complications.

[0045] Therefore, simplifying the assessment of target organ damage in hypertension has become a pressing technical problem in this field.

[0046] The image recognition method for fundus images provided in this application embodiment can be applied to the application environment shown in Figure 1. Figure 1 is an internal structural diagram of a computer device provided in this application embodiment. The computer device can be a server, and its internal structural diagram is as shown in Figure 1. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an image recognition method for fundus images.

[0047] Those skilled in the art will understand that the structure shown in Figure 1 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.

[0048] In one embodiment, as shown in FIG2, FIG2 is a schematic flowchart of an image recognition method for fundus images provided in an embodiment of the present application. The method can be applied to the computer device in FIG1, and the method includes the following steps S201 and S202.

[0049] S201, input the fundus image of the object to be detected into the feature extraction model to obtain the feature image corresponding to the fundus image.

[0050] In one embodiment, the feature extraction model is obtained by federated training of an initial masked auto-encoder (MAE) model based on historical fundus images. The feature extraction model consists of an embedding layer, an encoder, and a decoder.

[0051] For example, the original fundus image of the subject (3×512×512) can be converted into a 3-color-channel, 224×224 resolution fundus image using bicubic interpolation. Then, this 3-color-channel, 224×224 resolution fundus image is input into the embedding layer and convolved with a 16×16 kernel and a stride of 16 to obtain a length of... The first feature vector is then input into the encoder for feature abstraction, extracting a second feature vector of length 1024. This second feature vector is then input into the decoder to obtain a feature image with 3 color channels and a resolution of 224×224.

[0052] Optionally, during federated training of the initial Masked Auto-Encoders (MAE) model, each participant in the federated training can mask a first predetermined number of historical fundus images to obtain mask images. The initial MAE model is then trained based on these mask images to obtain a first intermediate MAE model, ensuring that the images generated by the model based on the mask images are as close as possible to the original images. Furthermore, each participant in the federated training can also annotate a second predetermined number of historical fundus images to obtain labeled images. The first intermediate MAE model is then semi-supervised trained based on these labeled images to obtain a second intermediate MAE model. Finally, the target model parameters can be obtained by combining the first model parameters of the second intermediate MAE model with the second model parameters of the other participants in the federated training. These target model parameters are then used as the model parameters for the feature extraction model.

[0053] S202, input the feature image and the clinical indicator information of the subject to be detected into the evaluation and recognition model to obtain the evaluation result.

[0054] The assessment results include: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0055] For example, the clinical indicators of the subject to be tested may include the subject's age, gender, body mass index (BMI), smoking status, diabetes, triglycerides (TG), low-density lipoprotein cholesterol (LDL-C), and high-density lipoprotein cholesterol (HDL-C).

[0056] In one embodiment, the assessment and identification model may include a diagnostic model and / or a risk prediction model. The feature image and the clinical indicators of the subject to be tested can be input into the diagnostic model to obtain a diagnostic result, which characterizes the target organ damage status of the subject's hypertension. Alternatively, the feature image and the clinical indicators of the subject to be tested can be input into the risk prediction model to obtain a risk prediction result, which characterizes the probability of target organ damage in the subject's hypertension.

[0057] Optionally, the above eight clinical indicators can be represented by an indicator vector a = (a 1 ,a 2,…,a 8 ) T Indicated. Where, a i This represents the scalar numerical value of each indicator. In other words, the indicator vector and the feature image vector can be used together as input to evaluate the recognition model to obtain the evaluation result.

[0058] For example, referring to Figure 3, which is a flowchart of a hypertension screening and prediction method provided by an embodiment of this application, opportunistic screening of the community population can be conducted through blood pressure measurement. If the screening result is positive, the resident is identified as a subject to be tested (i.e., a hypertensive patient). Then, the characteristic images and clinical indicators of the fundus images of the subject to be tested are used as inputs to a diagnostic model for cross-sectional auxiliary diagnosis of target organ damage and cardiovascular and cerebrovascular diseases caused by hypertension, and the characteristic images and clinical indicators of the fundus images of the subject to be tested are used as inputs to a risk prediction model for longitudinal risk prediction.

[0059] In this embodiment, the fundus image of the subject to be tested is input into a feature extraction model obtained by federated training of an initial mask autoencoder model based on historical fundus images. This yields a feature image corresponding to the fundus image. The feature image and the clinical indicator information of the subject are then input into an evaluation and recognition model to obtain the extent of hypertension target organ damage and / or the probability of damage to hypertension target organs. Because this embodiment only requires the fundus image and clinical indicator information of the subject to determine the diagnosis and prediction of hypertension, it eliminates the need to examine multiple organs involved in hypertension target organ damage. This simplifies the operational process of hypertension target organ damage assessment, reduces assessment costs, and improves the accuracy of hypertension target organ damage assessment by using the feature extraction model and the evaluation and recognition model. Furthermore, federated training breaks down data silos in the medical application field, expanding the amount of training data while protecting the data privacy of all parties, further improving the reliability and robustness of the feature extraction model.

[0060] In one embodiment, the feature extraction model includes an embedding layer, an encoder, and a decoder. Referring to Figure 4, which is a flowchart illustrating a feature image acquisition method according to an embodiment of this application, this embodiment relates to a possible implementation of how to input a fundus image of the object to be detected into a feature extraction model to obtain a feature image corresponding to the fundus image. Based on the above embodiment, step S201 includes steps S401 to S403.

[0061] S401, the fundus image of the object to be detected is input into the embedding layer for preliminary feature extraction to obtain the first feature vector.

[0062] For example, the original fundus image of the subject (3×512×512) can be converted into a 3-color-channel, 224×224 resolution fundus image using bicubic interpolation. Then, this 3-color-channel, 224×224 resolution fundus image is input into the embedding layer and convolved with a 16×16 kernel and a stride of 16 to obtain a length of... The first eigenvector.

[0063] S402, the first feature vector is input to the encoder for feature abstraction to obtain a second feature vector of a preset length.

[0064] The preset length is greater than the length of the first feature vector.

[0065] In this embodiment, a preset length can be set according to the length of the first feature vector, such that the preset length is greater than the length of the first feature vector, so as to perform feature abstraction on the first feature vector and reduce the complexity of the data.

[0066] For example, the first feature vector can be input into the encoder for feature abstraction to extract a second feature vector of length 1024.

[0067] S403, input the second feature vector into the decoder to obtain the feature image corresponding to the fundus image.

[0068] For example, the second feature vector can be input into the decoder to obtain a feature image with 3 color channels and a resolution of 224×224. This feature image with 3 color channels and a resolution of 224×224 is the feature image corresponding to the fundus image.

[0069] In this embodiment, the fundus image of the subject to be detected is input into the embedding layer for preliminary feature extraction to obtain a first feature vector. The first feature vector is then input into the encoder for feature abstraction to obtain a second feature vector of a preset dimension. The second feature vector is then input into the decoder to obtain a feature image corresponding to the fundus image. This enables the extraction of features from the fundus image of the subject to be detected through the feature extraction model, facilitating subsequent diagnosis and risk prediction of hypertension target organ damage based on the feature image, thereby improving the accuracy and reliability of diagnosis and risk prediction of hypertension target organ damage.

[0070] Referring to Figure 5, which is a flowchart illustrating a feature extraction model acquisition method according to an embodiment of this application, based on the above embodiments, an image recognition method for fundus images according to an embodiment of this application further includes the following steps: S501 to S503 before inputting the fundus image of the object to be detected into the feature extraction model.

[0071] S501, a first preset number of images in the historical fundus images are masked to obtain a mask image.

[0072] Optionally, historical fundus images of hypertensive patients and healthy individuals can be obtained from the internet or local databases.

[0073] In one embodiment, a first preset number of images can be randomly extracted from historical fundus images, and then a portion of the image blocks of the first preset number of images can be masked to obtain a masked image.

[0074] For example, a first preset number of images can be randomly extracted from historical fundus images, and then 75% of the image blocks of each of the first preset number of images can be randomly masked to obtain a masked image.

[0075] It should be noted that the first preset number can be less than or equal to the number of historical fundus images.

[0076] S502, label a second preset number of images in the historical fundus images to obtain labeled images.

[0077] In one embodiment, a second preset number of images can be randomly selected from historical fundus images, and then the hypertensive target organ damage areas in the second preset number of images can be labeled to obtain labeled images.

[0078] It should be noted that the second preset number can be less than or equal to the number of historical fundus images.

[0079] S503 uses masked images, labeled images, and historical fundus images to perform federated training on the initial masked autoencoder model to obtain a feature extraction model.

[0080] For example, assuming that 75% of the image patches of each image of a first preset number are randomly masked to obtain mask images, an initial mask autoencoder model can be trained based on the mask images and historical fundus images to obtain a first intermediate mask autoencoder model. Pixel loss is used as the training loss function to ensure that the image generated by the model based on the remaining 25% of the unmasked image patches is as close as possible to the original historical fundus image.

[0081] In one embodiment, a first intermediate mask autoencoder model can be trained based on labeled images and historical fundus images to obtain a second intermediate mask autoencoder model. This can be achieved by first fine-tuning a ResNet-50 model pre-trained on ImageNet using labeled images as a teacher model for semi-supervised knowledge distillation. TThen, the historical fundus images are input into the pre-trained ResNet-50 model for image recognition, and the image recognition results output by the ResNet-50 model are used as the historical fundus images x. j pseudo-tags Right now:

[0082] in, Let be the indicator function, and sigmoid(·) be the sigmoid function. Then, the pseudo-labels are... As a signal, a supervised learning classification task is set up to train the first intermediate mask autoencoder model, thereby obtaining the second intermediate mask autoencoder model.

[0083] Alternatively, cross entropy can be used as the loss function for the classification task. In the training process described above, based on the masked image, labeled image, and historical fundus images, the loss function Loss can be expressed as: Loss = L pixel +0.3·L CE

[0084] Among them, L pixel L CE These are the pixel loss function and the cross-entropy function, respectively.

[0085] In this embodiment of the application, there are multiple participants in the federated training. The process of training the initial mask autoencoder model based on the mask image, the labeled image, and the historical fundus image to obtain the first intermediate mask autoencoder model and the second intermediate mask autoencoder model needs to be performed separately by each participant in the federated training.

[0086] For example, assuming there are N participants in the federated training, each participant will train a second intermediate mask autoencoder model based on its local data. In other words, a total of N second intermediate mask autoencoder models with different model parameters will be obtained.

[0087] In one implementation, a participant in the federated training can aggregate the first model parameters of the local second intermediate mask autoencoder model with the second model parameters of other participants in the federated training to obtain the target model parameters, and determine the target model parameters as the model parameters of the feature extraction model.

[0088] In this embodiment, a first preset number of historical fundus images are masked to obtain mask images, and a second preset number of historical fundus images are labeled to obtain labeled images. Based on the mask images, labeled images, and historical fundus images, an initial masked autoencoder model is federatedly trained to obtain a feature extraction model. This enables the model to learn the characteristics of hypertension target organ damage based on historical fundus images, improving the accuracy of the feature extraction model. Furthermore, training based on the mask images and labeled images further improves the accuracy and robustness of the feature extraction model, facilitating subsequent diagnosis and risk prediction of hypertension target organ damage based on feature images, thus enhancing the accuracy and reliability of hypertension target organ damage diagnosis and risk prediction.

[0089] Referring to Figure 6, which is a flowchart illustrating a feature extraction model determination method provided in an embodiment of this application, this embodiment relates to a possible implementation of how to perform federated training on an initial mask autoencoder model based on a mask image, an annotated image, and historical fundus images to obtain a feature extraction model. Based on the above embodiment, step S503 includes the following steps S601 to S603.

[0090] S601, the initial mask autoencoder model is trained based on the mask image and historical fundus images to obtain the first intermediate mask autoencoder model.

[0091] For example, assuming that 75% of the image patches of each image of a first preset number are randomly masked to obtain mask images, an initial mask autoencoder model can be trained based on the mask images and historical fundus images to obtain a first intermediate mask autoencoder model. Pixel loss is used as the training loss function to ensure that the image generated by the model based on the remaining 25% of the unmasked image patches is as close as possible to the original historical fundus image.

[0092] S602, the first intermediate mask autoencoder model is trained based on the labeled image and historical fundus image to obtain the second intermediate mask autoencoder model.

[0093] In one embodiment, a first intermediate mask autoencoder model can be trained based on labeled images and historical fundus images to obtain a second intermediate mask autoencoder model. This can be achieved by first fine-tuning a ResNet-50 model pre-trained on ImageNet using labeled images as a teacher model for semi-supervised knowledge distillation. T Then, historical fundus images are input into the pre-trained ResNet-50 model for image recognition, and the image recognition results output by the ResNet-50 model are used as the historical fundus image x. j pseudo-tags Right now:

[0094] in, Let be the indicator function, and sigmoid(·) be the sigmoid function. Then, the pseudo-labels are... As a signal, a supervised learning classification task is set up to train the first intermediate mask autoencoder model and obtain the second intermediate mask autoencoder model.

[0095] S603 determines the feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of the other participants in the federated training.

[0096] The second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of other participants in the federated training.

[0097] In this embodiment of the application, there are multiple participants in the federated training. The process of training the initial mask autoencoder model based on the mask image, the labeled image, and the historical fundus image to obtain the first intermediate mask autoencoder model and the second intermediate mask autoencoder model needs to be performed separately by each participant in the federated training.

[0098] For example, assuming there are N participants in the federated training, each participant will train a second intermediate mask autoencoder model based on its local data. In other words, a total of N second intermediate mask autoencoder models with different model parameters will be obtained.

[0099] In one implementation, a participant in the federated training can aggregate the first model parameters of the local second intermediate mask autoencoder model with the second model parameters of other participants in the federated training to obtain the target model parameters, and determine the target model parameters as the model parameters of the feature extraction model.

[0100] In this embodiment, an initial masked autoencoder model is trained based on a masked image and historical fundus images to obtain a first intermediate masked autoencoder model. The first intermediate masked autoencoder model is then trained based on labeled images and historical fundus images to obtain a second intermediate masked autoencoder model. Based on the first model parameters of the second intermediate masked autoencoder model and the second model parameters of other participants in the federated training, a feature extraction model is determined. This enables the model to learn the characteristics of hypertension target organ damage based on historical fundus images, improving the accuracy of the feature extraction model. Furthermore, training based on masked images and labeled images further enhances the accuracy and robustness of the feature extraction model, facilitating subsequent diagnosis and risk prediction of hypertension target organ damage based on feature images, thus improving the accuracy and reliability of hypertension target organ damage diagnosis and risk prediction.

[0101] Referring to Figure 7, which is a flowchart illustrating another feature extraction model determination method provided in an embodiment of this application, this embodiment relates to a possible implementation of determining the feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training. Based on the above embodiment, step S603 includes the following steps: S710 to S703.

[0102] S701, obtain shared model parameters from the second model parameters.

[0103] Among them, the shared model parameters are model parameters with a preset proportion in the second model parameters.

[0104] In one embodiment, a preset ratio can be set as needed, and then the preset ratio of shared model parameters can be obtained from the devices of other participants in the federated training.

[0105] For example, a preset ratio of 15% can be set, in which case 15% of the model parameters can be obtained from each of the other participating devices in the federated training, and the obtained model parameters can be collectively determined as shared model parameters.

[0106] S702, aggregate the shared model parameters and the first model parameters to obtain the target model parameters.

[0107] In one possible implementation, the shared model parameters and the first model parameters can be weighted and summed to obtain the target model parameters.

[0108] Alternatively, the shared model parameters and the first model parameters can be weighted and summed to obtain a sum, which can then be used as the model parameters for the next round of training.

[0109] For example, referring to Figure 8, which is a schematic diagram of a federated training process provided in an embodiment of this application, in the (t+1)th training round of the model, each participant k in the federated training optimizes and updates the model parameters locally in parallel for one round, that is:

[0110] For each client k in parallel do

[0111] in, These are some of the model parameters obtained from aggregation in round t. Let n be the model parameters for the k-th participant in round t; k The model performance of the k-th participant is multiplied by the amount of data used to evaluate the model performance, n = ∑ k n k S is a filter for randomly selecting parameters that need to be shared and aggregated. It is implemented as follows: all parties share the same random number seed from the random library. In the (t+1)th round, the layer number of the shared parameter is the (t+1)th group of random non-negative integers generated by the random library with the seed. σ is the proportion of the number of parameters shared in each round to the total number of parameters, which is set to 15% here.

[0112] S703, determine the feature extraction model based on the target model parameters.

[0113] In one embodiment, the target model parameters can be determined as the model parameters of the feature extraction model, and then the model parameters of the initial mask autoencoder model can be updated based on the target model parameters to obtain the feature extraction model.

[0114] In this embodiment, shared model parameters are obtained from the second model parameters, and the shared model parameters and the first model parameters are aggregated to obtain the target model parameters. The feature extraction model is determined based on the target model parameters, thereby enabling the feature extraction model to be jointly obtained based on historical fundus data from different participating devices. This breaks down data silos in the medical application field, expands the amount of training data while protecting the data privacy of all parties, and improves the reliability and robustness of the feature extraction model. This facilitates subsequent diagnosis and risk prediction of hypertension target organ damage based on feature images, and improves the accuracy and reliability of hypertension target organ damage diagnosis and risk prediction.

[0115] The assessment and identification model includes a diagnostic model and / or a risk prediction model. Referring to Figure 9, which is a flowchart illustrating an assessment result acquisition method according to an embodiment of this application, this embodiment relates to a possible implementation of how to input feature images and clinical indicator information of the subject to be detected into the assessment and identification model to obtain assessment results. Based on the above embodiment, step S202 includes the following steps S901 and S902.

[0116] S901 inputs the feature image and the clinical indicator information of the object to be detected into the diagnostic model to obtain the diagnostic result.

[0117] The diagnostic results are used to characterize the target organ damage caused by hypertension in the subjects being tested.

[0118] Alternatively, the diagnostic model can be a multilayer perceptron neural network model.

[0119] In this embodiment of the application, clinical indicator information can be represented by an indicator vector a = (a 1 ,a 2 ,…,a 8 ) T The process involves concatenating the feature image and clinical indicator information of the same target object as a vector to obtain the model input vector Cat. This input vector Cat is then fed into the diagnostic model to obtain the diagnostic result.

[0120] Optionally, cross-entropy can be used as the loss function in the training process of the diagnostic model, and stochastic gradient descent (SGD) can be used as the optimizer for model training, thereby realizing a hypertension diagnostic model based on clinical indicator information and fundus feature images. At the same time, the hypertension diagnostic model can also fit the PWV value of the object through regression analysis.

[0121] S902, input the feature image and the clinical indicator information of the subject to be detected into the risk prediction model to obtain the risk prediction result.

[0122] Among them, the risk prediction results are used to characterize the probability of damage to the target organs of the subject with hypertension.

[0123] Alternatively, the diagnostic model can be a neural network model.

[0124] In this embodiment of the application, clinical indicator information can be represented by an indicator vector a = (a 1 ,a 2 ,…,a 8 ) TThe process involves concatenating the feature image and clinical indicator information of the same target object into a vector to obtain the model input vector. This model input vector is then fed into the risk prediction model to obtain the risk prediction result.

[0125] For example, the clinical indicators of the subjects to be tested may include age, gender, BMI, smoking status, diabetes, systolic blood pressure, diastolic blood pressure, TG, and LDL-C and HDL-C. These, together with the feature images extracted from the fundus images, constitute 9+1024 covariates, i.e. To describe a risk function:

[0126] Where, p T =(p 1 ,p 2 ,…,p 9+1024 ) T H is the regression coefficient of the independent variable, and H0(t) is the baseline risk, which is mathematically defined as H(0,t), that is, the risk function when the input is fixed at x=0.

[0127] In one embodiment, individuals with hypertension-negative status can be divided into low-risk and high-risk groups based on historical risk prediction scores from the training set. The dividing threshold could be, for example, the median of the historical risk prediction scores. Then, Kaplan-Meier survival analysis is used to plot survival curves for both groups. Finally, based on the risk prediction scores and the survival curves, the risk prediction outcome is determined.

[0128] In this embodiment, the feature image and the clinical indicator information of the subject to be tested are input into the diagnostic model to obtain the diagnostic result, and the feature image and the clinical indicator information of the subject to be tested are input into the risk prediction model to obtain the risk prediction result. Thus, the diagnosis and prediction result of hypertension of the subject to be tested can be determined based only on the fundus image and clinical indicator information of the subject to be tested, without the need to examine multiple organs involved in hypertension target organ damage. This simplifies the operation process of hypertension target organ damage assessment, reduces the assessment cost, and improves the accuracy of hypertension target organ damage assessment by using the feature extraction model and the assessment and recognition model.

[0129] Referring to Figure 10, which is a flowchart illustrating a method for obtaining risk prediction results according to an embodiment of this application, this embodiment relates to a possible implementation of how to input feature images and clinical indicator information of the subject to be detected into a risk prediction model to obtain risk prediction results. Based on the above embodiment, step S902 includes the following steps S1001 and S1002.

[0130] S1001, input the feature image and the clinical indicator information of the subject to be detected into the risk prediction model to obtain the risk prediction score.

[0131] In one embodiment, the feature image and the clinical indicator information of the object to be detected can be concatenated in the form of a vector to obtain the model input vector. Then, the model input vector is input into the risk function of the risk prediction model to obtain the risk prediction score.

[0132] S1002, based on the risk prediction score and survival curve, determines the risk prediction result.

[0133] The survival curve is determined based on historical risk prediction scores.

[0134] In one embodiment, the model input vectors of the study population can be used as the training set. Historical risk prediction scores are then determined based on each model input vector in the training set, and the median of multiple historical risk prediction scores is used as the splitting threshold. Individuals with negative hypertension are divided into low-risk and high-risk groups according to the splitting threshold. Kaplan-Meier survival analysis is used to plot survival curves for the two groups. Finally, based on the risk prediction scores and survival curves, the risk prediction result is determined. The survival function is calculated as follows:

[0135] Alternatively, the log-rank test can be used to compare the difference in disease risk between the low-risk and high-risk groups, and the time-dependent ROC curve can be used to quantify the model's predictive performance at different future times.

[0136] In this embodiment, the feature image and the clinical indicator information of the subject to be tested are input into the risk prediction model to obtain a risk prediction score. Based on the risk prediction score and the survival curve, the risk prediction result is determined, thereby quantifying the individual's risk of developing the disease in a specific future time period, improving the accuracy of the risk prediction result, and facilitating the obtaining of more accurate prevention advice based on the risk prediction result, in order to reduce the incidence, disability rate and mortality rate of hypertension.

[0137] Referring to Figure 11, which is a flowchart illustrating a method for diagnosing and providing risk warning of hypertension and target organ damage based on fundus images according to an embodiment of this application, the method includes the following steps S1101 to S1107.

[0138] S1101, a first preset number of images in the historical fundus images are masked to obtain masked images, and a second preset number of images in the historical fundus images are labeled to obtain labeled images.

[0139] S1102, the initial mask autoencoder model is trained based on the mask image and historical fundus images to obtain the first intermediate mask autoencoder model.

[0140] S1103, the first intermediate mask autoencoder model is trained based on the labeled image and historical fundus image to obtain the second intermediate mask autoencoder model.

[0141] S1104: Obtain shared model parameters from the second model parameters, aggregate the shared model parameters and the first model parameters to obtain the target model parameters, and determine the feature extraction model based on the target model parameters.

[0142] S1105, input the fundus image of the object to be detected into the feature extraction model to obtain the feature image corresponding to the fundus image.

[0143] S1106, input the feature image and the clinical indicator information of the object to be detected into the diagnostic model to obtain the diagnostic result.

[0144] S1107, input the feature image and the clinical indicator information of the subject to be detected into the risk prediction model to obtain the risk prediction result.

[0145] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0146] Based on the same inventive concept, this application also provides an image recognition device for fundus images to implement the image recognition method for fundus images described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more fundus image recognition device embodiments provided below can be found in the limitations of the image recognition method for fundus images described above, and will not be repeated here.

[0147] In one embodiment, as shown in FIG12, FIG12 is a structural block diagram of an image recognition device for fundus images provided in an embodiment of the present application. The device 1200 includes: a first determining module 1201 and a second determining module 1202.

[0148] The first determining module 1201 is used to input the fundus image of the object to be detected into the feature extraction model to obtain the feature image corresponding to the fundus image. In one embodiment, the feature extraction model is obtained by federated training of an initial mask autoencoder model based on historical fundus images.

[0149] The second determining module 1202 is used to input the feature image and the clinical indicator information of the subject to be detected into the evaluation and recognition model to obtain the evaluation result. In one embodiment, the evaluation result includes: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0150] In one embodiment, the first determining module 1201 includes: a first determining unit, a second determining unit, and a third determining unit.

[0151] The first determining unit is used to input the fundus image of the object to be detected into the embedding layer for preliminary feature extraction to obtain the first feature vector.

[0152] The second determining unit is used to input the first feature vector into the encoder for feature abstraction to obtain a second feature vector of a preset length; the preset length is greater than the length of the first feature vector.

[0153] The third determining unit is used to input the second feature vector into the decoder to obtain the feature image corresponding to the fundus image.

[0154] In one embodiment, the device 1200 further includes a third determining module, a fourth determining module, and a fifth determining module.

[0155] The third determining module is used to mask a first preset number of images in historical fundus images to obtain a masked image.

[0156] The fourth determination module is used to annotate a second preset number of images in historical fundus images to obtain annotated images.

[0157] The fifth determination module is used to perform federated training on the initial mask autoencoder model based on the mask image, labeled image, and historical fundus images to obtain the feature extraction model.

[0158] In one embodiment, the fifth determining module includes: a fourth determining unit, a fifth determining unit, and a sixth determining unit.

[0159] The fourth determining unit is used to train the initial mask autoencoder model based on the mask image and historical fundus images to obtain the first intermediate mask autoencoder model.

[0160] The fifth determining unit is used to train the first intermediate mask autoencoder model based on the labeled image and historical fundus image to obtain the second intermediate mask autoencoder model.

[0161] The sixth determining unit is used to determine the feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of the other participants in the federated training; wherein the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of the other participants in the federated training.

[0162] In one embodiment, the sixth determining unit includes: an acquisition subunit, a first determining subunit, and a second determining subunit.

[0163] Obtain sub-units to retrieve shared model parameters from the second model parameters; the shared model parameters are model parameters in the second model parameters at a preset ratio.

[0164] The first determining sub-unit is used to aggregate the shared model parameters and the first model parameters to obtain the target model parameters.

[0165] The second determining subunit is used to determine the feature extraction model based on the target model parameters.

[0166] In one embodiment, the second determining module 1202 includes a sixth determining module and a seventh determining module.

[0167] The sixth determination module is used to input the feature image and the clinical indicator information of the subject to be tested into the diagnostic model to obtain the diagnostic results; the diagnostic results are used to characterize the target organ damage of the subject to be tested due to hypertension.

[0168] The seventh determination module is used to input the feature image and the clinical indicator information of the subject to be tested into the risk prediction model to obtain the risk prediction result; the risk prediction result is used to characterize the probability of damage to the target organs of hypertension in the subject to be tested.

[0169] In one embodiment, the seventh determining module includes a seventh determining unit and an eighth determining unit.

[0170] The seventh determination unit is used to input the feature image and the clinical indicator information of the subject to be detected into the risk prediction model to obtain the risk prediction score.

[0171] The eighth determining unit is used to determine the risk prediction result based on the risk prediction score and the survival curve; the survival curve is determined based on the historical risk prediction score.

[0172] The modules in the aforementioned fundus image recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0173] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: inputting a fundus image of a subject to be detected into a feature extraction model to obtain a feature image corresponding to the fundus image; wherein the feature extraction model is obtained by federated training of an initial mask autoencoder model based on historical fundus images; inputting the feature image and clinical indicator information of the subject to be detected into an evaluation and recognition model to obtain an evaluation result; wherein the evaluation result includes: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0174] In one embodiment, when the processor executes the computer program, it further performs the following steps: inputting the fundus image of the object to be detected into the embedding layer for preliminary feature extraction to obtain a first feature vector; inputting the first feature vector into the encoder for feature abstraction to obtain a second feature vector of a preset length; the preset length is greater than the length of the first feature vector; and inputting the second feature vector into the decoder to obtain a feature image corresponding to the fundus image.

[0175] In one embodiment, when the processor executes the computer program, it further performs the following steps: masking a first preset number of images in historical fundus images to obtain mask images; labeling a second preset number of images in historical fundus images to obtain labeled images; and performing federated training on an initial mask autoencoder model based on the mask images, labeled images, and historical fundus images to obtain a feature extraction model.

[0176] In one embodiment, when the processor executes the computer program, it further performs the following steps: training an initial mask autoencoder model based on a mask image and historical fundus images to obtain a first intermediate mask autoencoder model; training the first intermediate mask autoencoder model based on annotated images and historical fundus images to obtain a second intermediate mask autoencoder model; determining a feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training; the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of other participants in the federated training.

[0177] In one embodiment, when the processor executes the computer program, it further performs the following steps: obtaining shared model parameters from the second model parameters; the shared model parameters are model parameters of a preset proportion in the second model parameters; aggregating the shared model parameters and the first model parameters to obtain target model parameters; and determining a feature extraction model based on the target model parameters.

[0178] In one embodiment, when the processor executes the computer program, it further performs the following steps: inputting the feature image and the clinical indicator information of the subject to be tested into a diagnostic model to obtain a diagnostic result; the diagnostic result is used to characterize the damage to the target organs of hypertension in the subject to be tested; inputting the feature image and the clinical indicator information of the subject to be tested into a risk prediction model to obtain a risk prediction result; the risk prediction result is used to characterize the probability of damage to the target organs of hypertension in the subject to be tested.

[0179] In one embodiment, when the processor executes the computer program, it further performs the following steps: inputting feature images and clinical indicator information of the subject to be detected into a risk prediction model to obtain a risk prediction score; determining the risk prediction result based on the risk prediction score and the survival curve; the survival curve is determined based on historical risk prediction scores.

[0180] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon. When executed by a processor, the computer program performs the following steps: inputting a fundus image of the subject to be detected into a feature extraction model to obtain a feature image corresponding to the fundus image; the feature extraction model is obtained by federated training of an initial mask autoencoder model based on historical fundus images; inputting the feature image and clinical indicator information of the subject to be detected into an evaluation and recognition model to obtain an evaluation result; the evaluation result includes: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0181] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting the fundus image of the object to be detected into the embedding layer for preliminary feature extraction to obtain a first feature vector; inputting the first feature vector into the encoder for feature abstraction to obtain a second feature vector of a preset length; the preset length is greater than the length of the first feature vector; and inputting the second feature vector into the decoder to obtain a feature image corresponding to the fundus image.

[0182] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: masking a first preset number of images in the historical fundus images to obtain mask images; labeling a second preset number of images in the historical fundus images to obtain labeled images; and performing federated training on an initial mask autoencoder model based on the mask images, labeled images, and historical fundus images to obtain a feature extraction model.

[0183] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: training an initial mask autoencoder model based on a mask image and historical fundus images to obtain a first intermediate mask autoencoder model; training the first intermediate mask autoencoder model based on annotated images and historical fundus images to obtain a second intermediate mask autoencoder model; determining a feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training; the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of other participants in the federated training.

[0184] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: obtaining shared model parameters from the second model parameters; the shared model parameters are model parameters of a preset proportion in the second model parameters; aggregating the shared model parameters and the first model parameters to obtain target model parameters; and determining a feature extraction model based on the target model parameters.

[0185] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting the feature image and the clinical indicator information of the subject to be tested into a diagnostic model to obtain a diagnostic result; the diagnostic result is used to characterize the damage to the target organs of hypertension in the subject to be tested; inputting the feature image and the clinical indicator information of the subject to be tested into a risk prediction model to obtain a risk prediction result; the risk prediction result is used to characterize the probability of damage to the target organs of hypertension in the subject to be tested.

[0186] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting feature images and clinical indicator information of the subject to be detected into a risk prediction model to obtain a risk prediction score; determining the risk prediction result based on the risk prediction score and the survival curve; the survival curve is determined based on historical risk prediction scores.

[0187] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps: inputting a fundus image of a subject to be detected into a feature extraction model to obtain a feature image corresponding to the fundus image; the feature extraction model is obtained by federated training of an initial mask autoencoder model based on historical fundus images; inputting the feature image and clinical indicator information of the subject to be detected into an evaluation and recognition model to obtain an evaluation result; the evaluation result includes: the extent of damage to target organs in hypertension, and / or the probability of damage to target organs in hypertension.

[0188] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting the fundus image of the object to be detected into the embedding layer for preliminary feature extraction to obtain a first feature vector; inputting the first feature vector into the encoder for feature abstraction to obtain a second feature vector of a preset length; the preset length is greater than the length of the first feature vector; and inputting the second feature vector into the decoder to obtain a feature image corresponding to the fundus image.

[0189] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: masking a first preset number of images in the historical fundus images to obtain mask images; labeling a second preset number of images in the historical fundus images to obtain labeled images; and performing federated training on an initial mask autoencoder model based on the mask images, labeled images, and historical fundus images to obtain a feature extraction model.

[0190] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: training an initial mask autoencoder model based on a mask image and historical fundus images to obtain a first intermediate mask autoencoder model; training the first intermediate mask autoencoder model based on annotated images and historical fundus images to obtain a second intermediate mask autoencoder model; determining a feature extraction model based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of other participants in the federated training; the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of other participants in the federated training.

[0191] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: obtaining shared model parameters from the second model parameters; the shared model parameters are model parameters of a preset proportion in the second model parameters; aggregating the shared model parameters and the first model parameters to obtain target model parameters; and determining a feature extraction model based on the target model parameters.

[0192] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting the feature image and the clinical indicator information of the subject to be tested into a diagnostic model to obtain a diagnostic result; the diagnostic result is used to characterize the damage to the target organs of hypertension in the subject to be tested; inputting the feature image and the clinical indicator information of the subject to be tested into a risk prediction model to obtain a risk prediction result; the risk prediction result is used to characterize the probability of damage to the target organs of hypertension in the subject to be tested.

[0193] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: inputting feature images and clinical indicator information of the subject to be detected into a risk prediction model to obtain a risk prediction score; determining the risk prediction result based on the risk prediction score and the survival curve; the survival curve is determined based on historical risk prediction scores.

[0194] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0195] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0196] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image recognition method for fundus images, comprising: The fundus image of the object to be detected is input into the feature extraction model to obtain the feature image corresponding to the fundus image; The feature extraction model mentioned above is obtained by federated training of an initial mask autoencoder model based on historical fundus images; The feature image and the clinical indicator information of the subject to be detected are input into the evaluation and recognition model to obtain the evaluation result; the evaluation result includes: the damage status of target organs in hypertension, and / or the probability of damage to target organs in hypertension.

2. The method according to claim 1, characterized in that, The feature extraction model includes an embedding layer, an encoder, and a decoder; the fundus image of the object to be detected is input into the feature extraction model to obtain the feature image corresponding to the fundus image, including: The fundus image of the object to be detected is input into the embedding layer for preliminary feature extraction to obtain a first feature vector; The first feature vector is input to the encoder for feature abstraction to obtain a second feature vector of a preset length; wherein the preset length is greater than the length of the first feature vector. The second feature vector is input into the decoder to obtain the feature image corresponding to the fundus image.

3. The method according to claim 2, characterized in that, Before inputting the fundus image of the object to be detected into the feature extraction model, the method further includes: A first preset number of images in the historical fundus images are masked to obtain a masked image; A second preset number of images in the historical fundus images are labeled to obtain labeled images; Based on the mask image, the labeled image, and the historical fundus image, the initial mask autoencoder model is federated to obtain the feature extraction model.

4. The method according to claim 3, characterized in that, Based on the mask image, the labeled image, and the historical fundus image, the initial mask autoencoder model is federatedly trained to obtain the feature extraction model, including: The initial mask autoencoder model is trained based on the mask image and the historical fundus image to obtain the first intermediate mask autoencoder model; The first intermediate mask autoencoder model is trained based on the labeled image and the historical fundus image to obtain the second intermediate mask autoencoder model. The feature extraction model is determined based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of the other participants in the federated training; the second model parameters are the model parameters of the second intermediate mask autoencoder model obtained from the devices of the other participants in the federated training.

5. The method according to claim 4, characterized in that, Based on the first model parameters of the second intermediate mask autoencoder model and the second model parameters of the other participants in the federated training, the feature extraction model is determined, including: Obtain shared model parameters from the second model parameters; the shared model parameters are model parameters in the second model parameters at a preset ratio; The shared model parameters and the first model parameters are aggregated to obtain the target model parameters; The feature extraction model is determined based on the target model parameters.

6. The method according to any one of claims 1-5, characterized in that, The assessment and identification model includes a diagnostic model and / or a risk prediction model. The step of inputting the feature image and the clinical indicator information of the subject to be tested into the assessment and identification model to obtain the assessment result includes: The feature image and the clinical indicator information of the subject to be tested are input into the diagnostic model to obtain a diagnostic result; the diagnostic result is used to characterize the target organ damage of the subject to be tested due to hypertension. The feature image and the clinical indicator information of the subject to be tested are input into the risk prediction model to obtain the risk prediction result; the risk prediction result is used to characterize the probability of damage to the target organs of hypertension in the subject to be tested.

7. The method according to claim 6, characterized in that, The feature image and the clinical indicator information of the subject to be detected are input into the risk prediction model to obtain the risk prediction result, including: The feature image and the clinical indicator information of the subject to be detected are input into the risk prediction model to obtain a risk prediction score; The risk prediction result is determined based on the risk prediction score and the survival curve; the survival curve is determined based on historical risk prediction scores.

8. The method according to claim 7, characterized in that, The feature image and the clinical indicator information of the subject to be detected are input into the risk prediction model to obtain the risk prediction score, including: The feature image and the clinical indicator information of the object to be detected are concatenated in vector form to obtain the model input vector; The model input vector is input into the risk function of the risk prediction model to obtain the risk prediction score.

9. The method according to claim 7 or 8, characterized in that, Based on the risk prediction score and the survival curve, the risk prediction result is determined, including: The model input vectors of the study population are used as the training set, and the historical risk prediction scores are determined based on each model input vector in the training set. The median of the multiple historical risk prediction scores is determined as the dividing threshold, and individuals with negative hypertension are divided into low-risk and high-risk groups based on the dividing threshold. The survival curves for the two groups were plotted using the Kaplan-Meier survival analysis method, and then the risk prediction results were determined based on the risk prediction scores and the survival curves.

10. The method according to any one of claims 1-9, characterized in that, Clinical indicators of the subjects to be tested may include, for example, the subject's age, gender, body mass index (BMI), smoking status, diabetes, triglycerides (TG), low-density lipoprotein cholesterol (LDL-C), and high-density lipoprotein cholesterol (HDL-C).

11. The method according to any one of claims 3-5, characterized in that, To obtain a masked image, a first preset number of images from the historical fundus images are masked. This includes: randomly selecting the first preset number of images from the historical fundus images and masking a portion of the image blocks of the first preset number of images to obtain a masked image, wherein the first preset number is less than or equal to the number of historical fundus images.

12. The method according to any one of claims 3-5, characterized in that, Labeling the second preset number of images in the historical fundus images to obtain the labeled images includes: randomly selecting the second preset number of images from the historical fundus images, and labeling the hypertensive target organ damage areas in the second preset number of images to obtain the labeled images; wherein the second preset number may be less than or equal to the number of historical fundus images.

13. The method according to claim 5, characterized in that, Aggregating the shared model parameters and the first model parameters to obtain target model parameters includes: performing a weighted summation of the shared model parameters and the first model parameters to obtain the target model parameters.

14. An image recognition device for fundus images, characterized in that, The device includes: The first determining module is used to input the fundus image of the object to be detected into the feature extraction model to obtain the feature image corresponding to the fundus image; the feature extraction model is obtained by federated training of an initial mask autoencoder model based on historical fundus images; The second determining module is used to input the feature image and the clinical indicator information of the subject to be detected into the evaluation and recognition model to obtain the evaluation result; the evaluation result includes: the damage status of target organs in hypertension, and / or the probability of damage to target organs in hypertension.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.