A feature extraction model construction method and device, electronic equipment and storage medium

By screening and preprocessing eye image data, and employing a masking scheme and self-supervised learning to train a feature extraction model, the problem of existing eye diagnosis models requiring a large number of samples for training was solved, achieving accurate disease diagnosis and improving model accuracy.

CN118692134BActive Publication Date: 2025-11-11CAPITALBIO CORP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411157162.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-11-11
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

Existing visual diagnosis models require a large number of samples to be collected and trained for each disease during the construction process, which makes data acquisition and annotation difficult, the model accuracy is poor, and it lacks information-based and visual presentation, resulting in ambiguity and subjectivity.

Method used

By preparing and preprocessing eye image data, a masking scheme is used for occlusion processing. Self-supervised learning is used for model training to build a feature extraction model that can extract eye image features. The model is given the ability to extract eye image features and is fine-tuned in downstream tasks to obtain an accurate eye diagnosis model.

Benefits of technology

This enables the creation of accurate disease diagnosis models that require only a small number of disease-specific samples in downstream tasks, breaking through the bottleneck of data annotation, improving the accuracy and robustness of the models, and reducing the need for a large number of samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692134B_ABST
    Figure CN118692134B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, electronic device, and storage medium for constructing a feature extraction model. The method and apparatus are applied to an electronic device to construct a feature extraction model capable of extracting features from eye images. Specifically, the method involves preparing eye image data, which includes multiple filtered and preprocessed eye images, comprising multiple normal eye images and multiple abnormal eye images containing anomalous features. The eye image data is then occluded using a masking scheme to obtain a training dataset. Based on a designed reconstruction goal, the model is trained using the training dataset to obtain the feature extraction model. This application endows the feature extraction model with the ability to extract eye image features by performing non-disease-specific model pre-training on eye images. Thus, in downstream tasks, only a small number of disease-specific samples are needed to fine-tune the model to obtain an eye diagnosis model capable of accurately diagnosing diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a method, apparatus, electronic device, and storage medium for constructing a feature extraction model. Background Technology

[0002] For a long time, traditional medicine, represented by Traditional Chinese Medicine (TCM), has focused on the "external" level of diagnosis, such as observation, auscultation and olfaction, inquiry, and palpation. Modern medicine, represented by Western medicine, focuses on the "internal" level, such as biochemistry, immunology, molecular biology, and genetics. However, Chinese medicine in the new era should be a precise diagnosis that is consistent with both the external and internal aspects, a precise diagnosis that combines TCM and Western medicine, and a precise diagnosis that integrates macroscopic and microscopic perspectives. Only by achieving true consistency between the external and internal aspects can we realize truly precise health management.

[0003] Traditional medicine believes that subtle color changes in the human body due to disease can be observed through the eyes. Modern medicine also believes that changes in the bulbar conjunctival microcirculation can reflect the overall condition of the body, with corresponding changes in the bulbar conjunctival microcirculation regardless of the disease. Therefore, eye diagnosis is not only an important technique for "prevention of disease" but also an excellent pathway for integrating traditional Chinese and Western medicine. However, due to the lack of a physical medium to represent the eye and the absence of informational and visual presentation methods, traditional eye diagnosis relied heavily on the doctor's naked eye observation and personal experience. This naturally led to the formation of various schools of thought in eye diagnosis, resulting in significant ambiguity, subjectivity, and instability. This has caused numerous inconveniences for clinical practice, teaching, and research in eye diagnosis, severely limiting its effectiveness.

[0004] Currently, researchers have applied artificial intelligence technology to eye diagnosis, achieving digitalization, objectification, and intelligence in this field, with promising results. However, existing methods require training a model for a specific disease by collecting a large number of eye images from case and control groups as training samples. When analyzing a different disease, it is necessary to collect a large number of samples for the different disease and train the model again. Therefore, acquiring and labeling a large number of samples is a prerequisite for model construction. In the biomedical field, sample acquisition and labeling have their own unique characteristics: 1) the number of patient samples is much smaller than the number of non-patient samples, resulting in data imbalance; 2) labeling medical data requires a high level of professional expertise and experience from doctors. Therefore, the current approach of building a separate eye diagnosis model for each disease is quite difficult, leading to generally poor accuracy in the resulting eye diagnosis models. Summary of the Invention

[0005] In view of this, this application provides a method, apparatus, electronic device and storage medium for constructing a feature extraction model, which is used to obtain a feature extraction model that can realize eye image feature extraction, enabling users to obtain a high-precision eye diagnosis model based on the feature extraction model in downstream tasks.

[0006] To achieve the above objectives, the following solution is proposed:

[0007] A method for constructing a feature extraction model, applied to electronic devices, for constructing a feature extraction model capable of extracting eye image features, the method comprising the following steps:

[0008] Prepare eye image data, which includes multiple filtered and preprocessed eye image images, including multiple normal eye image images and multiple abnormal eye image images containing abnormal features;

[0009] The eye image data is occluded according to the masking scheme to obtain the training dataset;

[0010] Based on the designed reconstruction goals, the model is trained using the training dataset to obtain the feature extraction model.

[0011] Optionally, the preparation of eye image data includes the following steps:

[0012] Multiple eye images are acquired, and each eye image undergoes quality control processing.

[0013] The eye image image that has undergone quality control processing is preprocessed to obtain the eye image data.

[0014] Optionally, the preparation of eye image data further includes the step of:

[0015] The eye image data is then augmented.

[0016] Optional steps may also be included:

[0017] The feature extraction model is used for downstream tasks to obtain an eye diagnosis model for a specific disease.

[0018] A feature extraction model construction apparatus, applied to an electronic device, is used to construct a feature extraction model capable of extracting eye image features. The construction apparatus includes:

[0019] The data preparation module is configured to prepare eye image data, which includes multiple filtered and preprocessed eye image images, including multiple normal eye image images and multiple abnormal eye image images containing abnormal features.

[0020] The masking module is configured to perform occlusion processing on the eye image data according to the masking scheme to obtain a training dataset.

[0021] The first training module is configured to train the model using the training dataset according to the designed reconstruction goal, so as to obtain the feature extraction model.

[0022] Optionally, the data preparation module includes:

[0023] The image acquisition unit is configured to acquire multiple eye images and perform quality control processing on each eye image;

[0024] The preprocessing unit is configured to preprocess the eye image image that has undergone quality control processing to obtain the eye image data.

[0025] Optionally, the data preparation module further includes:

[0026] The data augmentation unit is configured to augment the eye image data.

[0027] Optionally, the construction method further includes:

[0028] The second training module is configured to use the feature extraction model to perform downstream tasks and obtain an eye diagnosis model for a specific disease.

[0029] An electronic device includes at least one processor and a memory connected to the processor, wherein:

[0030] The memory is used to store computer programs or instructions;

[0031] The processor is used to execute the computer program or instructions to enable the electronic device to implement the feature extraction model construction method as described above.

[0032] A computer-readable storage medium is applied to an electronic device, the storage medium carrying one or more computer programs that can be executed by the electronic device, thereby enabling the electronic device to implement the feature extraction model construction method as described above.

[0033] As can be seen from the above technical solution, this application discloses a method, apparatus, electronic device, and storage medium for constructing a feature extraction model. This method and apparatus are applied to an electronic device to construct a feature extraction model capable of extracting features from eye images. Specifically, it involves preparing eye image data, which includes multiple filtered and preprocessed eye images, including multiple normal eye images and multiple abnormal eye images containing abnormal features; occluding the eye image data according to a masking scheme to obtain a training dataset; and training the model using the training dataset according to the designed reconstruction goal to obtain the feature extraction model. This application endows the feature extraction model with the ability to extract eye image features by performing non-disease-specific model pre-training on eye images. In this way, in downstream tasks, only a small number of disease-specific samples are needed to fine-tune the model to obtain an eye diagnosis model capable of accurately diagnosing diseases, greatly overcoming the bottleneck of data annotation. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart illustrating a method for constructing a feature extraction model according to an embodiment of this application;

[0036] Figure 2 A schematic diagram of eye images from multiple perspectives;

[0037] Figure 3 This is a schematic diagram of preprocessed eye image data in an embodiment of this application;

[0038] Figure 4 A flowchart illustrating another method for constructing a feature extraction model according to an embodiment of this application;

[0039] Figure 5 This is a block diagram of a feature extraction model construction apparatus according to an embodiment of this application;

[0040] Figure 6 This is a block diagram of an apparatus for constructing another feature extraction model according to an embodiment of this application;

[0041] Figure 7 This is a block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0043] Self-supervised learning aims to improve a model's feature extraction capabilities by designing auxiliary tasks to mine the inherent representational characteristics of unlabeled data as supervisory information. It has been widely applied in computer vision and natural language processing, demonstrating excellent results. In practical applications, the model can first be pre-trained using self-supervised learning, allowing it to focus on learning the features of the data. Then, the pre-trained model can be transferred to downstream tasks in various ways, thereby reducing the sample size requirements for downstream tasks and improving model accuracy. Based on this, this application proposes the following technical solution.

[0044] Figure 1 This is a flowchart illustrating a method for constructing a feature extraction model according to an embodiment of this application.

[0045] like Figure 1 As shown, the construction method provided in this embodiment is applied to an electronic device to construct a feature extraction model based on collected eye image data. This electronic device can be understood as a computer, server, or cloud platform with data computing and information processing capabilities. The construction method specifically includes the following steps:

[0046] S1. Prepare eye image data.

[0047] The eye image data here refers to multiple filtered and preprocessed eye images, including normal eye images and abnormal eye images containing anomalous features. The specific process is as follows:

[0048] First, multiple eye images were collected, including both normal eye images and abnormal eye images containing anomalous features. Specifically, during the collection process, images of both eyes of the sample participants were taken from multiple perspectives, such as... Figure 2 As shown. These include multiple perspectives, including but not limited to upward, downward, leftward, and rightward views.

[0049] After collection, the eye images need to undergo quality control processing. Quality control items include, but are not limited to, screening based on the completeness of the eye image, the exposed area of ​​the eyeball, the diagonal state, hue, brightness, and saturation, so as to remove eye images that do not meet the above conditions, so that the samples generated later can meet the requirements of model training.

[0050] Then, the eye images are preprocessed. Preprocessing includes white balance processing, iris and sclera segmentation, and image size normalization. White balance refers to normalizing the brightness channels of each pair of eye images; iris and sclera segmentation uses pre-trained deep neural networks, including but not limited to Unet, Swin-Unet, and Maxvit-Unet, to distinguish between the iris and sclera; image size normalization scales the image to 512x512 while maintaining the aspect ratio and pads the shorter sides with zeros to eliminate interference from other factors. In this application, when preprocessing the eye images, the brightness channels are first normalized for white balance, then the foreground of the iris and sclera is extracted to remove background areas such as skin, eyelashes, and edges of the eye image acquisition device, and the segmented image size is normalized to 512x512x3 pixels, thus obtaining the eye image data in this application, such as... Figure 3 As shown.

[0051] Furthermore, the aforementioned eye image data can be expanded to improve the model's robustness to feature learning. Specifically, data expansion schemes include data domain expansion and data augmentation. Data domain expansion can utilize task-independent eye images or non-eye images, including but not limited to natural images, biomedical images (CT images, MRT images, OCT images, or ultrasound images), and geographic images. Data augmentation methods include, but are not limited to, image correction (histogram equalization, normalization, white balance, grayscale correction and transformation, or image smoothing), illumination distortion (randomly changing image brightness, contrast, or saturation), geometric distortion (randomly scaling, cropping, flipping, or rotating images), and multi-color space processing (HIS, Lab, YCbCr), etc.

[0052] Specifically, data augmentation methods include not only adding newly collected or generated data to the database, but also fusing images from different domains. For example, an eye image can be divided into n uncovered blocks, and a block of image k < 50% * n can be randomly selected and replaced with an image from another domain. Alternatively, an eye image can be fused with images of the same size from other domains according to certain rules based on pixel values.

[0053] S2. Perform occlusion processing on the eye image data according to the masking scheme.

[0054] The training dataset is obtained by occluding eye images, and this dataset is used for subsequent unsupervised learning. The masking scheme involves masking rate, masking method, and noise addition method. Specifically, the masking rate ranges from 0.1 to 0.9; the masking method includes, but is not limited to, random masking, semantic information integration masking, block-wise masking, and multi-fold masking methods; the noise addition method includes, but is not limited to, adding fixed noise, random noise (salt-and-pepper noise, Gaussian noise, Poisson noise, or heteroscedastic Gaussian noise), etc.

[0055] S3. Use the above training dataset to train the model.

[0056] Based on the designed reconstruction goals, the model framework is trained using the aforementioned training dataset to obtain a feature extraction model capable of feature extraction. Specifically, the model framework includes, but is not limited to, Vision Transformer, CNN, and CNN-Vision Transformer hybrid structures; the reconstruction goals include, but are not limited to, the original image pixel values ​​and features calculated based on the original image pixel values, such as oriented gradient histograms; the batch size is an integer factored by 2 and 5; the learning rate ranges from (0.0000001, 1); and the optimization methods include, but are not limited to, Ranger, Adam, SGD, RMSProp, and AdaGrad.

[0057] As can be seen from the above technical solution, this application provides a method for constructing a feature extraction model. This method is applied to electronic devices and is used to construct a feature extraction model capable of extracting features from eye images. Specifically, it involves preparing eye image data, which includes multiple filtered and preprocessed eye images, including multiple normal eye images and multiple abnormal eye images containing abnormal features; occluding the eye image data according to a masking scheme to obtain a training dataset; and training the model using the training dataset according to the designed reconstruction goal to obtain the feature extraction model. This application endows the feature extraction model with the ability to extract eye image features by performing non-disease-specific model pre-training on eye images. In this way, in downstream tasks, only a small number of disease-specific samples are needed to fine-tune the model to obtain an eye diagnosis model capable of accurately diagnosing diseases.

[0058] In addition, such as Figure 4 As shown, in one specific embodiment of this application, the following steps are also included:

[0059] S4. Utilize the feature extraction model to perform downstream tasks.

[0060] Through downstream tasks with specific objectives, eye diagnosis models targeting specific diseases are obtained. Specifically, downstream tasks include, but are not limited to, classification, segmentation, reconstruction, and super-resolution tasks of ocular surface images. Classification tasks include predicting the type of disease a sample belongs to (or a healthy group) based on the disease; segmentation tasks include segmenting ocular surface structures, such as extracting the iris; reconstruction tasks include generating more realistic ocular surface images based on the currently acquired images; and super-resolution tasks include generating higher-resolution images from lower-resolution images. The feature extraction model structure and weights can be directly used as part or all of the backbone network for downstream tasks. This process does not require retraining or only requires simple fine-tuning, thus greatly breaking through the bottleneck of data annotation.

[0061] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0062] Although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous.

[0063] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0064] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer.

[0065] Figure 5 This is a block diagram of a feature extraction model construction apparatus according to an embodiment of this application.

[0066] like Figure 5 As shown, the construction device provided in this embodiment is applied to an electronic device for constructing a feature extraction model based on the collected eye image data. The electronic device can be understood as a computer, server, or cloud platform with data computing and information processing capabilities. The construction device specifically includes a data preparation module 10, a mask processing module 20, and a first training module 30.

[0067] The data preparation module is used to prepare eye image data.

[0068] The eye image data here refers to multiple filtered and preprocessed eye image images, including normal eye image images and abnormal eye image images containing anomalous features. This module includes an image acquisition unit and a preprocessing unit.

[0069] The image acquisition unit is used to acquire multiple eye images, including both normal eye images and abnormal eye images containing anomalous features. Specifically, during acquisition, images of the sample subjects' eyes from multiple perspectives are captured, such as... Figure 2 As shown. These include multiple perspectives, including but not limited to upward, downward, leftward, and rightward views.

[0070] After collection, the eye images need to undergo quality control processing. Quality control items include, but are not limited to, screening based on the completeness of the eye image, the exposed area of ​​the eyeball, the diagonal state, hue, brightness, and saturation, so as to remove eye images that do not meet the above conditions, so that the samples generated later can meet the requirements of model training.

[0071] The preprocessing unit is used to preprocess the eye images. Preprocessing includes white balance processing, iris and sclera segmentation, and image size normalization. White balance refers to normalizing the brightness channels of each binocular eye image; iris and sclera segmentation uses pre-trained deep neural networks, including but not limited to Unet, Swin-Unet, and Maxvit-Unet, to distinguish between the iris and sclera; image size normalization scales the image to 512x512 while maintaining the aspect ratio and pads the shorter sides with zeros to eliminate interference from other factors. In this application, when preprocessing the eye images, the brightness channels are first normalized for white balance, then the foreground of the iris and sclera is extracted to remove background areas such as skin, eyelashes, and edges of the eye image acquisition device, and the segmented image size is normalized to 512x512x3 pixels, thus obtaining the eye image data in this application, such as... Figure 3 As shown.

[0072] Additionally, a data augmentation unit may be included, which expands the aforementioned eye image data to improve the model's robustness to feature learning. Specifically, the data augmentation scheme includes data domain expansion and data augmentation. Data domain expansion can utilize task-independent eye images or non-eye images, including but not limited to natural images, biomedical images (CT images, MRT images, OCT images, or ultrasound images), and geographic images. Data augmentation methods include, but are not limited to, image correction (histogram equalization, normalization, white balance, grayscale correction and transformation, or image smoothing), illumination distortion (randomly changing image brightness, contrast, or saturation), geometric distortion (randomly scaling, cropping, flipping, or rotating images), and multi-color space processing (HIS, Lab, YCbCr), etc.

[0073] Specifically, data augmentation methods include not only adding newly collected or generated data to the database, but also fusing images from different domains. For example, an eye image can be divided into n uncovered blocks, and a block of image k < 50% * n can be randomly selected and replaced with an image from another domain. Alternatively, an eye image can be fused with images of the same size from other domains according to certain rules based on pixel values.

[0074] The masking module is used to occlude eye image data according to the masking scheme.

[0075] The training dataset is obtained by occluding eye images, and this dataset is used for subsequent unsupervised learning. The masking scheme involves masking rate, masking method, and noise addition method. Specifically, the masking rate ranges from 0.1 to 0.9; the masking method includes, but is not limited to, random masking, semantic information integration masking, block-wise masking, and multi-fold masking methods; the noise addition method includes, but is not limited to, adding fixed noise, random noise (salt-and-pepper noise, Gaussian noise, Poisson noise, or heteroscedastic Gaussian noise), etc.

[0076] The first training module is used to train the model using the training dataset mentioned above.

[0077] Based on the designed reconstruction goals, the model framework is trained using the aforementioned training dataset to obtain a feature extraction model capable of feature extraction. Specifically, the model framework includes, but is not limited to, Vision Transformer, CNN, and CNN-Vision Transformer hybrid structures; the reconstruction goals include, but are not limited to, the original image pixel values ​​and features calculated based on the original image pixel values, such as oriented gradient histograms; the batch size is an integer factored by 2 and 5; the learning rate ranges from (0.0000001, 1); and the optimization methods include, but are not limited to, Ranger, Adam, SGD, RMSProp, and AdaGrad.

[0078] As can be seen from the above technical solution, this application provides a feature extraction model construction device. This device is applied to an electronic device and is used to construct a feature extraction model capable of extracting features from eye images. Specifically, it involves preparing eye image data, which includes multiple filtered and preprocessed eye images, including multiple normal eye images and multiple abnormal eye images containing abnormal features; occluding the eye image data according to a masking scheme to obtain a training dataset; and training the model using the training dataset according to the designed reconstruction goal to obtain the feature extraction model. This application endows the feature extraction model with the ability to extract eye image features by performing non-disease-specific model pre-training on eye images. In this way, in downstream tasks, only a small number of disease-specific samples are needed to fine-tune the model to obtain an eye diagnosis model capable of accurately diagnosing diseases.

[0079] In addition, such as Figure 6 As shown, in one specific embodiment of this application, a second training module 40 is also included.

[0080] The second training module is used to perform downstream tasks using the feature extraction model.

[0081] Through downstream tasks with specific objectives, eye diagnosis models targeting specific diseases are obtained. Specifically, downstream tasks include, but are not limited to, classification, segmentation, reconstruction, and super-resolution tasks of ocular surface images. Classification tasks include predicting the type of disease a sample belongs to (or a healthy group) based on the disease; segmentation tasks include segmenting ocular surface structures, such as extracting the iris; reconstruction tasks include generating more realistic ocular surface images based on the currently acquired images; and super-resolution tasks include generating higher-resolution images from lower-resolution images. The feature extraction model structure and weights can be directly used as part or all of the backbone network for downstream tasks. This process does not require retraining or only requires simple fine-tuning, thus greatly breaking through the bottleneck of data annotation.

[0082] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0083] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.

[0084] Figure 7 This is a block diagram of an electronic device according to an embodiment of this application.

[0085] The following is for reference. Figure 7 This document illustrates a structural diagram suitable for implementing the electronic device in the embodiments of this disclosure. The terminal device in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. This electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.

[0086] The electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from an input device 706 into a random access memory (RAM) 703. The RAM also stores various programs and data required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0087] Typically, the following devices can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various devices are shown in the figures, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0088] This application also provides an embodiment of a computer-readable storage medium.

[0089] The aforementioned computer-readable storage medium is applied to an electronic device and carries one or more computer programs. When these computer programs are executed by the electronic device, the device prepares eye image data, which includes multiple filtered and preprocessed eye images, comprising multiple normal eye images and multiple abnormal eye images including anomalous features. The eye image data is then occluded according to a masking scheme to obtain a training dataset. Based on the designed reconstruction goal, the model is trained using the training dataset to obtain a feature extraction model. This application endows the feature extraction model with the ability to extract eye image features by performing non-disease-specific model pre-training on the eye images. Thus, in downstream tasks, only a small number of disease-specific samples are needed to fine-tune the model to obtain an eye diagnosis model capable of accurate disease diagnosis.

[0090] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0091] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), or any suitable combination thereof.

[0092] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0093] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0094] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0095] The technical solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of ​​the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for constructing a feature extraction model, applied to electronic devices, for constructing a feature extraction model capable of extracting eye image features, characterized in that, The construction method includes the following steps: Prepare eye image data, which includes multiple filtered and preprocessed eye images, including multiple normal eye images and multiple abnormal eye images containing abnormal features. The eye image data is occluded according to the masking scheme to obtain the training dataset; Based on the designed reconstruction goals, the model is trained using the training dataset to perform unsupervised learning, thereby obtaining the feature extraction model; Also includes step: The feature extraction model structure and weights are directly used as part or all of the backbone network of the downstream task. The feature extraction model is fine-tuned using disease-specific samples to obtain a visual diagnosis model for a specific disease. The preparation of eye image data includes the following steps: Multiple eye images are collected, and each eye image undergoes quality control processing. The eye image includes images from four perspectives of both eyes, namely, upward, downward, leftward, and rightward views. The quality control processing includes filtering based on the completeness of the eye image, the exposed area of ​​the eyeball, the diagonal state, hue, brightness, and saturation, to remove eye images that do not meet the criteria. The eye image image that has undergone quality control processing is preprocessed to obtain the eye image data. The preprocessing includes foreground extraction of the iris and sclera to remove the background area, which includes skin, eyelashes, and the edge of the eye image acquisition device. It also includes image size normalization, which means scaling the image to 512x512 while maintaining the aspect ratio and filling the short side with 0. The eye image data is augmented by, for example, image fusion from different domains. The fusion of images from different domains includes: dividing the eye image into n uncovered blocks, randomly selecting an image block where k < 50% * n, and replacing it with an image from another domain; or fusing the eye image with images from other domains of the same size according to certain rules based on pixel values.

2. A device for constructing a feature extraction model, applied to an electronic device, for constructing a feature extraction model capable of extracting eye image features, characterized in that, The construction apparatus includes: The data preparation module is configured to prepare eye image data, which includes multiple filtered and preprocessed eye image images, including multiple normal eye image images and multiple abnormal eye image images containing abnormal features. The masking module is configured to perform occlusion processing on the eye image data according to the masking scheme to obtain a training dataset. The first training module is configured to train an unsupervised learning model using the training dataset according to the designed reconstruction goal, so as to obtain the feature extraction model. The second training module is configured to directly use the feature extraction model structure and weights as part or all of the backbone network of the downstream task, and to fine-tune the feature extraction model using disease-specific samples to obtain a visual diagnosis model for a specific disease. The data preparation module includes: The image acquisition unit is configured to acquire multiple eye images and perform quality control processing on each eye image. The eye image includes images from four perspectives of both eyes, namely, upward, downward, leftward, and rightward gaze. The quality control processing includes filtering based on the completeness of the eye image, the exposed area of ​​the eyeball, the diagonal state, hue, brightness, and saturation to remove eye images that do not meet the criteria. The preprocessing unit is configured to preprocess the eye image image after quality control processing to obtain the eye image data. The preprocessing includes foreground extraction of the iris and sclera to remove the background area, which includes skin, eyelashes, and the edge of the eye image acquisition device. It also includes image size normalization processing, which means scaling the image to 512x512 while maintaining the aspect ratio and filling the short side with 0. The data augmentation unit is configured to augment the eye image data, including: image fusion of different domains; The fusion of images from different domains includes: dividing the eye image into n uncovered blocks, randomly selecting an image block where k < 50% * n, and replacing it with an image from another domain; or fusing the eye image with images from other domains of the same size according to certain rules based on pixel values.

3. An electronic device, characterized in that, The electronic device includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs or instructions; The processor is used to execute the computer program or instructions to enable the electronic device to implement the feature extraction model construction method as described in claim 1.

4. A computer-readable storage medium for use in electronic devices, characterized in that, The storage medium carries one or more computer programs that can be executed by the electronic device, thereby enabling the electronic device to implement the feature extraction model construction method as described in claim 1.

Citation Information

Patent Citations

  • Disease detection system and method based on eye image

    CN115496700A

  • Feature extraction model training method and device, electronic equipment and storage medium

    CN117218468A

  • Model training method and device and related equipment

    CN117273116A

  • Target detection method and device, computer equipment and storage medium

    CN117576364A

  • Training method of model for assisting eye disease diagnosis and related device thereof

    CN118380157A