A computer vision-based automatic identification method and system for children's fundus diseases
By combining multi-scale structural analysis and feature fusion of electroretinogram and fundus camera images, a topological perception graph structure is constructed, which solves the problem of insufficient accuracy in the diagnosis of children's fundus diseases and achieves automated and reliable diagnostic results.
Patent Information
- Application Number
- CN202510014190.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Existing fundus disease diagnosis technologies lack accuracy in children, especially due to the sparse vascular features, small blood vessel diameters and variable morphologies, making it difficult for existing technologies to meet the needs of accurate and reliable automated identification.
Combining electroretinogram and fundus camera images, through multi-scale structural analysis, feature enhancement and feature fusion, a topological perception graph structure is constructed, pattern matching and adaptive threshold segmentation are performed, and diagnosis is performed using computer vision technology and fuzzy logic reasoning.
It achieves more automated and reliable diagnosis of children's fundus diseases, reduces human errors, and improves diagnostic efficiency and accuracy.
Smart Images

Figure CN120088194B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroretinograms, and in particular to a method and system for automatically identifying fundus diseases in children based on computer vision. Background Art
[0002] In the field of medical imaging, early screening and diagnosis of fundus diseases are crucial for vision preservation, especially in children. Prompt detection of fundus lesions can effectively prevent irreversible vision damage. Secondly, computer vision, as a key branch of artificial intelligence, has played an increasingly important role in medical image analysis in recent years. Computer vision simulates the human visual system to automatically process and analyze information in images and videos, enabling features such as feature extraction, pattern recognition, and automated decision-making.
[0003] At present, the diagnosis of fundus diseases mainly relies on the subjective judgment of clinicians on fundus images, combined with analysis of physiological signals such as electroretinogram (ERG). However, traditional diagnostic methods have many limitations: First, the doctor's experience level directly affects the accuracy of the diagnosis. In particular, for children's fundus features with sparse vascular features, small blood vessel diameters, and variable morphology, manual diagnosis is difficult to ensure consistency and accuracy. Second, existing automatic diagnosis technologies mainly rely on models of adult fundus data and lack analysis methods specifically for children's fundus features, which leads to poor applicability of existing technologies in children's scenarios. Some current technologies deal with these problems through simple image enhancement and statistical analysis, but lack adaptability to children's unique developmental characteristics, and the diagnostic results are easily affected by noise and individual differences, making it difficult to meet the needs of accurate and reliable automated recognition.
[0004] Based on the above-mentioned shortcomings of the prior art, this application proposes a method and system for automatically identifying pediatric fundus diseases based on computer vision. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for automatically identifying pediatric fundus diseases based on computer vision to improve the above-mentioned problems. To achieve the above-mentioned purpose, the technical solution adopted by the present invention is as follows:
[0006] On the one hand, the present application provides a method for automatically identifying pediatric fundus diseases based on computer vision, comprising:
[0007] Acquiring first information and second information, wherein the first information is an electroretinogram of the child acquired by a non-invasive device, and the second information is an image of the child's retina captured by a fundus camera;
[0008] Performing feature extraction processing based on the first information to obtain an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time, and waveform characteristic parameters;
[0009] performing multi-scale structural analysis and feature enhancement processing based on the second information, extracting blood vessel diameter, branch point number, and blood vessel sparsity features and performing feature enhancement to obtain an image feature set;
[0010] Performing feature fusion processing based on the electroretinogram feature set and the image feature set, and obtaining a fusion feature by weightedly splicing the two types of features into a comprehensive feature vector;
[0011] Performing dimensionality reduction processing on the fused features to obtain a comprehensive feature representation, and constructing a topological perception map structure based on the comprehensive feature representation, wherein the topological perception map structure includes a vascular feature map, an electroretinogram signal feature map, a retinal morphology feature map, and a corresponding hierarchical feature organization;
[0012] Performing pattern matching and adaptive threshold segmentation processing based on the topological perception graph structure to obtain vascular pattern information by identifying the changing patterns of retinal blood vessels in children at different growth stages;
[0013] A diagnostic result is obtained by performing a diagnostic analysis based on the comprehensive feature representation and the blood vessel pattern information.
[0014] On the other hand, the present application also provides a device for automatically identifying pediatric fundus diseases based on computer vision, comprising:
[0015] an acquisition module, configured to acquire first information and second information, wherein the first information is an electroretinogram of the child acquired by a non-invasive device, and the second information is an image of the child's retina captured by a fundus camera;
[0016] An extraction module performs feature extraction processing based on the first information to obtain an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time and waveform characteristic parameters;
[0017] an enhancement module, configured to perform multi-scale structural analysis and feature enhancement processing based on the second information, extracting features of blood vessel diameter, number of branch points, and blood vessel sparsity and performing feature enhancement to obtain an image feature set;
[0018] a fusion module, configured to perform feature fusion processing based on the electroretinogram feature set and the image feature set, and obtain a fusion feature by weightedly concatenating the two types of features into a comprehensive feature vector;
[0019] a construction module for performing dimensionality reduction processing on the fused features to obtain a comprehensive feature representation, and constructing a topological perception map structure based on the comprehensive feature representation, wherein the topological perception map structure includes a vascular feature map, an electroretinogram signal feature map, a retinal morphology feature map, and a corresponding hierarchical feature organization;
[0020] a matching module for performing pattern matching and adaptive threshold segmentation processing based on the topological perception graph structure, and obtaining vascular pattern information by identifying the changing patterns of retinal blood vessels in children at different growth stages;
[0021] The diagnosis module performs a diagnosis analysis based on the comprehensive feature representation and the blood vessel pattern information to obtain a diagnosis result.
[0022] The beneficial effects of the present invention are:
[0023] This invention uses multi-scale structural analysis, convolutional neural networks, graph convolutional networks and other computer vision technologies to process and extract complex features in fundus images, accurately extracting children's fundus features from multiple angles such as blood vessel diameter, branching points, and sparsity, and combines algorithms such as dynamic time warping and fuzzy logic reasoning to comprehensively evaluate the changing patterns of disease characteristics, providing more automated and reliable diagnostic results, reducing human errors, and improving diagnostic efficiency and accuracy.
[0024] Other features and advantages of the present invention will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 A flowchart of a method for automatically identifying fundus diseases in children based on computer vision according to an embodiment of the present invention;
[0027] Figure 2 4 is a block diagram of the fuzzy controller described in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0029] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0030] Example 1:
[0031] This embodiment provides a method for automatically identifying pediatric fundus diseases based on computer vision.
[0032] See also Figure 1 , the figure shows that the method includes steps S100 to S700.
[0033] Step S100: obtaining first information and second information, wherein the first information is an electroretinogram of the child acquired by a non-invasive device, and the second information is an image of the child's retina captured by a fundus camera;
[0034] In data collection, the first thing to do is to obtain the child's electroretinogram (ERG) signal, which is an important fundus biological signal that can reflect the electrophysiological activity of the child's retina when stimulated by light. The acquisition of ERG data is carried out in a non-invasive manner to ensure the comfort and safety of the acquisition process for the child patient. Secondly, the acquired fundus color images are taken by a high-resolution fundus camera. These images contain the overall structural information of the retina, including key parts such as vascular morphology, macular area, and optic disc. The joint acquisition of retinal images and ERG signals can provide a rich data basis for the subsequent extraction and fusion of multimodal features, which helps to comprehensively evaluate the functional and structural status of the retina.
[0035] Step S200: performing feature extraction processing based on the first information to obtain an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time, and waveform feature parameters;
[0036] It is understood that the amplitudes of the A and B waves represent the response strengths of different retinal cell layers to light stimulation, respectively. These features can help identify functional retinal damage. Extracting these amplitude values requires first detecting the local minima and maxima of the electroretinogram signal. Signal processing algorithms, such as the second-order derivative method or threshold detection, can be used to precisely locate the peak positions of the A and B waves. Delay time reflects the time difference between the A and B waves of the signal. This feature can be used to identify delays caused by abnormal neural conduction or photoreceptor function. To extract the delay time, the time difference between the two peaks is calculated to obtain delay information. In addition, waveform feature parameters are extracted using time-frequency analysis techniques such as wavelet transform. Wavelet transform can decompose the electroretinogram signal into multi-scale time and frequency components, thereby extracting shape features, frequency variations, and signal symmetry. These features not only describe the overall morphology of the signal but also capture the local characteristics of abnormalities.
[0037] Step S300: performing multi-scale structural analysis and feature enhancement processing based on the second information, extracting blood vessel diameter, branch point number, and blood vessel sparsity features and performing feature enhancement to obtain an image feature set;
[0038] It's important to note that multi-scale structural analysis and feature enhancement can effectively extract and enhance vascular features in children's fundus images, particularly those small vessels and sparse areas that are difficult to identify in low-contrast or small-scale images. These processing steps significantly improve the usability and accuracy of image features, providing a solid foundation for subsequent feature fusion and disease classification, thereby improving the overall accuracy of automatic identification of pediatric fundus diseases.
[0039] Step S400: performing feature fusion processing based on the electroretinogram feature set and the image feature set, and obtaining a fused feature by weightedly concatenating the two types of features into a comprehensive feature vector;
[0040] It is understandable that the weighted splicing method can flexibly combine multimodal information and significantly enhance the accuracy of feature representation of fundus diseases. In the case of pediatric fundus diseases, in particular, since electrogram and image features show differences at different ages, dynamic weight adjustment can effectively adapt to feature changes at different developmental stages, providing more reliable feature input for subsequent pattern recognition and anomaly detection, ultimately improving diagnostic accuracy and reducing the risk of misdiagnosis.
[0041] Step S500: performing dimensionality reduction processing on the fused features to obtain a comprehensive feature representation, and constructing a topological perception graph structure based on the comprehensive feature representation. The topological perception graph structure includes a vascular feature map, an electroretinogram signal feature map, a retinal morphology feature map, and corresponding hierarchical feature organization;
[0042] It's important to note that the topologically aware map organizes different feature information in a hierarchical manner, forming a complete lesion map from local features to global structure. This structured representation not only captures the characteristic relationships between different lesion patterns in children's fundus, but also uses topological analysis to identify key feature regions that are highly correlated with disease types. This multi-level feature organization helps reveal the connections between fundus features and the complete picture of lesion patterns.
[0043] Step S600: performing pattern matching and adaptive threshold segmentation processing based on the topological perception graph structure, and obtaining vascular pattern information by identifying the changing patterns of the children's retinal blood vessels at different growth stages;
[0044] It can be understood that this step, by identifying specific patterns of vascular morphology and density in children's retinas, generates vascular pattern information that can reflect children's developmental characteristics and potential pathologies at different stages of growth. This pattern information can be further used to compare normal and abnormal patterns, identifying possible pathologies or early abnormalities, and thus supporting the accurate diagnosis of pediatric fundus diseases.
[0045] Step S700: Perform diagnostic analysis based on the comprehensive feature representation and blood vessel pattern information to obtain a diagnostic result.
[0046] In specific applications, by inputting the comprehensive feature representation and vascular pattern information into the diagnostic model, the model will predict the type of disease based on the feature relationship learned in the previous training data. For example, in the analysis, if the electroretinogram shows a specific amplitude abnormality, and the vascular pattern information shows that the vascular structure is sparse and the branches are reduced, it is predicted to be a retinal degenerative disease. By comprehensively analyzing these characteristic patterns, it is possible to identify abnormalities that may occur in children's retinas at different stages of growth, and provide early warnings for various types of fundus lesions. Preferably, in this embodiment, the diagnosis and analysis of children's fundus diseases are performed by fuzzy reasoning methods, and the composition block diagram of the fuzzy controller is as follows: Figure 2The input part represents the input data, including the comprehensive feature representation and the vascular pattern information, which provides important features about the retinal health status of children. The fuzzification interface module is responsible for converting the input feature data into fuzzy sets for subsequent processing. In this step, the feature values are mapped into fuzzy sets, reflecting the uncertainty and fuzziness of the features. The inference engine is the core part of the system, which analyzes the fuzzy sets through fuzzy reasoning rules. According to the feature relationships learned from the previous training data, the inference is carried out using fuzzy logic, and the possibility of each type of fundus lesion is obtained. The database part contains the knowledge base and rule base used to support the reasoning process. The knowledge base stores medical knowledge and feature relationships related to fundus diseases, while the rule base contains the logical rules required for fuzzy reasoning. The inference engine uses this information to perform comprehensive analysis. The defuzzification interface module is responsible for converting the results of fuzzy reasoning into specific and explicit diagnostic results, which are then presented through the output module for doctors to understand and apply.
[0047] Further, step S200 includes steps S210 to S240.
[0048] Step S210, according to the first information, carries out data extraction processing, uses the second derivative method to detect the local minimum value of the electroretinogram signal, identifies the peak position of A wave and B wave and calculates the amplitude to obtain A wave amplitude feature and B wave amplitude feature;
[0049] Specifically, the electroretinogram signal is a voltage curve that changes over time, reflecting the electrical physiological activity of the retina under light stimulation. A wave and B wave represent the responses of photoreceptor layer and bipolar cell layer in the retina, respectively. In order to accurately locate the peak positions of these waves, it is necessary to first smooth the original ERG signal to reduce the influence of noise on the signal. On the smoothed signal, the second derivative method can effectively identify the minimum and maximum values of the signal curve. The zero crossing points and extreme points of the second derivative can help find the positions of A wave and B wave, especially the minimum value of A wave and the maximum value of B wave, because these wave peaks represent the peak response of the retina to light stimulation. Once the positions of A wave and B wave are determined, their corresponding amplitudes can be calculated. The amplitude calculation is based on the difference from the baseline to the peak, and these amplitude values reflect the response intensity of different cell layers of the retina. For example, the amplitude of A wave mainly reflects the functional status of photoreceptors in the retina, while the amplitude of B wave represents the health degree of bipolar cells. In the diagnosis of children's fundus, these features are particularly important, because the function of children's retina will have greater fluctuations during development, and the change of amplitude can directly reflect whether the function of the retina is abnormal.
[0050] Step S220, according to the first information, carries out delay time calculation processing, and obtains delay time data by calculating the time delay between A wave and B wave;
[0051] It is to be explained that the A-wave and the B-wave represent the electrophysiological activity in the photoreceptor and bipolar cell layers of the retina, respectively, the A-wave typically appearing as a negative peak of the signal and the B-wave being a subsequent positive peak. By calculating the time difference between the A-wave and the B-wave, the conduction delay of the light signal in the retinal layers can be quantified. Specifically, first the peak positions of the A-wave and the B-wave need to be precisely located, which are typically detected by the second derivative method in the previous step. Then, by calculating the time interval from the A-wave minimum to the B-wave maximum, the so-called latency time is obtained. For the child's retina, the latency time can reveal the efficiency of the light signal transmission from the receptors to the inner retinal layers. If the neural signal conduction is blocked, for example due to developmental abnormalities of the retina, poor blood supply or genetic retinal diseases, the latency time will increase accordingly. The child's electroretinogram shows large individual differences at different stages of development, so this feature of the latency time is particularly important for the diagnosis of early dysfunction, which can help to assess whether there is an abnormal signal conduction in the retina at different stages of development, so as to identify early diseases before they show obvious structural changes.
[0052] In step S230, waveform feature parameter extraction processing is performed according to the first information, wavelet transform results are obtained by using time-frequency analysis of wavelet transform, and waveform feature parameters are generated by extracting shape features, symmetry and frequency components from the wavelet transform results;
[0053] As can be understood, the wavelet transform is a multi-resolution analysis tool suitable for processing non-stationary signals, particularly time-varying and transient signals like the ERG. Unlike the Fourier transform, the wavelet transform provides information in both the time and frequency domains. By decomposing the ERG signal into wavelet coefficients of varying scales, local details and overall changes in the signal can be observed at different frequencies. For children's ERG signals, the wavelet transform can effectively capture signal variations caused by developmental stages, such as the strengthening or weakening of specific frequency components. From the wavelet transform results, waveform shape features are first extracted, including local peaks and valleys, waveform smoothness, and the slope of the signal. Shape features reflect the intensity of the retinal response to light stimulation and its changing trends, helping to identify potential pathological changes in the signal. Second, symmetry features are extracted, primarily assessing the degree of waveform symmetry at different time points. This is important for identifying abnormal responses in the signal. Changes in symmetry may indicate functional abnormalities in certain retinal regions. Especially during the developmental period of children, retinal symmetry often reflects developmental deviations. Finally, the frequency components are extracted to analyze the intensity distribution of different frequencies in the signal. By observing the energy changes in specific frequency bands, specific abnormalities related to retinal nerve function can be identified. For example, abnormalities in the lower frequency part may be related to poor nerve conduction.
[0054] Step S240: Integrate the A-wave amplitude feature, the B-wave amplitude feature, the delay time data, and the waveform feature parameters to obtain an electroretinogram feature set.
[0055] It should be noted that in the feature integration process, all features are uniformly processed and arranged to form a multi-dimensional feature vector. The integration process involves feature standardization to ensure that different types of features have consistent scales, so that in subsequent analysis, some features will not be excessively affected by scale differences. For A-wave and B-wave amplitude features, they can be integrated according to their amplitude values, and for delay time, it is added as a feature representing conduction efficiency. For waveform feature parameters, they are a subset of multi-dimensional features, such as the frequency component features extracted from wavelet transform forming a set of feature vectors, which occupy a part of the dimension in the feature set. The electroretinogram feature set can comprehensively reflect the multi-level response of the child's retina under light stimulation, integrating multiple feature dimensions from the whole to the local and from the time domain to the frequency domain. This feature integration helps to ensure that important information related to retinal function is not missed in subsequent diagnostic analysis. This is particularly important for children, as their retinas are still in the development stage, and there may be large individual differences in neural function and electrophysiological characteristics. By constructing such a comprehensive feature set, the physiological response of the retina under different development states can be more accurately captured, thereby improving the overall sensitivity and accuracy of fundus disease diagnosis.
[0056] Further, the step S300 comprises steps S310 to S340.
[0057] Step S310, according to the second information, a multi-scale structure analysis process is performed, by layering the retinal image and gradually applying Gaussian blur and downsampling, a multi-level image representation under different resolutions is obtained;
[0058] It can be understood that multi-scale structure analysis is an effective image processing technique that can help the model extract local and global information from images under different resolutions. First, Gaussian blur and downsampling are applied to the retinal image to form a multi-level image pyramid. Gaussian blur is used to smooth the image and reduce noise interference, so that important features of different scales can be focused on at each level. Through downsampling operation, the image is gradually reduced, thereby forming a series of image versions from high resolution to low resolution. These versions can effectively represent different levels of detail of the retina, from large structures to fine structures, which is particularly helpful in capturing small blood vessels and local lesion areas in children's retinas.
[0059] The blood vessels in children's retinas are typically thinner and sparsely distributed. These features may be easily masked by background noise in high-resolution images, but in appropriately blurred low-resolution images, these vascular features will be more obvious. Therefore, through layered processing, structural information of retinal images can be obtained from different scales, which helps to further identify the trunks and branches of blood vessels, as well as other important retinal disease features. Multi-level image representations of different resolutions provide rich basic data for subsequent feature extraction. For example, the fine structure of blood vessels can be better captured in high-resolution images, while low-resolution images help to understand the global morphology and distribution trends of retinal blood vessels. Through this multi-level image representation, the system can more accurately locate and segment small blood vessels and potential disease areas in children's retinas.
[0060] Step S320: Performing vascular feature extraction based on the multi-level image representation, identifying the main structure of the blood vessels by binarizing and skeletonizing the image, and analyzing the node connections in the skeleton graph using a graph analysis algorithm. The vessel diameter is measured and branch points with a connection number greater than or equal to 2 are identified to obtain vessel diameter features and branch point number features.
[0061] Furthermore, step S320 includes steps S321 to S324.
[0062] Step S321: binarize the multi-level image representation to obtain a binarized image.
[0063] As you can understand, this step converts the original image into a binary image consisting of only black and white, extracting the vascular regions while suppressing the background. To ensure accurate representation of vascular structure in the binary image, an adaptive thresholding method is used to process different regions differently. This is particularly important in children's retinal images, where blood vessels are sparsely distributed and small, and a uniform threshold cannot fully capture all details.
[0064] Step S322: Skeletonize the binary image to obtain a blood vessel skeleton image.
[0065] It's important to note that skeletonization simplifies the vascular region into a single-pixel centerline, known as the vascular skeleton. This step preserves the vascular topology while removing redundant peripheral pixels, simplifying subsequent analysis. Skeletonization allows for a clearer visualization of the vascular path and branching structure.
[0066] Step S323: Based on the vascular skeleton graph, a graph analysis algorithm is used to analyze the nodes and connections in the graph to obtain vascular diameter features and branch point number features.
[0067] Firstly, the generation of the vascular skeleton map simplifies the vascular network into a "skeleton" form with single-pixel width, making the quantitative analysis of structural features of the vascular network more direct and clear. The graph analysis algorithm starts with the skeleton map and specifically analyzes the nodes and connections. Nodes in the skeleton map represent the intersection, termination or branching points of blood vessels, while connections represent the connectivity between the main stem and branches of blood vessels. By analyzing these node and connection information, the geometric and topological features of the vascular network can be revealed.
[0068] In the calculation of the vascular diameter, the graph analysis algorithm estimates the diameter by the vertical projection method. At each node of the skeleton map, the width of the blood vessel is estimated along the pixel distance perpendicular to the skeleton direction. This method can dynamically adjust the measurement direction according to the specific location of each node to capture the actual diameter changes of the blood vessels. Since the blood vessels of children are usually more delicate than adults, and the width of the blood vessels varies significantly in different regions, the flexible direction projection method can more accurately reflect these changes.
[0069] For the extraction of the number of branching points, the graph analysis algorithm analyzes the number of connections of each node. Branching points are nodes with multiple connections, which represent the bifurcation positions in the vascular network. Specifically, the system traverses all nodes in the skeleton map and calculates the number of adjacent nodes for each node. Nodes with a connection number greater than or equal to 2 are identified as branching points. The number and location of branching points can directly reflect the complexity of the vascular network and the blood supply situation in the local area. In the retina of children, a decrease in the number of branching points or abnormal distribution may indicate the risk of retinal vascular dysplasia or early lesions.
[0070] Step S330, local density estimation processing is performed according to the multi-level image representation, and the vascular sparseness feature is obtained by calculating the number of blood vessel pixels in each region;
[0071] This step provides an important indicator for early identification of pediatric fundus diseases by analyzing the density of blood vessel distribution in the retinal image. The retinal blood vessels of children exhibit sparse and small features, and accurate estimation of this feature can reveal the health status of vascular development. The calculation formula is:
[0072] ;
[0073] wherein, represents the region number; represents the vascular sparseness feature in the region ; represents the area of the region ; represents the indicator function for judging whether a pixel is a blood vessel pixel; represents the coordinates of the pixel point; Represents the coordinates of the pixel points in the field; represents the weight function; represents the balance parameter, which is used to adjust the influence of neighboring blood vessel pixels on the sparsity feature; Indicates area field.
[0074] This method accurately calculates the density of local vascular pixels and, combined with neighborhood information, adjusts the sparsity feature, resulting in more reliable sparsity feature data. In children's retinas, the vascular distribution pattern exhibits significant variations in sparsity due to different growth stages. This sparsity estimate can reflect the vascular structural characteristics of children under normal development and pathological conditions.
[0075] Step S340: Perform feature fusion processing based on the blood vessel diameter feature, the branch point number feature, and the blood vessel sparsity feature to obtain an image feature set.
[0076] First, these features are based on different properties of blood vessels: the vascular diameter feature reflects the width of the blood vessels, which is important for identifying vascular narrowing, dilation and other pathological changes; the branch point number feature represents the branch complexity of the vascular network, which can indicate the health status of retinal blood vessels, especially in children, where the fundus may have abnormal branch numbers due to development or pathological changes; the vascular sparsity feature describes the distribution density of blood vessels, which is crucial for determining whether the retina has characteristic lesions such as vascular sparsity.
[0077] Specifically, during the feature fusion process, these features are first normalized to ensure that they have the same magnitude, thereby preventing any one feature from dominating the fusion result. Then, using a weighted concatenation method, each feature is assigned a different weight based on its importance to disease identification and combined into a comprehensive feature vector. This weighted process gives features with higher diagnostic value a greater influence in the feature set, helping to improve overall diagnostic accuracy.
[0078] This fusion process integrates vascular features from different dimensions into a unified image feature set. This image feature set is highly representative and information-dense, fully reflecting the structural characteristics of children's retinas. The synergistic effect of these different features makes the fused feature set more sensitive to subtle lesions in children's fundus and can capture abnormalities that may be overlooked during normal development.
[0079] Furthermore, step S400 includes steps S410 to S440.
[0080] Step S410, according to the electroretinogram feature set and the image feature set, a feature correlation analysis process is performed, the correlation score between the electroretinogram features and the image features is evaluated by using the maximum information coefficient algorithm, and a feature correlation matrix is obtained;
[0081] It should be noted that the maximum information coefficient (MIC) is an algorithm that can capture linear and nonlinear relationships, suitable for analyzing complex dependency relationships between multi-modal features. In the context of children's electroretinogram and image features, this complex relationship is due to the complementarity and coupling of different features. Through the maximum information coefficient algorithm, not only can the degree of mutual influence between electroretinogram features (such as A-wave amplitude, B-wave amplitude, etc.) and image features (such as blood vessel diameter, branch point number, etc.) be quantified, but also subtle nonlinear correlations between the two can be revealed. For example, changes in certain blood vessel features may be accompanied by fluctuations in specific ERG features, which are often difficult to capture through simple correlation coefficients.
[0082] In actual operation, first, the electroretinogram feature set and the image feature set are combined pairwise for correlation calculation. The maximum information coefficient algorithm is automatically adjusted to adapt to different relationship patterns between data, whether linear or nonlinear. The calculation result will be reflected in the form of correlation scores in the feature correlation matrix. Each element in the matrix represents the correlation between a pair of electroretinogram features and image features. A high score indicates a strong correlation between the features, and a low score indicates a weak or non-existent correlation.
[0083] Step S420, according to the feature correlation matrix, a feature denoising and selection process is performed, sparse principal component analysis is used to constrain the short-time mutation signal in the children's electroretinogram, and the importance of the features in the children's retinal image is weighted and sparsified, to obtain a reduced feature set;
[0084] First, sparse principal component analysis (Sparse PCA) converts feature data into a new space, striving to minimize the redundancy of feature data in the process and retaining only key features. Unlike traditional principal component analysis, sparse principal component analysis introduces a sparsity constraint in the dimensionality reduction process, selectively eliminates features that contribute less to information by adjusting the importance weight of the features. In children's electroretinogram data, this sparsification can effectively remove noise in short-time mutation signals and retain features that are stable and important to light stimulus response. For example, accidental peaks in the electroretinogram signal do not reflect the overall function of the retina, and these irrelevant signals can be effectively filtered out through sparse principal component analysis.
[0085] Meanwhile, in the sparse processing of image features, sparse principal component analysis can select features by weighting according to the importance scores in the feature correlation matrix. This matrix is previously generated by the maximum information coefficient algorithm, representing the correlation between different features, so when dimensionality reduction, according to the importance score of each feature, weighted sparse. This can ensure that the remaining features have the most influence on disease recognition as a whole, such as vascular sparsity, branch point number or diameter characteristics, which can significantly reveal the specific changes of children's retinal structure. Through this sparse processing, the final feature set contains features with high disease relevance, reducing information redundancy.
[0086] Step S430, according to the reduced feature set, perform feature weighting fusion processing, assign weights to each feature through a preset attention mechanism weighting model, automatically learn the importance of features in the fusion process, and obtain a weighted fusion feature vector;
[0087] Specifically, the attention mechanism weighting model will analyze each item of the input reduced feature set, and adjust the weight according to the correlation of each feature with the diagnostic target. The model automatically obtains a higher weight for the features that have a greater impact on the diagnostic result through optimization processes such as back propagation and gradient descent, while the relatively secondary features are assigned to a lower weight. Through learning and adjustment, it can adapt to the performance differences of different feature combinations in different pathological conditions, so that the fused feature vector is more targeted and accurate for diagnosis. This automatic learning of weights can significantly improve the model's recognition accuracy of children's fundus diseases. The feature performance of children's retinas usually varies from individual to individual, and through the attention mechanism weighting model, the feature fusion strategy can be optimized in real time based on the data of each child, ensuring the maximum information utility in diagnosis.
[0088] Step S440, according to the weighted fusion feature vector, perform feature generation processing, construct a third-order tensor representing the weighted electroretinogram features, image features and fusion weights respectively through a multi-modal tensor decomposition fusion algorithm, and perform tensor decomposition processing to obtain a fusion feature based on the dynamic change rule between multi-modal features.
[0089] It should be noted that the construction of the third-order tensor is based on the combined representation of electroretinogram features and image features as different dimensions, and the weighted fusion vector is introduced as the third dimension. Specifically, each element in the third-order tensor contains the influence of electroretinogram features, image features and weights at the same time, which can reflect the interaction between these features in the fusion process. The purpose of tensor decomposition is to further simplify and optimize the multimodal feature set, thereby retaining the core information that is most valuable for disease diagnosis, while reducing the impact of data redundancy and noise. Preferably, the tensor decomposition method includes CANDECOMP / PARAFAC decomposition (CP decomposition) and Tucker decomposition, which can flexibly adapt to the association patterns between different features.
[0090] During the decomposition process, the tensor decomposition method extracts the latent representation of features by analyzing the mutual influence of multimodal data across different dimensions. For children's retinal data, it is able to capture the complex associations between electroretinogram features and image features while retaining important dynamic change patterns. For example, a specific electroretinogram response pattern may be closely related to certain vascular morphological features. This latent feature relationship is revealed and preserved through tensor decomposition. The decomposed fusion feature retains the important correlation between the electroretinogram and image features, forming a fused feature vector that is most expressive for diagnostic tasks. This multimodal tensor decomposition and fusion algorithm integrates the dynamic interactive relationships of different features into a compact and expressive feature representation.
[0091] Furthermore, step S500 includes steps S510 to S530.
[0092] Step S510: Graph construction is performed based on the fused features. Feature adjacency relationships are calculated based on a k-nearest neighbor graph, different features are used as graph nodes, and the connection relationship between nodes is determined by feature similarity to obtain an initial feature graph.
[0093] It can be understood that each feature in the fused feature vector is regarded as a node in the graph, representing different electroretinogram and image features. Then, the k-nearest neighbor graph algorithm is used to establish the connection relationship between these nodes. The k-nearest neighbor graph algorithm searches for the k most similar nodes in the local neighborhood of each node and establishes connections with these similar nodes. The measure of similarity is based on Euclidean distance or cosine similarity to ensure that nodes with highly similar features are connected together. For the diagnosis of pediatric fundus diseases, this connection method can reveal possible implicit associations between functional and structural features, such as the relationship between vascular morphology changes and electroretinogram signal abnormalities.
[0094] This process generates an initial feature map that not only demonstrates direct similarity relationships between multimodal features but also captures underlying patterns and regularities within the fused features. The structure of the feature map reveals the distribution of disease-related features and, through the connectivity of the graph structure, helps identify clusters of features that play a key role in diagnosis. For diagnosing retinal diseases in children, this initial map can reflect the relative positions of individual features in multimodal data and their mutual influence, providing a structured data foundation for subsequent pattern recognition and anomaly detection.
[0095] Step S520: Graph structure generation is performed based on the initial feature graph. By applying the persistent homology method, the topological features of nodes and edges in the graph structure at different scales are analyzed, stable structures in the features are identified, and noise and unstable nodes are removed to obtain a topologically aware graph structure.
[0096] It's important to note that persistent homology is a topological data analysis technique that effectively analyzes complex morphological and topological features in graph structures. It examines the connectivity between nodes and edges at different scales to reveal how features vary across multiple scales. For example, in the context of pediatric fundus data, variations in certain features may be related to age, developmental stage, or specific lesion type. Persistent homology can identify the stability of these features across multiple scales.
[0097] In the specific processing process, the nodes and their adjacency relationships in the initial feature map are first analyzed using the persistent homology method. This method gradually increases the connection radius in the graph and observes the disappearance and merging of nodes and edges at different scales. Structures with high persistence indicate that they exist and are stable at multiple scales. These stable nodes and edges often represent feature information that is more meaningful for disease identification. In contrast, nodes and edges with low persistence are regarded as noise, which may be introduced due to randomness or local instability, and are therefore removed during the optimization process. This process of removing noisy and unstable nodes can significantly improve the overall robustness of the graph structure, ensuring that the final topological perception graph structure is more focused on the core features related to fundus diseases. By screening out structures with high persistence, the map can better demonstrate the stable relationship between disease-related features and exclude irrelevant features that may lead to misjudgment.
[0098] Step S530: hierarchically organize the topological perception map structure, generate a hierarchical structure by hierarchically dividing the vascular features, electroretinogram signal features, and morphological features, and organize the vascular features and electroretinogram features at different hierarchical levels according to the changes in the children's growth characteristics to obtain a topological perception map structure.
[0099] Specifically, hierarchical organization begins by analyzing the impact of different features on fundus diseases. Multimodal features are divided into several layers, including vascular characteristics (such as vessel diameter, branching number, and sparsity), electroretinogram signal characteristics (such as A-wave amplitude, B-wave amplitude, and latency), and morphological characteristics (such as the overall morphology of the retinal structure). Each layer represents a set of relatively independent but biologically closely related feature types. Through this hierarchical division, disease-related patterns can be identified at different levels.
[0100] Next, the organization of features at different hierarchical levels was adjusted to account for changes in children's growth characteristics. For example, children's retinal vasculature and electrogram signals change with age. This physiological change may cause some features to stabilize during development, while others exhibit significant fluctuations at certain stages. To more accurately reflect these growth characteristic changes, developmentally dependent features (such as vascular diameter and electrogram waveforms) were organized at specific hierarchical levels to reflect age-related differences in the graph structure.
[0101] This hierarchical structure captures the different dimensional relationships of features at multiple levels, enabling the system to identify changes in the characteristics of children's fundus at different stages of physiological development. For the diagnosis of pediatric fundus diseases, this hierarchical organization provides a more flexible and targeted analysis approach. For example, at a low level, the system can focus on identifying early vascular abnormalities in younger children, while at a higher level, it can identify structural changes that may occur during growth.
[0102] Furthermore, step S600 includes steps S610 to S640.
[0103] Step S610: Aggregate local information of different feature nodes in the topology perception graph structure based on a preset graph convolutional network, convolve the vascular feature nodes layer by layer, capture the local variation characteristics of the vascular features in their neighborhood, and obtain feature vector representations of all nodes in the graph;
[0104] It should be noted that the main function of a graph convolutional network is to aggregate the information of each node with that of its neighboring nodes through convolution operations, thereby capturing local features in the graph structure. In a topologically aware graph structure, each node represents a feature (such as vessel diameter, number of branch points, sparsity, etc.), and the connection relationship between nodes reflects the similarity or correlation between features. The graph convolutional network transmits feature information to neighboring nodes through layer-by-layer convolution and generates a feature representation that includes local feature changes. For children's fundus images, local changes in vascular features often imply pathological information, such as abnormal branching or increased vascular sparsity. The graph convolutional network can effectively capture these subtle changes.
[0105] Specifically, at each layer, a graph convolutional network (GCN) performs a weighted average of the information from the central node and its neighboring nodes through convolution operations to update the node's feature representation. This process is performed recursively, layer by layer, so that the features of the central node not only contain information from its immediate neighbors but also gradually integrate features from nodes in more distant layers. Ultimately, the feature vector representation of each node not only describes its own characteristics but also reflects the combined information of the features of its connected neighbors. For example, in the post-convolution feature vector, changes in blood vessel diameter may be combined with branch point features from neighboring nodes to form a more discernible graph structure representation. In the diagnosis of pediatric retinal lesions, the application of GCNs can capture microscopic changes in vascular structure through this aggregation of neighborhood features, thereby supporting the identification of complex pathological patterns. Due to developmental factors, retinal blood vessels in children exhibit significant individual variability and dynamic changes. During the convolution process, GCNs automatically learn correlations between features, enabling the system to adapt to these developmental characteristics and detect potential abnormalities.
[0106] Step S620: Perform pattern matching based on the feature vector representation. A dynamic time warping algorithm is used to calculate the similarity of the feature vectors at different time points. The vascular features collected from the same child at different times are then compared to identify vascular development trends or abnormal patterns, thereby obtaining a change pattern.
[0107] It's understandable that the Dynamic Time Warping (DTW) algorithm is an algorithm used to measure the similarity between two time series. It's particularly well-suited for analyzing time series with different time scales. In analyzing retinal vascular patterns in children, the developmental trends and pathological changes exhibited by vascular characteristics over time are inconsistent. The Dynamic Time Warping algorithm can flexibly adjust the matching paths of the time series to capture similar patterns. This is particularly important for capturing changes in vascular structure across children's developmental stages, as changes in physiological characteristics with age often do not follow a linear pattern.
[0108] During the specific operation, the vascular feature vectors collected from the same child at different time points are matched. The dynamic time warping algorithm calculates the shortest path to align the two sets of feature vectors in the time dimension to minimize the cumulative distance between them, thereby identifying developmental trends or abnormal patterns. For example, by comparing the current features with the historical feature vectors, abnormal trends such as the gradual narrowing of the blood vessel diameter or the decrease in the number of branch points can be detected, and even patterns such as increased sparsity that may indicate lesions. The advantage of the dynamic time warping algorithm is that it can adapt to the asynchronous changes that occur during a child's development, which means that even if the rate of change of vascular characteristics varies due to individual differences or the degree of lesions, the matching can still be flexibly adjusted to capture these complex change patterns.
[0109] Step S630: performing local feature adaptive segmentation processing based on the change pattern, estimating the probability distribution of vascular features using a Gaussian mixture model, and adaptively selecting thresholds to segment different types of vascular regions based on the morphological differences of retinal vessels in children of different age groups to obtain segmentation results;
[0110] It should be noted that the Gaussian mixture model is a probabilistic model used to estimate multimodal data distributions. By modeling data as a combination of multiple Gaussian distributions, it can flexibly describe the probability distribution of complex feature data. In the case of pediatric retinal vascular features, vascular structures vary significantly across age groups. The Gaussian mixture model can capture the distribution of these features across different regions, providing a basis for selecting adaptive thresholds.
[0111] Specifically, based on the results of the variation pattern analysis, vascular feature data is first input into a Gaussian mixture model. The model iteratively estimates the probability distribution of vascular features using an expectation-maximization (EM) algorithm and divides the distribution of vascular features into multiple categories. Preferably, the Gaussian mixture model segments the vessels into larger, smaller, and abnormal vascular regions, reflecting the differences in their characteristics under different distribution patterns. Subsequently, based on the estimated distribution, a segmentation threshold is adaptively selected to apply different segmentation strategies to different regions. For different age groups in pediatric fundus images, the threshold can be adjusted according to the developmental stage to accommodate the natural variations in vascular structure. This adaptive segmentation process ensures that the segmentation results accurately reflect the different vascular regions in the child's retina and capture abnormal structures.
[0112] Step S640 : performing topological feature evaluation based on the segmentation results, and obtaining vascular pattern information by calculating the complexity of vascular branches, the number of loops, and connectivity of the blood vessels.
[0113] First, vascular branching complexity reflects the frequency and density of bifurcations in the vascular network. This step quantifies branching complexity by calculating the number and distribution of vascular branch points. A higher branching complexity generally indicates a well-developed vascular network, while in certain pathological conditions, branching complexity may decrease or increase abnormally. Retinal vascular branching in children undergoes dynamic changes during development, so this characteristic can reveal specific patterns in retinal development or pathology. Next, loop count is analyzed, which refers to the number of closed loops in the vascular network. Loop characteristics are related to the redundancy and connectivity of the vascular network. In healthy retinal vascular networks, the number of loops fluctuates within a reasonable range, but an abnormal increase or decrease may indicate underlying structural pathology. By calculating the number of loops in the vascular network graph, abnormal circulatory structures that may be caused by pathology can be identified. Furthermore, vascular connectivity analysis examines the overall connectivity of the vascular network. Graph analysis methods are used to calculate the connectivity between nodes to understand the overall structural coherence of the vascular network. Poor retinal vascular connectivity can reflect conditions such as insufficient blood supply or localized hypoxia. Especially in fundus examinations of children, insufficient connectivity may indicate the presence of developmental abnormalities or lesion areas. Ultimately, these topological features together constitute vascular pattern information, thereby providing richer data support for subsequent disease identification. By evaluating branch complexity, number of loops, and connectivity, comprehensive topological features of children's retinal blood vessels can be generated. In terms of technical effect, this topology-based assessment greatly enhances the sensitivity to retinal vascular abnormalities in children. Vascular pattern information can reveal the macroscopic structural characteristics of the vascular network, provide more in-depth structural diagnostic information, and improve the detection accuracy and diagnostic precision of children's fundus diseases.
[0114] Furthermore, step S700 includes steps S710 to S740.
[0115] Step S710: Input the comprehensive feature representation and vascular pattern information into a preset extreme gradient boosting tree model, use the tree structure to gradually split the data and calculate the contribution score of each feature to the classification result to obtain the feature importance ranking result;
[0116] It should be noted that Extreme Gradient Boosting (XGBoost) is an ensemble learning algorithm that improves the model's predictive capabilities by building a series of decision trees. In each iteration, the XGBoost model generates a new tree and learns based on the residuals (i.e., prediction errors) of the previous tree, ultimately constructing a classifier. Given the comprehensive input feature representation (including multimodal data) and vascular pattern information, XGBoost utilizes its weighted splitting mechanism to accurately identify the contribution of each feature to the final classification.
[0117] Specifically, the extreme gradient boosting tree model associates different features with specific disease states by gradually splitting the data. During the splitting process, the selection of each feature node is based on metrics such as information gain or the Gini coefficient to ensure that impurities are minimized in each split. By calculating the contribution of each feature to the classification result and accumulating the scores, an importance ranking of each feature is finally generated. This ranking reflects the relative influence of the feature in disease classification. For the diagnosis of pediatric fundus diseases, the model can accurately identify which vascular features or electrogram features have the highest explanatory power for identifying specific lesions. The feature importance ranking results provide a list of features for priority analysis. The use of the extreme gradient boosting tree model can efficiently and accurately screen out the features with the most diagnostic value. By ranking the importance of features, resources and analysis efforts can be concentrated on the features that are most critical for disease diagnosis, thereby enhancing overall diagnostic accuracy.
[0118] Step S720: Classify the features according to the importance ranking results. Input the features with feature importance scores greater than the threshold into a preset naive Bayes model, assume conditional independence between the features, and calculate the probability of each disease to obtain the classification results of childhood retinal diseases.
[0119] Specifically, this step uses features with feature importance scores greater than a preset threshold as input data, and selects high-weight features for disease classification. The naive Bayes model calculates the conditional probability of each feature under different disease categories and combines these probabilities to estimate the overall category probability. The most likely disease category is then selected as the classification result based on the maximum posterior probability. For example, if the complexity of vascular branching and specific waveform features of the electroretinogram show high conditional probability in the probability model of a specific disease, the classification result will tend to be assigned to that disease category. This process can effectively use high-weight, important features for classification, thereby reducing unnecessary information redundancy and improving classification efficiency and accuracy. The probability output of the naive Bayes model can also provide confidence levels for different disease categories, helping to understand classification decisions and probability distributions. This intuitive probability estimate provides data support for subsequent medical decisions and interventions.
[0120] Step S730: performing abnormal feature detection based on the comprehensive feature representation and the vascular pattern information, calculating the Mahalanobis distance between each sample and the center of the overall feature distribution, and detecting abnormal vascular distribution or abnormal electroretinogram response to obtain abnormal feature information;
[0121] It should be noted that the Mahalanobis distance is a method for measuring the distance between a multi-dimensional data point and the center of the overall distribution, which can effectively consider the correlation and covariance between features, and is particularly suitable for detecting abnormal points of data. Compared with the Euclidean distance, the Mahalanobis distance can better process multi-dimensional data with complex covariance structure, so in the case of more feature dimensions and mutual relationship between features, the Mahalanobis distance is a more robust distance measurement method. For children's fundus data, since the multi-modal features (such as blood vessel features, electrogram features, etc.) have complex interdependence, the Mahalanobis distance can better capture this correlation and effectively detect abnormal patterns. The calculation formula is:
[0122] ;
[0123] wherein, represents the sample vector to be detected; represents the Mahalanobis distance of the sample ; represents the mean vector of the overall feature distribution; represents the sample covariance matrix; represents the feature serial number; represents the total number of features; represents the importance weight of the feature ; represents the variance of the feature ; represents the transpose of the evidence.
[0124] The abnormality detection method based on the Mahalanobis distance can accurately locate the abnormality of multi-modal feature data, and combined with the feature importance weight, it ensures that the changes in key features are more sensitive. This way greatly improves the abnormality detection ability of the system when facing complex fundus data, so that it can identify the signs of lesions in the children's retina earlier and more accurately.
[0125] Step S740, according to the classification result and the abnormal feature information, the inference processing is carried out, the blood vessel sparseness, the wave amplitude abnormality of the electroretinogram and the delay feature of the electroretinogram are converted into fuzzy sets, and the diagnosis result is obtained through comprehensive analysis by the preset fuzzy inference rule.
[0126] Specifically, these key features (vascular sparsity, ERG amplitude abnormalities, and delay characteristics) are first converted into fuzzy sets. Fuzzy sets are a method for handling imprecise information and can effectively describe concepts with uncertain boundaries. For example, vascular sparsity can be described as "sparse," "moderately sparse," or "dense," and amplitude abnormalities can be expressed as "mildly abnormal," "significantly abnormal," and so on. These fuzzy sets allow for a more natural representation of feature variations, rather than rigidly classifying features into discrete categories. This approach is particularly suitable for pediatric fundus diagnosis, as children's physiological characteristics vary imprecisely with development, and fuzzy sets can more flexibly adapt to these changes. These fuzzy sets are then comprehensively analyzed using pre-set fuzzy inference rules. Fuzzy inference rules are a set of "if-then" rules based on experience and medical knowledge, simulating the way doctors make comprehensive judgments about features during the diagnostic process. For example, a rule might be: "If the degree of vascular rarefaction is 'significantly rarefaction' and the electroretinogram amplitude abnormality is 'significantly abnormal,' then the disease risk is 'high.'" The core of the fuzzy inference process is to assess the likelihood and risk level of a disease by comprehensively considering multiple fuzzy features. This process is similar to how a doctor makes a comprehensive judgment based on a combination of different symptoms when examining a patient's condition. When applying fuzzy inference rules, operations such as the intersection and union of fuzzy sets and weighted reasoning are used to obtain a comprehensive evaluation of the features and convert it into a specific diagnosis.
[0127] Example 2:
[0128] This embodiment provides a computer vision-based automatic identification system for children's fundus diseases, including:
[0129] an acquisition module, configured to acquire first information and second information, wherein the first information is an electroretinogram of a child acquired by a non-invasive device, and the second information is an image of the child's retina captured by a fundus camera;
[0130] An extraction module performs feature extraction processing based on the first information, and obtains an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time, and waveform characteristic parameters;
[0131] an enhancement module, configured to perform multi-scale structural analysis and feature enhancement processing based on the second information, extracting features such as blood vessel diameter, number of branch points, and sparsity of blood vessels and performing feature enhancement to obtain an image feature set;
[0132] A fusion module is used to perform feature fusion processing based on the electroretinogram feature set and the image feature set, and obtain the fusion feature by weighted splicing of the two types of features into a comprehensive feature vector;
[0133] The construction module is configured to perform dimension reduction processing on the fusion features to obtain comprehensive feature representations, and construct a topology-aware graph structure based on the comprehensive feature representations, the topology-aware graph structure including a blood vessel feature map, a retinal electrogram signal feature map, a retinal morphological feature map, and corresponding hierarchical feature organizations;
[0134] The matching module is configured to perform pattern matching and adaptive threshold segmentation processing based on the topology-aware graph structure, to obtain blood vessel pattern information by identifying change patterns of retinal blood vessels of children in different growth stages;
[0135] The diagnosis module is configured to perform diagnostic analysis based on the comprehensive feature representations and the blood vessel pattern information to obtain a diagnosis result.
[0136] In one specific embodiment disclosed in the present application, the extraction module includes:
[0137] The first extraction unit is configured to perform data extraction processing based on the first information, to detect local minimum values of the retinal electrogram signal using a second derivative method, to identify peak value positions of A waves and B waves and calculate amplitudes thereof to obtain A wave amplitude features and B wave amplitude features;
[0138] The first calculation unit is configured to perform delay time calculation processing based on the first information, to obtain delay time data by calculating a time delay between the A wave and the B wave;
[0139] The second extraction unit is configured to perform waveform feature parameter extraction processing based on the first information, to obtain wavelet transform results by performing time-frequency analysis using wavelet transform, and to extract shape features, symmetry, and frequency components from the wavelet transform results to generate waveform feature parameters;
[0140] The first integration unit is configured to integrate the A wave amplitude features, the B wave amplitude features, the delay time data, and the waveform feature parameters to obtain a retinal electrogram feature set.
[0141] In one specific embodiment disclosed in the present application, the enhancement module includes:
[0142] The first analysis unit is configured to perform multi-scale structure analysis processing based on the second information, to obtain multi-level image representations at different resolutions by performing hierarchical processing on the retinal image, and gradually applying Gaussian blur and down-sampling;
[0143] The third extraction unit is configured to perform blood vessel feature extraction processing based on the multi-level image representations, to identify a main stem structure of the blood vessels by performing binarization and skeletonization processing on the image, and to measure diameters of the blood vessels and identify branch points with a connection number greater than or equal to 2 by analyzing node connection conditions in the skeleton map using a graph analysis algorithm, to obtain blood vessel diameter features and branch point number features;
[0144] The second calculation unit is used to perform local density estimation processing according to the multi-level image representation, and obtain the vascular sparsity feature by calculating the number of vascular pixels in each area;
[0145] The first fusion unit is used to perform feature fusion processing based on the blood vessel diameter feature, the branch point number feature and the blood vessel sparsity feature to obtain an image feature set.
[0146] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0147] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for automatically identifying children's fundus diseases based on computer vision, characterized in that: include: Acquiring first information and second information, wherein the first information is an electroretinogram of the child acquired by a non-invasive device, and the second information is an image of the child's retina captured by a fundus camera; Performing feature extraction processing based on the first information to obtain an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time, and waveform characteristic parameters; performing multi-scale structural analysis and feature enhancement processing based on the second information, extracting blood vessel diameter, branch point number, and blood vessel sparsity features and performing feature enhancement to obtain an image feature set; Performing feature fusion processing based on the electroretinogram feature set and the image feature set, and obtaining a fusion feature by weightedly splicing the two types of features into a comprehensive feature vector; Performing dimensionality reduction processing on the fused features to obtain a comprehensive feature representation, and constructing a topological perception map structure based on the comprehensive feature representation, wherein the topological perception map structure includes a vascular feature map, an electroretinogram signal feature map, a retinal morphology feature map, and a corresponding hierarchical feature organization; Performing pattern matching and adaptive threshold segmentation processing based on the topological perception graph structure to obtain vascular pattern information by identifying the changing patterns of retinal blood vessels in children at different growth stages; Performing a diagnostic analysis based on the comprehensive feature representation and the vascular pattern information to obtain a diagnostic result; The feature fusion processing is performed based on the electroretinogram feature set and the image feature set, and the fusion feature is obtained by weightedly splicing the two types of features into a comprehensive feature vector, including: Performing feature correlation analysis on the electroretinogram feature set and the image feature set, and obtaining a feature correlation matrix by evaluating the correlation scores between the electroretinogram features and the image features using a maximum information coefficient algorithm; Performing feature denoising and selection processing based on the feature correlation matrix, performing sparse constraints on short-term mutation signals in the children's electroretinogram through sparse principal component analysis, and performing weighted sparseness on the importance of features in the children's retinal images to obtain a streamlined feature set; Performing feature weighted fusion processing based on the streamlined feature set, assigning weights to each feature through a preset attention mechanism weighted model, automatically learning the importance of features in the fusion process, and obtaining a weighted fusion feature vector; Feature generation processing is performed based on the weighted fusion feature vector, and third-order tensors representing the weighted electroretinogram features, image features and fusion weights are constructed through a multimodal tensor decomposition and fusion algorithm. Tensor decomposition processing is performed, and feature fusion is performed on the basis of retaining the dynamic change rules between multimodal features to obtain fused features.
2. The method for automatically identifying children's fundus diseases based on computer vision according to claim 1, characterized in that , performing feature extraction processing based on the first information, and obtaining an electroretinogram feature set by extracting A wave amplitude, B wave amplitude, delay time and waveform feature parameters, including: Performing data extraction and processing based on the first information, detecting local minima of the electroretinogram signal using a second-order derivative method, identifying the peak positions of the A wave and the B wave and calculating their amplitudes to obtain an A wave amplitude feature and a B wave amplitude feature; Performing delay time calculation processing based on the first information, and obtaining delay time data by calculating the time delay between wave A and wave B; Performing waveform feature parameter extraction processing based on the first information, performing time-frequency analysis using wavelet transform to obtain a wavelet transform result, and extracting shape features, symmetry, and frequency components from the wavelet transform result to generate waveform feature parameters; The A-wave amplitude feature, the B-wave amplitude feature, the delay time data and the waveform feature parameters are integrated to obtain an electroretinogram feature set.
3. The method for automatically identifying children's fundus diseases based on computer vision according to claim 1, characterized in that ,According to the second information, multi-scale structural analysis and feature enhancement processing are performed, by extracting the blood vessel diameter, branch point number and blood vessel sparsity features and performing feature enhancement, an image feature set is obtained, including: performing multi-scale structural analysis based on the second information, performing layered processing on the retinal image, gradually applying Gaussian blur and downsampling, and obtaining multi-level image representations at different resolutions; performing vascular feature extraction processing based on the multi-level image representation, identifying the main structure of the blood vessels by binarizing and skeletonizing the image, analyzing the node connections in the skeleton graph using a graph analysis algorithm, measuring the diameter of the blood vessels, and identifying branch points with a connection number greater than or equal to 2, thereby obtaining vascular diameter features and branch point number features; Performing local density estimation processing according to the multi-level image representation, and obtaining a vascular sparsity feature by calculating the number of vascular pixels in each region; Feature fusion processing is performed based on the blood vessel diameter feature, the branch point number feature, and the blood vessel sparsity feature to obtain an image feature set.
4. The method for automatically identifying children's fundus diseases based on computer vision according to claim 1, characterized in that , performing dimensionality reduction processing according to the fusion features to obtain a comprehensive feature representation, and constructing a topological perception graph structure based on the comprehensive feature representation, including: Performing graph construction processing based on the fusion features, calculating feature adjacency based on a k-nearest neighbor graph, using different features as graph nodes, and determining the connection relationship between nodes by feature similarity to obtain an initial feature graph; Performing graph structure generation processing based on the initial feature graph, analyzing the topological features of nodes and edges in the graph structure at different scales by applying a persistent homology method, identifying stable structures in the features and removing noise and unstable nodes, thereby obtaining a topologically aware graph structure; Hierarchical organization processing is performed according to the topological perception map structure, and a hierarchical structure is generated by hierarchically dividing the vascular features, electroretinogram signal features, and morphological features. The vascular features and electroretinogram features are organized at different hierarchical levels according to changes in the growth characteristics of the children to obtain a topological perception map structure.
5. The method for automatically identifying children's fundus diseases based on computer vision according to claim 1, characterized in that ,According to the topological perception graph structure, pattern matching and adaptive threshold segmentation processing are performed, and vascular pattern information is obtained by identifying the change pattern of children's retinal blood vessels at different growth stages, including: Based on a preset graph convolutional network, local information of different feature nodes in the topology-aware graph structure is aggregated. By convolving the vascular feature nodes layer by layer, local variation characteristics of vascular features in their neighborhoods are captured to obtain feature vector representations of all nodes in the graph. performing pattern matching processing based on the feature vector representation, calculating the similarity between the feature vectors at different time points using a dynamic time warping algorithm, and comparing the vascular features collected from the same child at different times to identify vascular development trends or abnormal patterns and obtain a change pattern; Performing local feature adaptive segmentation processing based on the change pattern, estimating the probability distribution of vascular features by using a Gaussian mixture model, and adaptively selecting thresholds to segment different types of vascular regions based on the morphological differences of retinal blood vessels in children of different age groups to obtain segmentation results; A topological feature evaluation is performed based on the segmentation results, and vascular pattern information is obtained by calculating the vascular branching complexity, the number of loops, and the connectivity of the blood vessels.
6. The method for automatically identifying children's fundus diseases based on computer vision according to claim 1, characterized in that , performing diagnostic analysis based on the comprehensive feature representation and the vascular pattern information to obtain a diagnostic result, including: Inputting the comprehensive feature representation and the vascular pattern information into a preset extreme gradient boosting tree model, using a tree structure to gradually split the data and calculate the contribution score of each feature to the classification result, thereby obtaining a feature importance ranking result; Performing classification processing based on the importance ranking results, inputting features with feature importance scores greater than a threshold into a preset naive Bayes model, assuming conditional independence between features and calculating the probability of each disease, to obtain classification results for childhood retinal diseases; performing abnormal feature detection processing based on the comprehensive feature representation and the vascular pattern information, calculating the Mahalanobis distance between each sample and the center of the overall feature distribution, and detecting abnormal distribution of blood vessels or abnormal response of the electroretinogram to obtain abnormal feature information; Inference processing is performed based on the classification results and the abnormal feature information, by converting the degree of vascular rarefaction, the amplitude abnormality of the electroretinogram and the delay feature of the electroretinogram into fuzzy sets, and performing comprehensive analysis through preset fuzzy inference rules to obtain a diagnosis result.
7. A computer vision-based automatic identification system for children's fundus diseases, characterized by: include: An acquisition module is used to acquire first information and second information, wherein the first information is the electroretinogram of the child acquired by a non-invasive device, and the second information is the retinal image of the child taken by a fundus machine. An extraction module performs feature extraction processing based on the first information to obtain an electroretinogram feature set by extracting A-wave amplitude, B-wave amplitude, delay time and waveform characteristic parameters; an enhancement module, configured to perform multi-scale structural analysis and feature enhancement processing based on the second information, extracting features such as blood vessel diameter, number of branch points, and blood vessel sparsity and performing feature enhancement to obtain an image feature set; a fusion module, configured to perform feature fusion processing based on the electroretinogram feature set and the image feature set, and obtain a fusion feature by weightedly concatenating the two types of features into a comprehensive feature vector; a construction module for performing dimensionality reduction processing on the fused features to obtain a comprehensive feature representation, and constructing a topological perception map structure based on the comprehensive feature representation, wherein the topological perception map structure includes a vascular feature map, an electroretinogram signal feature map, a retinal morphology feature map, and a corresponding hierarchical feature organization; a matching module for performing pattern matching and adaptive threshold segmentation processing based on the topological perception graph structure, and obtaining vascular pattern information by identifying the changing patterns of retinal blood vessels in children at different growth stages; a diagnosis module, performing a diagnosis analysis based on the comprehensive feature representation and the vascular pattern information to obtain a diagnosis result; The feature fusion processing is performed based on the electroretinogram feature set and the image feature set, and the fusion feature is obtained by weightedly splicing the two types of features into a comprehensive feature vector, including: Performing feature correlation analysis on the electroretinogram feature set and the image feature set, and obtaining a feature correlation matrix by evaluating the correlation scores between the electroretinogram features and the image features using a maximum information coefficient algorithm; Performing feature denoising and selection processing based on the feature correlation matrix, performing sparse constraints on short-term mutation signals in the children's electroretinogram through sparse principal component analysis, and performing weighted sparseness on the importance of features in the children's retinal images to obtain a streamlined feature set; Performing feature weighted fusion processing based on the streamlined feature set, assigning weights to each feature through a preset attention mechanism weighted model, automatically learning the importance of features in the fusion process, and obtaining a weighted fusion feature vector; Feature generation processing is performed based on the weighted fusion feature vector, and third-order tensors representing the weighted electroretinogram features, image features and fusion weights are constructed through a multimodal tensor decomposition and fusion algorithm. Tensor decomposition processing is performed, and feature fusion is performed on the basis of retaining the dynamic change rules between multimodal features to obtain fused features.
8. The computer vision-based automatic identification system for children's fundus diseases according to claim 7 is characterized in that , the extraction module includes: a first extraction unit, configured to perform data extraction processing based on the first information, detect a local minimum of the electroretinogram signal using a second-order derivative method, identify the peak positions of the A wave and the B wave, and calculate their amplitudes to obtain an A wave amplitude feature and a B wave amplitude feature; a first calculation unit, configured to perform delay time calculation processing based on the first information, and obtain delay time data by calculating the time delay between wave A and wave B; a second extraction unit, configured to perform waveform feature parameter extraction processing based on the first information, perform time-frequency analysis using wavelet transform to obtain a wavelet transform result, and extract shape features, symmetry, and frequency components from the wavelet transform result to generate waveform feature parameters; The first integration unit is used to integrate the A-wave amplitude feature, the B-wave amplitude feature, the delay time data and the waveform feature parameter to obtain an electroretinogram feature set.
9. The computer vision-based automatic identification system for children's fundus diseases according to claim 7 is characterized in that , the enhancement module includes: a first analysis unit, configured to perform multi-scale structural analysis processing based on the second information, by layering the retinal image and gradually applying Gaussian blur and downsampling to obtain multi-level image representations at different resolutions; a third extraction unit, configured to perform vascular feature extraction processing based on the multi-level image representation, identify the main structure of the blood vessels by binarizing and skeletonizing the image, analyze the node connections in the skeleton graph using a graph analysis algorithm, measure the diameter of the blood vessels, and identify branch points with a connection number greater than or equal to 2, thereby obtaining a vascular diameter feature and a branch point number feature; a second computing unit, configured to perform local density estimation processing according to the multi-level image representation, and obtain a vascular sparsity feature by calculating the number of vascular pixels in each region; The first fusion unit is configured to perform feature fusion processing according to the blood vessel diameter feature, the branch point number feature, and the blood vessel sparsity feature to obtain an image feature set.
Citation Information
Patent Citations
Eye fundus image retinal vessel segmentation method based on graph convolutional neural network
CN116862928A
Method for constructing neonatal retinopathy prediction model based on image technology
CN117838041A