Multi-modal intelligent diagnosis model generation method and device for myopic maculopathy

Through the multimodal intelligent diagnostic model generation method, a myopic macular degeneration diagnostic model is constructed using multimodal data, which solves the problem of low efficiency of manual diagnosis in existing technologies, realizes intelligent and efficient lesion diagnosis, and improves the accuracy and reliability of diagnosis.

CN120809163APending Publication Date: 2025-10-17ANHUI PROVINCIAL HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510959251.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the existing technology, the diagnosis of myopic maculopathy mainly relies on manual diagnosis, which has the problem of low diagnostic efficiency and makes it difficult to achieve intelligent and efficient diagnosis.

Method used

A multimodal intelligent diagnostic model generation method is adopted to obtain multimodal data of target users, including medical records, fundus photographs, OCT data and OCTA data, and perform data preprocessing, analysis and feature extraction to build a multimodal intelligent diagnostic model and realize lesion prediction.

Benefits of technology

It improves the diagnostic intelligence and accuracy of myopic macular degeneration, enhances the reliability and efficiency of diagnosis, adapts to complex diagnostic scenarios, and reduces diagnostic costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809163A_ABST
    Figure CN120809163A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal intelligent diagnosis model generation method and device for myopic maculopathy, and the method comprises the steps: obtaining multi-modal data, carrying out the data preprocessing operation of each piece of obtained multi-modal data, obtaining target multi-modal data, and generating a target multi-modal database based on all the target multi-modal data; performing data analysis operation on each piece of target multi-modal data to obtain a data analysis result; and according to the data analysis result of all the target multi-modal data, constructing a target data chart, extracting target feature information from the target data chart, and based on all the target feature information, generating target feature mapping so as to construct and obtain a multi-modal intelligent diagnosis model. Visibly, the multi-mode intelligent diagnosis model of the myopic maculopathy can be intelligently generated by implementing the method and the system, so that the intelligence of diagnosing the myopic maculopathy through the model is improved, and the accuracy and the reliability of diagnosing the myopic maculopathy are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of model generation, in particular to a myopic maculopathy multi-modal intelligent diagnosis model generation method and device. BACKGROUND

[0002] At present, myopia has shown a trend of outbreak similar to an epidemic, and its incidence has been increasing year by year. According to the latest epidemiological survey, the global prevalence rate of myopia in all age groups is about 30%, of which 163 million people (accounting for 2.7% of the total population) suffer from high myopia. It is estimated that by 2050, the global myopia population will reach 4.8 billion (accounting for 49.8% of the global population). In addition to harming eye health, myopia also has a significant impact on public health and social and economic well-being. According to the progression and pathological changes of myopia, myopia can be divided into simple myopia and pathological myopia (PM). Compared with simple myopia, PM shows more serious clinical symptoms and has become the second leading cause of blindness in China.

[0003] With the rapid development of data scale, algorithm model and computing performance, artificial intelligence (AI) technology based on deep learning has been applied in many fields such as computer vision, natural language processing and speech analysis, and has played an important role in medical image analysis and processing. However, most of the current diagnosis of myopic maculopathy is still based on human experience, which not only has the problem of low diagnosis efficiency. Therefore, it is particularly important to provide a new diagnosis model for myopic maculopathy to improve the intelligence and efficiency of diagnosis. SUMMARY

[0004] The present application provides a myopic maculopathy multi-modal intelligent diagnosis model generation method and device, which can intelligently generate a myopic maculopathy multi-modal intelligent diagnosis model, and is beneficial to improve the intelligence and efficiency of diagnosing myopic maculopathy, and further beneficial to improve the accuracy and reliability of diagnosing myopic maculopathy.

[0005] In order to solve the above technical problems, the first aspect of the present application discloses a myopic maculopathy multi-modal intelligent diagnosis model generation method, which comprises:

[0006] Obtaining user data of a target user; wherein the user data at least includes multi-modal data of the target user;

[0007] Extracting target data from the user data, and generating a data analysis result based on the target data and a pre-determined data analysis model;

[0008] performing a data matching operation on the data analysis result and a predetermined target data result to obtain a data matching result, and determining data feature information based on the data matching result;

[0009] generating a lesion prediction result according to the data feature information and a pre-constructed multi-modal intelligent diagnosis model.

[0010] As an optional implementation form, before the step of generating the lesion prediction result according to the data feature information and the pre-constructed multi-modal intelligent diagnosis model, the method further includes:

[0011] performing a data processing operation on the data feature information to obtain image data feature information and text data feature information, wherein each image data feature information has corresponding text data feature information;

[0012] For each image data feature information, performing a data conversion operation on the image data feature information to obtain a data conversion result of the image data feature information, and generating target data feature information of the image data feature information according to the data conversion result of the image data feature information and the text data feature information corresponding to the image data feature information;

[0013] The step of generating the lesion prediction result according to the data feature information and the pre-constructed multi-modal intelligent diagnosis model includes:

[0014] generating the lesion prediction result according to all the target data feature information and the pre-constructed multi-modal intelligent diagnosis model.

[0015] As an optional implementation form, after the step of generating the lesion prediction result according to the data feature information and the pre-constructed multi-modal intelligent diagnosis model, the method further includes:

[0016] performing a verification operation on the lesion prediction result to obtain a prediction verification result of the lesion prediction result, and performing a model verification operation on the pre-determined data analysis model based on the prediction verification result to obtain a model verification result;

[0017] determining whether the model verification result meets a preset model running condition;

[0018] When it is determined that the model verification result does not meet the preset model running condition, performing a model optimization operation on the pre-determined data analysis model based on the model verification result and the preset model running condition to update the data analysis model.

[0019] As an optional implementation, in the first aspect of the present application, the generating of the data analysis result based on the target data and the pre-determined data analysis model comprises:

[0020] inputting the target data into the pre-determined data analysis model to perform data decoupling operation on the target data by the pre-determined data analysis model to obtain a data decoupling result corresponding to the target data, wherein the data decoupling result comprises at least data semantic features corresponding to the target data;

[0021] generating the data analysis result corresponding to the target data based on the data decoupling result corresponding to the target data.

[0022] As an optional implementation, in the first aspect of the present application, the performing of the data matching operation on the data analysis result and the pre-determined target data result to obtain a data matching result comprises:

[0023] extracting first key information in the data analysis result and second key information in the pre-determined target data result, performing information matching operation on the first key information and the second key information to obtain an information matching result;

[0024] generating data distribution information corresponding to the data analysis result based on the information matching result, and generating a data matching result according to the data distribution information.

[0025] As an optional implementation, in the first aspect of the present application, the generating of the lesion prediction result according to all the target data feature information and the pre-constructed multi-modal intelligent diagnosis model comprises:

[0026] inputting all the target data feature information into the pre-constructed multi-modal intelligent diagnosis model to perform feature extraction operation on all the target data feature information by a feature extraction layer in the multi-modal intelligent diagnosis model to obtain key feature data, perform feature fusion operation on all the key feature data to obtain a feature fusion result, and generate target key feature data based on the feature fusion result and a full connection layer in the multi-modal intelligent diagnosis model;

[0027] generating the lesion prediction result based on the target key feature data.

[0028] As an optional implementation, in the first aspect of the present application, for each image data feature information, the generating of the target data feature information of the image data feature information according to the data conversion result of the image data feature information and the literal data feature information corresponding to the image data feature information comprises:

[0029] Determining data distribution parameters according to a data conversion result of the image data feature information, wherein the data distribution parameters include prior distribution parameters and conditional distribution parameters;

[0030] generating target data feature information of the image data feature information based on the data distribution parameter, the data conversion result of the image data feature information, and the text data feature information corresponding to the image data feature information;

[0031] The step of generating target data feature information of the image data feature information based on the data distribution parameter, the data conversion result of the image data feature information, and the text data feature information corresponding to the image data feature information includes:

[0032] Determining distribution characteristic information of the text data characteristic information corresponding to the image data characteristic information based on the prior distribution parameter, and determining the text data characteristic information corresponding to the image data characteristic information and hidden characteristic information of the image data characteristic information based on the conditional distribution parameter;

[0033] According to the distribution feature information and the hidden feature information, comprehensive data feature information of the image data feature information is determined, and based on the comprehensive data feature information of the image data feature information, target data feature information of the image data feature information is generated.

[0034] A second aspect of the present invention discloses a device for generating a multimodal intelligent diagnostic model for myopic maculopathy, the device comprising:

[0035] An acquisition module, configured to acquire user data of a target user; wherein the user data at least includes multimodal data of the target user;

[0036] An extraction module, configured to extract target data from the user data;

[0037] A generating module, configured to generate a data analysis result based on the target data and a predetermined data analysis model;

[0038] A matching module is used to perform a data matching operation on the data analysis result and a predetermined target data result to obtain a data matching result;

[0039] A determination module, configured to determine data feature information based on the data matching result;

[0040] The generation module is further used to generate a lesion prediction result based on the data feature information and a pre-built multimodal intelligent diagnosis model.

[0041] As an optional implementation, in the second aspect of the present application, the device further comprises:

[0042] a processing module, configured to perform a data processing operation on the data feature information to obtain image data feature information and text data feature information before the generating module generates the lesion prediction result according to the data feature information, wherein each image data feature information has corresponding text data feature information;

[0043] a conversion module, configured to perform a data conversion operation on each image data feature information to obtain a data conversion result of the image data feature information;

[0044] the generating module is further configured to generate target data feature information of the image data feature information according to the data conversion result of the image data feature information and the text data feature information corresponding to the image data feature information;

[0045] The specific manner in which the generating module generates the lesion prediction result according to the data feature information and the pre-constructed multi-modal intelligent diagnosis model comprises:

[0046] generating a lesion prediction result according to all the target data feature information and the pre-constructed multi-modal intelligent diagnosis model.

[0047] As an optional implementation, in the second aspect of the present application, the device further comprises:

[0048] a verification module, configured to perform a verification operation on the lesion prediction result after the generating module generates the lesion prediction result according to the data feature information and the pre-constructed multi-modal intelligent diagnosis model to obtain a prediction verification result of the lesion prediction result, and perform a model verification operation on the pre-determined data analysis model based on the prediction verification result to obtain a model verification result;

[0049] a judgment module, configured to judge whether the model verification result meets a preset model running condition;

[0050] an updating module, configured to perform a model optimization operation on the pre-determined data analysis model based on the model verification result and the preset model running condition to update the data analysis model when the judgment module judges that the model verification result does not meet the preset model running condition.

[0051] As an optional implementation, in the second aspect of the present application, the specific manner in which the generating module generates the data analysis result based on the target data and the pre-determined data analysis model comprises:

[0052] inputting the target data into a pre-determined data analysis model to perform a data decoupling operation on the target data through the pre-determined data analysis model, to obtain a data decoupling result corresponding to the target data, wherein the data decoupling result at least includes a data semantic feature corresponding to the target data;

[0053] generating a data analysis result corresponding to the target data based on the data decoupling result corresponding to the target data.

[0054] As an optional implementation, in the second aspect of the present application, the specific manner in which the matching module performs a data matching operation on the data analysis result and the pre-determined target data result to obtain a data matching result includes:

[0055] extracting first key information in the data analysis result, and extracting second key information in the pre-determined target data result, performing an information matching operation on the first key information and the second key information to obtain an information matching result;

[0056] generating data distribution information corresponding to the data analysis result based on the information matching result, and generating a data matching result according to the data distribution information.

[0057] As an optional implementation, in the second aspect of the present application, the specific manner in which the generating module generates a lesion prediction result according to all the target data feature information and a pre-constructed multi-modal intelligent diagnosis model includes:

[0058] inputting all the target data feature information into a pre-constructed multi-modal intelligent diagnosis model to perform a feature extraction operation on all the target data feature information through a feature extraction layer in the multi-modal intelligent diagnosis model to obtain key feature data, perform a feature fusion operation on all the key feature data to obtain a feature fusion result, and generate target key feature data based on the feature fusion result and a full connection layer in the multi-modal intelligent diagnosis model;

[0059] generating a lesion prediction result based on the target key feature data.

[0060] As an optional implementation, in the second aspect of the present application, for each image data feature information, the specific manner in which the generating module generates target data feature information of the image data feature information according to a data conversion result of the image data feature information and text data feature information corresponding to the image data feature information includes:

[0061] According to the data conversion result of the image data feature information, determine data distribution parameters, wherein the data distribution parameters include prior distribution parameters and conditional distribution parameters;

[0062] Based on the data distribution parameters, the data conversion result of the image data feature information and the literal data feature information corresponding to the image data feature information, generate the target data feature information of the image data feature information;

[0063] Wherein, the generation module based on the data distribution parameters, the data conversion result of the image data feature information and the literal data feature information corresponding to the image data feature information, generate the target data feature information of the image data feature information, the specific way includes:

[0064] According to the prior distribution parameters, determine the distribution feature information of the literal data feature information corresponding to the image data feature information, and according to the conditional distribution parameters, determine the hidden feature information of the image data feature information and the literal data feature information corresponding to the image data feature information;

[0065] According to the distribution feature information and the hidden feature information, determine the comprehensive data feature information of the image data feature information, and generate the target data feature information of the image data feature information based on the comprehensive data feature information of the image data feature information.

[0066] The third aspect of the present application discloses another kind of myopia maculopathy multi-modal intelligent diagnosis model generation device, the device includes:

[0067] Memory storing executable program code;

[0068] The processor is coupled with the memory;

[0069] The processor calls the executable program code stored in the memory, and executes the myopia maculopathy multi-modal intelligent diagnosis model generation method disclosed in the first aspect of the present application.

[0070] The fourth aspect of the present application discloses a computer storage medium, the computer storage medium stores computer instructions, the computer instructions are called to execute the myopia maculopathy multi-modal intelligent diagnosis model generation method disclosed in the first aspect of the present application.

[0071] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0072] In the embodiment of the present application, multi-modal data is acquired, data preprocessing operations are performed on each acquired multi-modal data to obtain target multi-modal data, a target multi-modal database is generated based on all target multi-modal data, data analysis operations are performed on each target multi-modal data to obtain data analysis results, target data charts are constructed according to the data analysis results of all target multi-modal data, target feature information is extracted from the target data charts, and a target feature mapping is generated based on all target feature information to construct a multi-modal intelligent diagnosis model. It can be seen that the present application can intelligently generate a multi-modal intelligent diagnosis model for myopic maculopathy, which is beneficial to improve the intelligence of diagnosing myopic maculopathy through the model, and further beneficial to improve the accuracy and reliability of diagnosing myopic maculopathy. BRIEF DESCRIPTION OF DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0074] Figure 1 is a flowchart of a myopic maculopathy multi-modal intelligent diagnosis model generation method disclosed by the embodiment of the present application;

[0075] Figure 2 is a flowchart of another myopic maculopathy multi-modal intelligent diagnosis model generation method disclosed by the embodiment of the present application;

[0076] Figure 3 is a structural diagram of a myopic maculopathy multi-modal intelligent diagnosis model generation device disclosed by the embodiment of the present application;

[0077] Figure 4 is a structural diagram of another myopic maculopathy multi-modal intelligent diagnosis model generation device disclosed by the embodiment of the present application;

[0078] Figure 5 is a structural diagram of another myopic maculopathy multi-modal intelligent diagnosis model generation device disclosed by the embodiment of the present application. DETAILED DESCRIPTION

[0079] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0080] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or end including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product, or end.

[0081] In this paper, the term "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0082] The present application discloses a kind of myopia maculopathy multi-modal intelligent diagnosis model generation method and device, can intelligently generate the multi-modal intelligent diagnosis model of myopia maculopathy, it is favorable to improve the intelligence of myopia maculopathy by model and diagnose, in turn it is also favorable to improve the precision and reliability of myopia maculopathy and diagnose. The following are described in detail respectively.

[0083] Embodiment one

[0084] Please refer to Figure 1 , Figure 1 It is a kind of myopia maculopathy multi-modal intelligent diagnosis model generation method flow diagram disclosed in the embodiment of the present application. Among them, Figure 1 The described myopia maculopathy multi-modal intelligent diagnosis model generation method can be applied to myopia maculopathy multi-modal intelligent diagnosis model generation device, wherein the myopia maculopathy multi-modal intelligent diagnosis model generation device can be integrated in cloud server or local server, and the embodiment of the present application is not limited. As shown in Figure 1 The myopia maculopathy multi-modal intelligent diagnosis model generation method can include the following operations:

[0085] 101. Acquire multimodal data, perform data preprocessing on each acquired multimodal data to obtain target multimodal data, and generate a target multimodal database based on all target multimodal data.

[0086] In an embodiment of the present invention, optionally, the multimodal data can be acquired in real time, or can be acquired periodically according to a preset time period, or can be generated when a multimodal database or an intelligent diagnostic model needs to be generated, and the embodiment of the present invention does not make specific limitations.

[0087] In an embodiment of the present invention, optionally, the multimodal data may include multimodal data corresponding to several myopic patients, wherein the multimodal data may include one or more of the medical records, fundus photographs, OCT data, OCTA data, and optometry examination data of each myopic patient, which is not specifically limited in the embodiment of the present invention. Among them, OCT (Optical Coherence Tomography) is a high-resolution biological imaging technology that can detect the structure and physiological function inside the living eye in a non-invasive and contactless manner. Furthermore, in ophthalmology, OCT technology is widely used in retinal imaging, which can clearly observe the structure and pathological conditions of each layer of the human retina; and OCTA (Optical Coherence Tomography Angiography) is a non-invasive ophthalmic imaging technology that allows the three-dimensional structure of retinal blood vessels to be presented at a micron-level resolution, thereby observing the blood flow of the retina without the use of contrast agents. OCTA can provide more accurate images and data and is an extremely valuable imaging tool for ophthalmology.

[0088] In the embodiment of the present invention, optionally, the target multimodal database includes at least all target multimodal data.

[0089] 102. For each target multimodal data included in the target multimodal database, perform a data analysis operation on the target multimodal data to obtain a data analysis result of the target multimodal data.

[0090] In an embodiment of the present invention, each target multimodal data may optionally have a corresponding data analysis result, wherein the data analysis result of each target multimodal data may include at least an analysis result corresponding to the text data and an analysis result corresponding to the image data.

[0091] 103. Based on the data analysis results of all target multimodal data, a target data chart is constructed, target feature information is extracted from the target data chart, and a target feature map is generated based on all target feature information.

[0092] In the embodiment of the application, optionally, the constructing the target data chart according to the data analysis result of all target multi-modal data can comprise: constructing EHR data modeling according to the data analysis result of all target multi-modal data, and constructing the target data chart based on the EHR data modeling, wherein the target data chart at least comprises historical information of EHR.

[0093] In the embodiment of the application, optionally, the electronic health record (EHR) is an important medical information tool, which aims to record, store, manage and share personal health information in an electronic way; further, the EHR is an electronic history record with archival value formed directly by people in health-related activities. It is a lifelong personal health record stored in a computer system, providing services for individuals, and having security and privacy performance. The EHR data modeling can comprise EHR personal health information corresponding to each target patient, and the historical information of EHR comprises historical health information corresponding to each target patient.

[0094] 104. Based on all target feature mapping, a multi-modal intelligent diagnosis model is constructed.

[0095] In the embodiment of the application, optionally, the multi-modal intelligent diagnosis model can be used for intelligent diagnosis of myopic maculopathy of each patient.

[0096] It can be seen that, by implementing the method and the device, the multi-modal intelligent diagnosis model can be constructed based on the multi-modal data of each target patient, and the multi-modal intelligent diagnosis model can be used for intelligent diagnosis of myopic maculopathy of each patient. Figure 1The described myopic maculopathy multi-modal intelligent diagnosis model generation method can obtain multi-modal data and perform data preprocessing operations to obtain target multi-modal data, generate a target multi-modal database based on all target multi-modal data, perform data analysis operations on each target multi-modal data included in the target multi-modal database to obtain data analysis results, construct a target data chart according to the data analysis results of all target multi-modal data, and extract target feature information from the target data chart to generate a target feature mapping. Based on all target feature mappings, a multi-modal intelligent diagnosis model is constructed. By obtaining multi-modal data, more comprehensive and rich information can be provided, and by performing data preprocessing operations, the accuracy and consistency of the data can be improved, which can bring convenience for subsequent data analysis and model training. The target feature information obtained through data analysis and icon extraction can accurately reflect the internal characteristics and structure of the data, providing strong support for constructing an efficient and accurate multi-modal intelligent diagnosis model. The multi-modal intelligent diagnosis model constructed based on the target feature mapping can integrate multi-modal information from multiple aspects to realize cross-modal information fusion and complementarity, which is conducive to improving the accuracy and reliability of the subsequent model training and obtaining an intelligent diagnosis model. Compared with traditional single-modal diagnosis methods, the multi-modal intelligent diagnosis model has stronger adaptability and generalization ability, can cope with more complex and variable diagnosis scenarios, and can improve the accuracy and efficiency of diagnosis, reduce the cost of diagnosis, and bring better experience and benefits to users. In turn, it is also conducive to improving the diagnosis intelligence and efficiency of myopic maculopathy, and further improving the accuracy and reliability of diagnosing myopic maculopathy through an intelligent diagnosis model.

[0097] Embodiment two

[0098] Please refer to Figure 2 , Figure 2 is another flowchart of the myopic maculopathy multi-modal intelligent diagnosis model generation method disclosed in the embodiments of the present application. Among them, Figure 2 The described myopic maculopathy multi-modal intelligent diagnosis model generation method can be applied to a myopic maculopathy multi-modal intelligent diagnosis model generation device, wherein the myopic maculopathy multi-modal intelligent diagnosis model generation device can be integrated in a cloud server or a local server, and the embodiments of the present application are not limited. As shown in Figure 2 The described myopic maculopathy multi-modal intelligent diagnosis model generation method can be applied to a myopic maculopathy multi-modal intelligent diagnosis model generation device, wherein the myopic maculopathy multi-modal intelligent diagnosis model generation device can be integrated in a cloud server or a local server, and the embodiments of the present application are not limited. As shown in

[0099] 201, obtain multi-modal data, perform data preprocessing operations on each multi-modal data obtained, obtain target multi-modal data, and generate a target multi-modal database based on all target multi-modal data.

[0100] 202. For each target multimodal data included in the target multimodal database, perform a data analysis operation on the target multimodal data to obtain a data analysis result of the target multimodal data.

[0101] 203. Based on the data analysis results of all target multimodal data, a target data chart is constructed, target feature information is extracted from the target data chart, and a target feature map is generated based on all target feature information.

[0102] In the embodiment of the present invention, for the detailed description of steps 201 to 203 , please refer to the other descriptions of steps 101 to 103 in the first embodiment, and the embodiment of the present invention will not be repeated.

[0103] 204. Perform a convolution operation on all target feature maps to obtain feature operation results, and determine the output channel and the average pooling layer based on the feature operation results.

[0104] In an embodiment of the present invention, the convolution operation optionally includes a two-dimensional convolution operation. The two-dimensional convolution operation is an important operation in computer vision and image processing, and is typically used for tasks such as feature extraction, image filtering, and edge detection. The above-mentioned convolution operation performed on all target feature maps to obtain a feature operation result may be performed on a time axis to extract time information to obtain the feature operation result.

[0105] In an embodiment of the present invention, optionally, the above-mentioned determination of the output channel and the average pooling layer based on the feature operation results can be to generate a higher-level patient feature map through the feature operation results, and determine the channel attention layer corresponding to each ST-GCN unit based on each ST-GCN unit included in the patient feature map, and determine the global average pooling layer through all channel attention layers corresponding to all ST-GCN units. For example, each ST-GCN unit is followed by a channel attention layer to help the model focus on channels with more meaningful functions. The first two, middle two, and last two ST-GCN units have 64, 128, and 256 output channels, respectively, and each channel is followed by a global average pooling layer.

[0106] 205. Generate an output feature map based on all output channels, the average pooling layer, and the feature operation results.

[0107] In the embodiment of the present invention, optionally, generating an output feature map based on all output channels, the average pooling layer, and the feature operation results may include:

[0108] According to the feature operation result and all output channels, first feature data is generated, and according to the feature operation result and the average pooling layer, second feature data is generated; a feature fusion operation is performed on the first feature data and the second feature data to generate an output feature map.

[0109] In the embodiment of the application, for example, the outputs of the pooling layers are connected together to realize feature fusion and generate a final patient feature map. Further, the outputs of the first two pooling operations are not transmitted to the following units.

[0110] 206, based on the output feature map and the predetermined full connection layer, a multi-modal intelligent diagnosis model is constructed.

[0111] In the embodiment of the application, the output feature map is connected to the full connection layer.

[0112] It can be seen that the implementation Figure 2 The described multi-modal intelligent diagnosis model generation method for myopic maculopathy can perform convolution operation on all target feature maps to obtain feature operation results and determine output channels and average pooling layers, generate an output feature map based on all output channels, average pooling layers and feature operation results, and construct a multi-modal intelligent diagnosis model based on the output feature map and the full connection layer. The method can further extract and enhance key features in the data by performing convolution operation on the target feature map, can make the model capture the local spatial relationship of the data, and can improve the expression ability of the features. At the same time, the output feature map of the convolution operation combines the information of different channels, realizes the integration of cross-modal features, and can use the complementary information in the multi-modal data to improve the accuracy and reliability of the subsequently constructed multi-modal intelligent diagnosis model. By combining convolution operation, average pooling layer and full connection layer, the multi-modal intelligent diagnosis model can fully utilize the advantages of multi-modal data, improve the accuracy and efficiency of diagnosis, so that the model has strong generalization ability and can adapt to the needs of different fields and scenes. Therefore, the effectiveness of feature extraction and enhancement, the flexibility of feature integration, and the optimization of output channels and average pooling layers can optimize the model structure, improve the accuracy and reliability of the constructed multi-modal intelligent diagnosis model, improve the model performance of the multi-modal intelligent diagnosis model, and further improve the diagnosis intelligence and efficiency for myopic maculopathy. Further, the accuracy and reliability of diagnosing myopic maculopathy by the intelligent diagnosis model are further improved.

[0113] In an optional embodiment, data preprocessing is performed on each acquired multi-modal data to obtain target multi-modal data, including:

[0114] For each multi-modal data obtained, a data target processing operation is performed on the multi-modal data to obtain pre-processed multi-modal data corresponding to the multi-modal data, wherein the data target processing operation comprises a data normalization operation and / or a data enhancement operation.

[0115] For each multi-modal data, it is determined whether a data format corresponding to the pre-processed multi-modal data matches a preset reference data format according to the pre-processed multi-modal data corresponding to the multi-modal data.

[0116] When it is determined that the data format corresponding to the pre-processed multi-modal data does not match the preset reference data format, a data conversion operation is performed on the pre-processed multi-modal data to obtain target pre-processed data corresponding to the pre-processed multi-modal data; when it is determined that the data format corresponding to the pre-processed multi-modal data matches the preset reference data format, the pre-processed multi-modal data is determined as the target pre-processed data.

[0117] Based on all the target pre-processed data, target multi-modal data is generated.

[0118] In the optional embodiment, optionally, when the data target processing operation comprises the data normalization operation, the above-mentioned, for each multi-modal data obtained, performing a data target processing operation on the multi-modal data to obtain pre-processed multi-modal data corresponding to the multi-modal data, can comprise:

[0119] For each multi-modal data obtained, a data normalization operation is performed on a data format corresponding to the multi-modal data to obtain pre-processed multi-modal data corresponding to the multi-modal data; wherein the data format of the pre-processed multi-modal data corresponding to the multi-modal data matches a preset data format.

[0120] In the optional embodiment, optionally, for example, in order to unify the format and size of the same modality data from various shooting devices, it is necessary to pre-process the image so that the data format and data size of the image data included in all multi-modal data meet the requirements.

[0121] In the optional embodiment, optionally, when the data target processing operation comprises the data enhancement operation, the above-mentioned, for each multi-modal data obtained, performing a data target processing operation on the multi-modal data to obtain pre-processed multi-modal data corresponding to the multi-modal data, can comprise:

[0122] For each multi-modal data obtained, a data enhancement operation is performed on the multi-modal data to obtain pre-processed multi-modal data corresponding to the multi-modal data; wherein the pre-processed multi-modal data corresponding to the multi-modal data can comprise augmented data corresponding to the multi-modal data.

[0123] In the optional embodiment, optionally, for example, data enhancement can bring regularization effect, which can alleviate the problem of overfitting caused by too few images, and make the model more suitable for data diversity, and will not cause classification result error due to shooting angle or color deviation. The data enhancement algorithm to be adopted includes cropping area, random rotation, random inversion, random contrast change, etc.

[0124] In the optional embodiment, further optionally, before performing the data target processing operation on each multi-modal data to obtain the pre-processed multi-modal data corresponding to the multi-modal data, the above method can further include:

[0125] performing a data cleaning operation on all multi-modal data to update each multi-modal data;

[0126] The data cleaning operation includes one or more of a data deduplication operation, a data deletion operation, and a data conversion operation.

[0127] In the optional embodiment, optionally, for example, the data cleaning operation can include deleting sensitive information of the target patient, for example, removing data with obvious missing patient disease information and obviously poor image quality; and the removal process can be through first determining a standard scanning mode for fundus color images, OCT images and the like, and converting other scans to the standard mode through image post-processing; secondly, for optometry examinations with different measurement or recording methods, selecting authoritative formulas to convert various formats, processing according to uniform processing of OCT image sizes, and re-standardizing patient medical records, and unifying patient condition indicators and descriptive fields.

[0128] In the optional embodiment, optionally, for each multi-modal data, the above performing a data conversion operation on the pre-processed multi-modal data to obtain target pre-processed data corresponding to the pre-processed multi-modal data can include: performing a format conversion operation on the pre-processed multi-modal data to make the data format of the pre-processed multi-modal data match a preset reference data format; wherein the format conversion operation can include one or more of an optical degree conversion operation and an optical angle conversion operation.

[0129] In this optional embodiment, further optionally, the patient's examination interval, examination type, and examination equipment can change in the long-term follow-up. Therefore, the long-term follow-up data not only have inter-individual differences, but also have intra-individual differences. In the present study, the long-term follow-up data mainly include medical records, fundus photographs, OCT, OCTA, optometry, and other multi-modal data, while the equipment model, scanning mode, or recording method of the same modal data is different. First, for medical images such as fundus color maps and OCT, a standard scanning mode is determined, and other scans are converted to the standard mode through image post-processing; second, for optometry examinations with different measurement or recording methods, corresponding formulas are selected to convert various formats, and patient medical records will be re-standardized and arranged, wherein various formats are converted, and different types of optical degrees or angles are converted to a standard format to ensure consistency and comparability between different equipment or systems.

[0130] It can be seen that implementing the optional embodiment can perform data target processing operations on each acquired multi-modal data to obtain preprocessed multi-modal data, and determine whether the data format corresponding to the preprocessed multi-modal data matches the preset reference data format. If not, perform data conversion operations on the preprocessed multi-modal data to obtain corresponding target preprocessed data. If it matches, the preprocessed data is determined as the target preprocessed data, and the target multi-modal data is generated based on all target preprocessed data. It can convert data of different scales or units to the same scale or range, avoid certain features occupying too much weight in the model, improve the generalization ability of the model, expand the data set by enhancing the diversity of data, improve the robustness of model learning and model training, reduce the probability of overfitting of the model, improve the quality of data and the performance of the model by performing data target processing operations, and ensure the consistency of data format by determining whether the data format matches the preset reference data format. It can reduce the risk of errors and model performance degradation caused by data format differences. If it is determined that the data format does not match, data conversion operations can be performed to ensure that all input data meet the requirements of the model, further improving the stability and performance of the model, thereby improving the accuracy and reliability of generating target multi-modal data, and improving the intelligence and efficiency of generating target multi-modal data. It is also beneficial to improve the quality and consistency of data, providing better input for subsequent model training, enabling the model to more effectively learn information in multi-modal data, improving the accuracy and generalization ability of model training, improving the accuracy and reliability of the constructed multi-modal intelligent diagnosis model, thereby improving the model performance of the multi-modal intelligent diagnosis model, and further improving the diagnosis intelligence and efficiency of myopia maculopathy. Further, it is also beneficial to improve the accuracy and reliability of diagnosing myopia maculopathy through an intelligent diagnosis model.

[0131] In another optional embodiment, for each target multi-modal data included in the target multi-modal database, data analysis operations are performed on the target multi-modal data to obtain data analysis results of the target multi-modal data, including:

[0132] For all target multi-modal data included in the target multi-modal database, data classification operations are performed to obtain target text data and target image data.

[0133] For each target text data, split operations are performed on the target text data to obtain target split data corresponding to the target text data, and text processing operations are performed on the target split data to obtain target vector data of the target text data.

[0134] For each target image data, an image compression operation is performed on the target image data to obtain image output data corresponding to the target image data.

[0135] Based on the target vector data corresponding to each target text data and the image output data corresponding to each target image data, a data splicing operation is performed to obtain target splicing data, and the target splicing data is input into a predetermined data analysis model to obtain a data analysis result of each target multi-modal data.

[0136] In this optional embodiment, optionally, the target text data can include word data; and the target image data can include image data.

[0137] In this optional embodiment, optionally, the above-mentioned, for each target text data, performing a splitting operation on the target text data to obtain target splitting data corresponding to the target text data, and performing a text processing operation on the target splitting data to obtain target vector data of the target text data, can include:

[0138] For each target text data, a splitting operation is performed on the target text data by a preset splitting model to obtain target splitting data corresponding to the target text data, a text encoding operation is performed on the target splitting data to obtain text encoding data, and a completion operation is performed on the text encoding data to obtain target vector data of the target text data.

[0139] In this optional embodiment, optionally, for example, an electronic case and the like text is split into words. The word is called token, the process of splitting the text into tokens is called tokenization, and the model or tool used for tokenization is called tokenizer. After completing tokenization, the text is encoded by an encoder to obtain a maximum of 256 text tokens, all word vectors are padded to the same length by padding, and then a vectorization is performed by an embedding layer; wherein, tokenizer is called tokenizer in Chinese, which is to divide a sentence into small word blocks (tokens), generate a word table, and learn a better representation through a model. Byte-Pair Encoding (BPE) is the most widely used subword tokenizer.

[0140] In this optional embodiment, optionally, the above-mentioned, for each target image data, an image compression operation is performed on the target image data to obtain image output data corresponding to the target image data, can include:

[0141] For each target image data, determine the image format corresponding to the target image data, perform image compression operation on the target image data based on the image format corresponding to the target image data, and obtain image output data corresponding to the target image data.

[0142] In this optional embodiment, optionally, for example, by training a VAE to compress each 256x256 image picture into a 32x32 picture token, each position has 8192 possible values, that is, the encoder output of the VAE is logits with a dimension of 32x32x8192, and then the logits index the features of the codebook for combination, and the embedding of the codebook is learnable.

[0143] In this optional embodiment, optionally, the above performing data splicing operation based on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and inputting the target splicing data into the pre-determined data analysis model to obtain the data analysis result of each target multi-modal data can include:

[0144] Performing data splicing operation based on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and inputting the target splicing data into the pre-determined data analysis model to obtain the data analysis result of each target multi-modal data through supervised training of the data analysis model; wherein the target splicing data at least includes the target vector data corresponding to each target text data and the image output data corresponding to each target image data; the pre-determined data analysis model can include a Transformer model.

[0145] In this optional embodiment, optionally, the target vector data corresponding to the target text data can include text token data; and the image output data corresponding to the target image data can include image token data. For example, in order to obtain effective image tokens, nearest neighbor mapping and straight-through estimation, Gumbel sampling and straight-through estimation, Gumbel sampling and Softmax approximation are introduced. The text token and the image token are spliced to obtain data with a length of N, and finally the spliced data is input into the Transformer for supervised training.

[0146] It can be seen that by implementing the optional embodiment, the data classification operation can be performed on the target multi-modal data to obtain target text data and target image data, the splitting operation can be performed on each target text data to obtain corresponding target split data, the text processing operation can be performed to obtain target vector data, the image compression operation can be performed on each target image data to obtain corresponding image output data, the data splicing operation can be performed on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and the data analysis model can be input to obtain the data analysis result of each target multi-modal data. The data classification operation can be performed on the target multi-modal data, which is beneficial to efficiently identifying and separating the target text data and the target image data, and is beneficial to improving the intelligent processing of data corresponding to each data type, improving the efficiency and intelligence of the overall data processing, extracting key information from the target text data through the splitting and text processing operations, capturing semantic information in the text, providing more valuable data basis for subsequent data analysis, reducing the storage space and transmission bandwidth of data through the image compression operation on the target image data, improving the efficiency of processing image data and reducing the consumption of computing and storage resources, forming a data structure containing multi-modal information through data splicing, fully utilizing the complementarity between different modal data to provide more comprehensive and accurate data analysis results, and obtaining the data analysis result of each target multi-modal data through the data analysis model processing the target splicing data. The model-based analysis method can automatically process a large amount of data and produce reliable analysis results, improving the efficiency and accuracy of data analysis. Comprehensive analysis of the target multi-modal data can provide more comprehensive and accurate information support for decision-making, which is beneficial to improving the accuracy and reliability of the multi-modal intelligent diagnosis model constructed, thereby improving the model performance of the multi-modal intelligent diagnosis model, and further improving the diagnosis intelligence and efficiency of myopic maculopathy, and further improving the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0147] In another optional embodiment, before constructing the target data chart according to the data analysis results of all target multi-modal data, the method further comprises:

[0148] performing a data feature distribution evaluation operation on all target multi-modal data to obtain a feature distribution evaluation result, and determining whether the feature distribution evaluation result meets a preset data drift processing condition;

[0149] When it is judged that the feature distribution evaluation result meets the preset data drift processing condition, at least one to-be-processed data is determined from all the target multi-modal data, and a data drift processing operation is performed on each to-be-processed data to update each to-be-processed data and update all the target multi-modal data.

[0150] Based on the updated all target multi-modal data, it is judged whether all the target multi-modal data meets the preset modal data integrity condition.

[0151] When it is judged that all the target multi-modal data does not meet the preset modal data integrity condition, a feature splicing operation is performed on all the target multi-modal data to obtain feature splicing data, and based on the feature splicing data, the data analysis result is updated, and the operation of constructing a target data chart according to the data analysis result of all the target multi-modal data is triggered.

[0152] In this optional embodiment, optionally, the feature distribution evaluation result can include the distribution of data features, wherein the distribution of data features can include the structure and trend of the data set.

[0153] In this optional embodiment, optionally, the above judging whether the feature distribution evaluation result meets the preset data drift processing condition can include:

[0154] According to the feature distribution evaluation result, the data feature dispersion degree is determined, and it is judged whether the data feature dispersion degree is greater than or equal to a preset dispersion degree threshold value.

[0155] When it is judged that the data feature dispersion degree is greater than or equal to the preset dispersion degree threshold value, it is determined that the feature distribution evaluation result meets the preset data drift processing condition; when it is judged that the data feature dispersion degree is less than the preset dispersion degree threshold value, it is determined that the feature distribution evaluation result does not meet the preset data drift processing condition.

[0156] In this optional embodiment, further optionally, when it is judged that the feature distribution evaluation result does not meet the preset data drift processing condition, the operation of constructing a target data chart according to the data analysis result of all the target multi-modal data can be directly triggered.

[0157] In this optional embodiment, optionally, the to-be-processed data can include one or more of domain adaptation data, instance weight adjustment data, and feature selection data.

[0158] In this optional embodiment, optionally, the data drift processing operation can include one or more of feature alignment operation, domain alignment operation, and adversarial training operation.

[0159] In this optional embodiment, optionally, the data drift handling operation can be implemented by the following steps: first, the data of the source domain (e.g. the original dataset) and the target domain (the new dataset) need to be collected; the source domain is the domain that the model has seen during training, while the target domain is the new domain that the model will be applied to; before starting to solve the problem, it is necessary to understand the data drift between the source domain and the target domain, which may be caused by different data collection devices, environmental conditions, distribution bias, etc.; according to the characteristics of data drift, appropriate transfer learning techniques are selected, which may include domain adaptation, instance weight adjustment, feature selection, even layer migration in deep neural networks, etc.; a base model is trained using the source domain data, then the model is fine-tuned or transferred using the target domain data, which can be achieved through various methods such as feature alignment, domain alignment, adversarial training, etc. to adjust the model parameters to adapt to the data of the target domain; evaluate the performance of the adjusted model on the target domain, which includes comparing the performance difference of the model between the source domain and the target domain, and checking whether the model has successfully adapted to the data distribution of the target domain; according to the evaluation results, the model and the transfer learning process may need to be further adjusted and optimized, which may include changing the transfer learning technique, adjusting the model architecture or hyperparameters, etc.; through this process, the data drift problem can be effectively solved, and the model can achieve good performance in the new domain or environment.

[0160] In this optional embodiment, optionally, the above-mentioned performing feature splicing operation on all target multi-modal data to obtain feature splicing data, and updating the data analysis result based on the feature splicing data can include: performing feature splicing operation on all target multi-modal data to obtain feature splicing data, performing fusion collocation operation on the feature splicing data to obtain data fusion results corresponding to different feature splicing data, and determining an optimal fusion strategy based on all data fusion results, and updating the data analysis result according to the data fusion result corresponding to the optimal fusion strategy.

[0161] In the optional embodiment, optionally, for example, different multi-modal data fusion collocations and strategies are tried and compared, such as input layer splicing, deep feature splicing and the like, and the optimal strategy is selected to solve the problem of missing modal data; wherein, when the fusion collocation operation includes the input layer splicing method, the data of different modalities are spliced into a larger feature vector in the input layer, which means that the features of all modalities are connected to form a longer input vector, and then the vector is input into the model, which can be simple and direct, easy to implement, and can fully utilize all available modal information; when the fusion collocation operation includes the deep feature splicing method, in the deep network, the features of different modalities are input into respective branch networks for processing, and then the features of different branch networks are spliced or fused to form a mixed feature representation, which is then input into the subsequent network layer for learning and prediction, which can make the model better learn the correlation between different modalities, and the processing of missing data is more flexible, each branch network can independently process the features of different modalities, and can selectively fuse the features of different modalities. Further, for selecting the optimal strategy, the characteristics of the data, the architecture of the model, the computing resources and the like need to be considered: if the correlation between different modal data is weak or the data is missing, the input layer splicing method can be considered to simplify the model structure and reduce the computational complexity. If there is a strong correlation between different modal data, or the model is expected to better handle missing data, the deep feature splicing method can be considered to better utilize the information of different modalities. In actual application, the optimal strategy usually needs to be determined through experiments and verification, and is selected according to the specific problem requirements and the characteristics of the data.

[0162] In the optional embodiment, further optionally, when it is judged that all target multi-modal data meet the preset modal data integrity condition, an operation of constructing a target data chart according to a data analysis result of all the target multi-modal data can be directly triggered to be executed.

[0163] It can be seen that implementing the optional embodiment can perform data feature distribution evaluation operation on all target multi-modal data to obtain feature distribution evaluation result and judge whether the preset data drift processing condition is met. If it is met, at least one to-be-processed data is determined from all target multi-modal data, and data drift processing operation is performed to update each to-be-processed data and update all target multi-modal data. It is judged whether the modal data integrity condition is met based on the updated all target multi-modal data. If it is not met, the feature splicing operation is performed on all target multi-modal data to obtain feature splicing data and update the data analysis result, and then the operation of constructing the target data chart according to the data analysis result of all target multi-modal data is triggered. By performing the data feature distribution evaluation operation, the possible abnormalities or deviations in the data can be identified, thereby improving the overall quality of the data. When the feature distribution evaluation result shows that the data has drift, timely data drift processing can ensure the stability and consistency of the data, can reduce the influence of data drift on the data analysis result, and can improve the accuracy of the data chart. If the data does not meet the integrity condition, the missing part in the data can be made up through the feature splicing operation, thereby ensuring the comprehensiveness and accuracy of data analysis. Based on the updated target multi-modal data, the data analysis result will be more accurate and reliable, which is conducive to improving the accuracy and reliability of the data analysis result, and can be conducive to ensuring the accuracy and integrity of the data, which is conducive to improving the quality of the data analysis result, and is conducive to improving the accuracy and reliability of the multi-modal intelligent diagnosis model constructed, thereby being conducive to improving the model performance of the multi-modal intelligent diagnosis model, and further being conducive to improving the diagnosis intelligence and efficiency of myopic maculopathy, and further being conducive to improving the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0164] In yet another optional embodiment, target feature information is extracted from the target data chart, and a target feature map is generated based on all target feature information, including:

[0165] The target data chart is subjected to chart normalization processing to obtain a key data chart, and the key data chart is subjected to feature extraction operation to obtain target feature information, wherein the target feature information includes spatial domain feature information and / or time domain feature information;

[0166] Based on all target feature information, the node dimension of the extension node is determined, the feature propagation operation is performed on the target feature information based on the extension node and the node dimension to obtain a feature propagation result;

[0167] According to the target feature information and the feature propagation result, feature aggregation information is obtained, and a target feature map is generated based on the feature aggregation information.

[0168] In the optional embodiment, optionally, the performing graph normalization on the target data graph to obtain the key data graph can include: performing normalization processing on the input feature matrix contained in the target data graph by applying a batch normalization layer to obtain the key data graph.

[0169] In the optional embodiment, optionally, the performing feature extraction operation on the key data graph to obtain the target feature information can be performed by multiple ST-GCN units, and the number of ST-GCN units can be six, and the ST-GCN unit is used to extract the features of the patient graph in the spatial domain and the time domain.

[0170] In the optional embodiment, optionally, the node dimension of the extended node can be determined by each ST-GCN unit, wherein the ST-GCN unit is composed of three layers. The first layer performs a conventional two-dimensional (2D) convolution operation to expand the dimension of the input node feature.

[0171] In the optional embodiment, optionally, the performing feature propagation operation can be performed by propagating the features of the extended node along the edges of the graph by graph convolution. The GCN is a deep learning model specially designed for processing graph structured data. It can utilize the relationships between nodes in the graph to perform information transmission and feature extraction. The "features of the extended node" refer to updating or expanding the node features to contain more information or better reflect the state of the node. GCN utilizes the structure of the graph to propagate information, by propagating the node features along the edges of the graph, the nodes can exchange and update information with their adjacent nodes. In this way, each node can gather information from surrounding nodes to update its own feature representation.

[0172] In the optional embodiment, optionally, the obtaining feature aggregation information according to the target feature information and the feature propagation result can include: performing feature generation operation according to the target feature information and the feature propagation result to generate a feature map containing the aggregation information of the node and its neighbors, and generating the feature aggregation information based on the feature map corresponding to all nodes. Further optionally, a feature map containing the aggregation information of the node and its neighbors can be generated, and the last layer is similar to the first layer but has a different kernel size.

[0173] It can be seen that the optional embodiment can perform graph normalization processing on the target data graph to obtain a key data graph, perform feature extraction on the key data graph to obtain target feature information, determine the node dimension of the extended node based on all target feature information to perform feature propagation on the target feature information to obtain a feature propagation result, obtain feature aggregation information according to the target feature information and the feature propagation result, and generate a target feature mapping. By performing normalization processing on the target data graph, a key data graph can be obtained, which helps to eliminate the deviation caused by different units, scales or ranges between different data graphs, thereby improving the accuracy of subsequent feature extraction. The spatial domain feature information and the time domain feature information can comprehensively reflect the distribution, change and trend of the data in the data graph. Through rich feature information, the actual problem represented by the data graph can be more accurately understood and analyzed. Based on the extended node and the node dimension, the feature propagation operation is performed on the target feature information, which can ensure that the propagation of the feature information in the data graph is effective and accurate, and is beneficial to capturing the correlation of different regions or different time periods in the data graph, thereby improving the integrity and accuracy of the feature mapping. The generated target feature mapping can be used as the input of the machine learning or deep learning model. By fully utilizing the spatial and temporal information in the graph, the prediction accuracy and generalization ability of the model can be improved, thereby improving the accuracy, richness and reliability of data processing, improving the accuracy and reliability of the constructed multi-modal intelligent diagnosis model, and improving the model performance of the multi-modal intelligent diagnosis model, thereby improving the diagnosis intelligence and efficiency of myopic maculopathy, and further improving the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0174] In yet another optional embodiment, based on the output feature mapping and the predetermined fully connected layer, a multi-modal intelligent diagnosis model is constructed, comprising:

[0175] Based on the output feature mapping, a convolution unit is determined, and a channel attention layer corresponding to each convolution unit is determined;

[0176] Based on each convolution unit and the channel attention layer corresponding to each convolution unit, a target output channel is determined, and a global average pooling layer is generated based on the target output channel;

[0177] According to the global average pooling layer and the predetermined fully connected layer, a multi-modal intelligent diagnosis model is constructed.

[0178] In the optional embodiment, optionally, the convolution unit can comprise an ST-GCN unit; further optionally, there is a channel attention layer corresponding to each convolution unit, wherein the channel attention layer corresponding to each convolution unit is used to help the model pay attention to more meaningful and functional channels. Further, there is a global average pooling layer after each channel, and the outputs of these pooling layers are connected together to realize feature fusion and generate the final patient feature map.

[0179] In the optional embodiment, optionally, the multi-modal intelligent diagnosis model constructed according to the global average pooling layer and the pre-determined fully connected layer can comprise:

[0180] The global average pooling layer and the pre-determined fully connected layer are subjected to a splicing operation, and a multi-modal intelligent diagnosis model is constructed based on a pre-determined sigmoid activation function.

[0181] In the optional embodiment, optionally, for example, the fully connected layer receives all nodes of the previous layer as input, multiplies them with weights and adds a bias, and then passes them to an activation function. The sigmoid function is often used in the output layer because it can interpret the output as a probability, representing the probability of the positive class. In the output layer of the model, through the fully connected layer and the sigmoid activation function, the final output of the diagnosis prediction is generated. This means that the output of the model is a value between 0 and 1, representing the probability of predicting the positive class (e.g. the existence of a certain disease).

[0182] It can be seen that implementing the optional embodiment can determine the convolution unit based on the output feature map and determine the channel attention layer corresponding to each convolution unit, determine the target output channel generation global average pooling layer based on each convolution unit and the channel attention layer corresponding to each convolution unit, and construct the multi-modal intelligent diagnosis model according to the global average pooling layer and the full connection layer. Through the convolution unit, the model can efficiently extract key features from multi-modal data, and the introduction of the channel attention layer enables the model to focus on the feature channel that is most valuable for diagnosis, further improving the relevance of feature extraction. The global average pooling layer can integrate the output of each convolution unit into global information, and such integration of global information is of great significance to improving the accuracy and robustness of diagnosis. Through the pre-determined full connection layer, the features learned by the previous layer can be combined to perform classification or other tasks. In the multi-modal intelligent diagnosis model, the full connection layer is responsible for fusing the learned multi-modal features, and making accurate diagnosis decisions based on these fused features. By combining multi-modal data and deep learning technology, the multi-modal intelligent diagnosis model has stronger generalization ability, thereby improving the model performance of the multi-modal intelligent diagnosis model, and further improving the diagnosis intelligence and efficiency of myopic maculopathy. Further, it is also beneficial to improve the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0183] Embodiment three

[0184] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of a myopic maculopathy multi-modal intelligent diagnosis model generation device disclosed by an embodiment of the application. As shown in Figure 3 , the myopic maculopathy multi-modal intelligent diagnosis model generation device can include:

[0185] The preprocessing module 301 is configured to obtain multi-modal data, perform data preprocessing operations on each obtained multi-modal data, and obtain target multi-modal data.

[0186] The generation module 302 is configured to generate a target multi-modal database based on all target multi-modal data.

[0187] The analysis module 303 is configured to perform data analysis operations on each target multi-modal data included in the target multi-modal database, and obtain a data analysis result of the target multi-modal data.

[0188] The construction module 304 is configured to construct a target data chart according to the data analysis results of all target multi-modal data.

[0189] The generating module 302 is further configured to extract target feature information from the target data graph and generate a target feature map based on all target feature information;

[0190] The construction module 304 is further used to construct a multimodal intelligent diagnosis model based on all target feature maps.

[0191] It can be seen that implementation Figure 3 The described device can acquire multimodal data and perform data preprocessing operations to obtain target multimodal data, generate a target multimodal database based on all target multimodal data, perform data analysis operations on each target multimodal data included in the target multimodal database to obtain data analysis results, construct a target data chart based on the data analysis results of all target multimodal data, and extract target feature information from the target data chart to generate a target feature map, and construct a multimodal intelligent diagnostic model based on all target feature maps. It can provide more comprehensive and richer information by acquiring multimodal data, and can improve the accuracy and consistency of data by performing data preprocessing operations, which can bring convenience to subsequent data analysis and model training, and the target feature information obtained through data analysis and icon extraction can accurately reflect the content of the data. In terms of features and structure, it provides strong support for building an efficient and accurate multimodal intelligent diagnostic model, and the multimodal intelligent diagnostic model constructed based on target feature mapping can integrate multimodal information from multiple aspects to achieve cross-modal information fusion and complementarity, which is conducive to improving the subsequent model training and the accuracy and reliability of the intelligent diagnostic model. Compared with the traditional single-modality diagnosis method, the multimodal intelligent diagnostic model has stronger adaptability and generalization ability, and can cope with more complex and more variable diagnostic scenarios. By applying the multimodal intelligent diagnostic model, the accuracy and efficiency of diagnosis can be improved, the diagnosis cost can be reduced, and a better experience and benefits can be brought to users, which is also conducive to improving the diagnostic intelligence and efficiency of myopic macular degeneration, and further conducive to improving the accuracy and reliability of the diagnosis of myopic macular degeneration through the intelligent diagnostic model.

[0192] In an optional embodiment, if Figure 4 As shown, the device also includes:

[0193] The calculation module 305 is used to perform a convolution operation on all target feature maps to obtain a feature operation result before the construction module 304 constructs a multimodal intelligent diagnosis model based on all target feature maps;

[0194] A first determination module 306 is used to determine the output channel and the average pooling layer based on the feature operation result;

[0195] The generation module 302 is further used to generate an output feature map based on all output channels, the average pooling layer and the feature operation results;

[0196] The specific method of constructing the multimodal intelligent diagnosis model by the construction module 304 based on all target feature maps includes:

[0197] Based on the output feature map and the predetermined fully connected layer, a multimodal intelligent diagnosis model is constructed.

[0198] It can be seen that implementation Figure 4 The described device can perform convolution operations on all target feature maps to obtain feature operation results and determine output channels and average pooling layers, generate output feature maps based on all output channels, average pooling layers and feature operation results, and construct a multimodal intelligent diagnosis model based on the output feature maps and fully connected layers. It can further extract and enhance key features in the data by performing convolution operations on the target feature maps, enable the model to capture the local spatial relationship of the data, and help improve the expressive power of features. At the same time, it can combine information from different channels through the output feature maps of the convolution operation, realize the integration of cross-modal features, and use the complementary information in the multimodal data to improve the accuracy and reliability of the subsequently constructed multimodal intelligent diagnosis model. By combining convolution operations, average pooling layers and fully connected layers, the constructed multimodal intelligent diagnosis model can fully utilize the advantages of multimodal data and improve the accuracy and efficiency of diagnosis, so that the model has strong generalization ability and can adapt to the needs of different fields and scenarios. It can optimize the model structure through the effectiveness of feature extraction and enhancement, the flexibility of feature integration, and the optimization of output channels and average pooling layers, which is conducive to improving the accuracy and reliability of the constructed multimodal intelligent diagnosis model, thereby improving the model performance of the multimodal intelligent diagnosis model, and further helping to improve the diagnostic intelligence and efficiency of myopic maculopathy, and further helping to improve the accuracy and reliability of the diagnosis of myopic maculopathy through the intelligent diagnostic model.

[0199] In another optional embodiment, Figure 4 As shown, the preprocessing module 301 performs data preprocessing operations on each acquired multimodal data, and the specific manner of obtaining the target multimodal data includes:

[0200] For each acquired multimodal data, performing a data target processing operation on the multimodal data to obtain preprocessed multimodal data corresponding to the multimodal data, wherein the data target processing operation includes a data normalization operation and / or a data enhancement operation;

[0201] For each multi-modal data, it is judged whether the data format corresponding to the pre-processed multi-modal data matches the preset reference data format according to the pre-processed multi-modal data corresponding to the multi-modal data;

[0202] When it is judged that the data format corresponding to the pre-processed multi-modal data does not match the preset reference data format, a data conversion operation is performed on the pre-processed multi-modal data to obtain target pre-processed data corresponding to the pre-processed multi-modal data; when it is judged that the data format corresponding to the pre-processed multi-modal data matches the preset reference data format, the pre-processed multi-modal data is determined as the target pre-processed data;

[0203] Based on all the target pre-processed data, target multi-modal data is generated.

[0204] It can be seen that the implementation Figure 4 The described device can perform data target processing operations on each multi-modal data obtained to obtain pre-processed multi-modal data, judge whether the data format corresponding to the pre-processed multi-modal data matches the preset reference data format, if not, perform a data conversion operation on the pre-processed multi-modal data to obtain corresponding target pre-processed data, if so, determine the pre-processed as the target pre-processed data, and generate target multi-modal data based on all the target pre-processed data, which can convert data of different scales or units into the same scale or range, avoid certain features occupying too large a weight in the model, improve the generalization ability of the model, expand the data set by enhancing the diversity of the data, improve the robustness of model learning and model training, reduce the probability of overfitting of the model, improve the quality of the data and the performance of the model by performing data target processing operations, and ensure the consistency of the data format by judging whether the data format matches the preset reference data format, reduce the risk of errors and model performance decline caused by data format differences, perform data conversion operations if the data format is not matched, ensure that all input data meet the requirements of the model, further improve the stability and performance of the model, thereby improving the accuracy and reliability of generating target multi-modal data, improving the intelligence and efficiency of generating target multi-modal data, improving the quality and consistency of the data, providing better input for subsequent model training, enabling the model to more effectively learn information in multi-modal data, improving the accuracy and generalization ability of model training, improving the accuracy and reliability of the multi-modal intelligent diagnosis model constructed, thereby improving the model performance of the multi-modal intelligent diagnosis model, and further improving the diagnosis intelligence and efficiency of myopic maculopathy, further improving the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0205] In yet another optional embodiment, as shown in FIG. 3B, the analysis module 303 performs the data analysis operation on each target multi-modal data included in the target multi-modal database in the following manner: Figure 4

[0206] performing the data classification operation on all the target multi-modal data included in the target multi-modal database to obtain target text data and target image data;

[0207] for each target text data, performing a splitting operation on the target text data to obtain target splitting data corresponding to the target text data, and performing a text processing operation on the target splitting data to obtain target vector data of the target text data;

[0208] for each target image data, performing an image compression operation on the target image data to obtain image output data corresponding to the target image data;

[0209] performing a data splicing operation based on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and inputting the target splicing data into a pre-determined data analysis model to obtain a data analysis result of each target multi-modal data.

[0210] It can be seen that the implementation of the data analysis operation in the analysis module 303 is more efficient and more accurate. Figure 4 ​The described device can perform data classification operations on target multi-modal data to obtain target text data and target image data, perform splitting operations on each target text data to obtain corresponding target split data, perform text processing operations to obtain target vector data, perform image compression operations on each target image data to obtain corresponding image output data, perform data splicing operations on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and input the target splicing data into a data analysis model to obtain data analysis results of each target multi-modal data. The device can perform data classification operations on target multi-modal data, which is beneficial to efficiently identifying and separating target text data and target image data, improving intelligent processing of data corresponding to each data type, improving the efficiency and intelligence of overall data processing, and performing splitting and text processing operations on target text data to extract key information and convert it into vector data, capturing semantic information in the text and providing more valuable data basis for subsequent data analysis. By performing image compression operations on target image data, the device can reduce data storage space and transmission bandwidth, improve the efficiency of processing image data, and reduce the consumption of computing and storage resources. By splicing data, the device can form a data structure containing multi-modal information, which is beneficial to fully utilizing the complementarity between different modal data to provide more comprehensive and accurate data analysis results. By processing target splicing data through a data analysis model, the device can obtain data analysis results of each target multi-modal data, automatically process a large amount of data through a model-based analysis method, and produce reliable analysis results, improving the efficiency and accuracy of data analysis. By comprehensively analyzing target multi-modal data, the device can provide more comprehensive and accurate information support for decision-making, which is beneficial to improving the accuracy and reliability of the constructed multi-modal intelligent diagnosis model, thereby improving the model performance of the multi-modal intelligent diagnosis model, further improving the diagnosis intelligence and efficiency of myopic macular degeneration, and further improving the accuracy and reliability of diagnosing myopic macular degeneration through the intelligent diagnosis model.

[0211] In yet another optional embodiment, as shown in Figure 4 The device further includes:

[0212] The evaluation module 307 is configured to perform data feature distribution evaluation operations on all target multi-modal data to obtain feature distribution evaluation results before the construction module 304 constructs a target data graph based on data analysis results of all target multi-modal data.

[0213] The judgment module 308 is configured to determine whether the feature distribution evaluation results meet a preset data drift processing condition.

[0214] The second determination module 309 is configured to determine at least one to-be-processed data from all the target multi-modal data, and perform a data drift processing operation on each to-be-processed data to update each to-be-processed data and update all the target multi-modal data.

[0215] The judgment module 308 is further configured to judge whether all the target multi-modal data satisfy a preset modal data integrity condition based on the updated all the target multi-modal data.

[0216] The splicing module 310 is configured to perform a feature splicing operation on all the target multi-modal data to obtain feature splicing data when the judgment module 308 judges that all the target multi-modal data do not satisfy the preset modal data integrity condition, update the data analysis result based on the feature splicing data, and trigger the operation of constructing the target data chart according to the data analysis result of all the target multi-modal data performed by the construction module 304.

[0217] It can be seen that the implementation Figure 4 The described device can perform a data feature distribution evaluation operation on all the target multi-modal data to obtain a feature distribution evaluation result and judge whether it satisfies a preset data drift processing condition, if it satisfies, determine at least one to-be-processed data from all the target multi-modal data and perform a data drift processing operation to update each to-be-processed data and update all the target multi-modal data, judge whether it satisfies a modal data integrity condition based on the updated all the target multi-modal data, if it does not satisfy, perform a feature splicing operation on all the target multi-modal data to obtain feature splicing data and update the data analysis result, and then trigger the operation of constructing the target data chart according to the data analysis result of all the target multi-modal data, can recognize the possible abnormalities or deviations in the data by performing the data feature distribution evaluation operation, thereby improving the overall quality of the data, when the feature distribution evaluation result shows that the data has drift, timely data drift processing can ensure the stability and consistency of the data, can reduce the influence of data drift on the data analysis result, improve the accuracy of the data chart, if the data does not satisfy the integrity condition, the missing part in the data can be made up through the feature splicing operation, thereby ensuring the comprehensiveness and accuracy of the data analysis, the data analysis result will be more accurate and reliable based on the updated target multi-modal data, which is conducive to improving the accuracy and reliability of the data analysis result, and can be conducive to ensuring the accuracy and integrity of the data, which is conducive to improving the quality of the data analysis result, and is conducive to improving the precision and reliability of the multi-modal intelligent diagnosis model constructed, thereby being conducive to improving the model performance of the multi-modal intelligent diagnosis model, and further being conducive to improving the diagnosis intelligence and efficiency of myopic maculopathy, and further being conducive to improving the precision and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0218] In another optional embodiment, Figure 4 As shown, the generation module 302 extracts target feature information from the target data chart and generates a target feature map based on all target feature information in the following manner:

[0219] Performing a chart normalization process on the target data chart to obtain a key data chart, and performing a feature extraction operation on the key data chart to obtain target feature information, wherein the target feature information includes spatial domain feature information and / or temporal domain feature information;

[0220] Based on all target feature information, determine the node dimension of the expanded node, and based on the expanded node and the node dimension, perform feature propagation operation on the target feature information to obtain feature propagation results;

[0221] According to the target feature information and the feature propagation result, feature aggregation information is obtained, and a target feature map is generated based on the feature aggregation information.

[0222] It can be seen that implementation Figure 4 The described device can perform chart normalization processing on the target data chart to obtain a key data chart, and perform feature extraction operations on the key data chart to obtain target feature information, determine the node dimension of the extended node based on all target feature information, and thus perform feature propagation operations on the target feature information to obtain feature propagation results, obtain feature aggregation information based on the target feature information and the feature propagation results, and then generate a target feature map. By performing normalization processing on the target data chart, a key data chart can be obtained, which helps to eliminate the deviations between different data charts due to different units, scales or ranges, thereby improving the accuracy of subsequent feature extraction. The extracted spatial domain feature information and time domain feature information can fully reflect the distribution, changes and trends of data in the data chart. The rich feature information helps to more accurately understand and analyze the actual problems represented by the data chart. Performing feature propagation operations on target feature information based on extended nodes and node dimensions can ensure that the propagation of feature information in the data graph is effective and accurate, which is conducive to capturing the correlation between different regions or different time periods in the data graph, thereby improving the integrity and accuracy of the feature map. The generated target feature map can be used as the input of the machine learning or deep learning model. By making full use of the spatial and temporal information in the graph, the prediction accuracy and generalization ability of the model can be improved, which is conducive to improving the accuracy, richness and reliability of data processing, and is conducive to improving the accuracy and reliability of the constructed multimodal intelligent diagnostic model, thereby helping to improve the model performance of the multimodal intelligent diagnostic model, and further helping to improve the diagnostic intelligence and efficiency of myopic maculopathy, and further helping to improve the accuracy and reliability of the diagnosis of myopic maculopathy through the intelligent diagnostic model.

[0223] In yet another optional embodiment, as shown in Figure 4 The specific way of constructing the multi-modal intelligent diagnosis model by the construction module 304 based on the output feature map and the pre-determined fully connected layer includes:

[0224] Based on the output feature map, determine the convolution unit, and determine the channel attention layer corresponding to each convolution unit;

[0225] Based on each convolution unit and the channel attention layer corresponding to each convolution unit, determine the target output channel, and generate the global average pooling layer based on the target output channel;

[0226] According to the global average pooling layer and the pre-determined fully connected layer, the multi-modal intelligent diagnosis model is constructed.

[0227] It can be seen that the apparatus described in the embodiments Figure 5 The apparatus described in the embodiments can determine the convolution unit based on the output feature map and determine the channel attention layer corresponding to each convolution unit, determine the target output channel based on each convolution unit and the channel attention layer corresponding to each convolution unit to generate the global average pooling layer, and construct the multi-modal intelligent diagnosis model according to the global average pooling layer and the fully connected layer. Through the convolution unit, the model can efficiently extract key features from multi-modal data, and the introduction of the channel attention layer enables the model to focus on the most valuable feature channels for diagnosis, further improving the relevance of feature extraction. The global average pooling layer can integrate the output of each convolution unit into global information, which is of great significance for improving the accuracy and robustness of diagnosis. Through the pre-determined fully connected layer, the features learned by the previous layers can be combined to perform classification or other tasks. In the multi-modal intelligent diagnosis model, the fully connected layer is responsible for fusing the learned multi-modal features and making accurate diagnosis decisions based on these fused features. By combining multi-modal data and deep learning technology, the multi-modal intelligent diagnosis model constructed has stronger generalization ability, thereby improving the model performance of the multi-modal intelligent diagnosis model, and further improving the diagnosis intelligence and efficiency of myopic maculopathy. Furthermore, it is also beneficial to improve the accuracy and reliability of diagnosing myopic maculopathy through the intelligent diagnosis model.

[0228] Embodiment Four

[0229] Please refer to Figure 5 , Figure 5 is another structural schematic diagram of a multi-modal intelligent diagnosis model generation apparatus for myopic maculopathy disclosed in an embodiment of the present application. As shown in ​ The multi-modal intelligent diagnosis model generation apparatus for myopic maculopathy can include:

[0230] The memory 401 stores executable program codes;

[0231] The processor 402 is coupled with the memory 401.

[0232] The processor 402 invokes the executable program codes stored in the memory 401 to perform the steps in the myopic maculopathy multi-modal intelligent diagnosis model generation method described in the embodiment one or the embodiment two of the present application.

[0233] Embodiment five

[0234] The embodiment of the present application discloses a computer storage medium, which stores computer instructions, and when the computer instructions are invoked, the computer instructions are used to perform the steps in the myopic maculopathy multi-modal intelligent diagnosis model generation method described in the embodiment one or the embodiment two of the present application.

[0235] Embodiment six

[0236] The embodiment of the present application discloses a computer program product, which includes a non-transitory computer readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the myopic maculopathy multi-modal intelligent diagnosis model generation method described in the embodiment one or the embodiment two.

[0237] The above-described device embodiments are only schematic, and the modules described as separate components can or can not be physically separate, and the components shown as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed on multiple network modules. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0238] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and the necessary general hardware platform through the specific description of the above embodiments, and of course, the various embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other computer readable medium that can be used to carry or store data.

[0239] Finally, it should be noted that: the multi-modal intelligent diagnosis model generation method and device for myopia maculopathy disclosed by the embodiments of the present application are only the preferred embodiments of the present application, and are used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that; it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating a multimodal intelligent diagnostic model for myopic maculopathy, characterized in that: The method comprises: Acquire multimodal data, perform data preprocessing on each acquired multimodal data to obtain target multimodal data, and generate a target multimodal database based on all the target multimodal data; For each target multimodal data included in the target multimodal database, performing a data analysis operation on the target multimodal data to obtain a data analysis result of the target multimodal data; Constructing a target data graph based on data analysis results of all the target multimodal data, extracting target feature information from the target data graph, and generating a target feature map based on all the target feature information; Based on all the target feature maps, a multimodal intelligent diagnosis model is constructed.

2. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 1, wherein: Before constructing the multimodal intelligent diagnosis model based on all the target feature maps, the method further includes: Performing a convolution operation on all the target feature maps to obtain feature operation results, and determining an output channel and an average pooling layer based on the feature operation results; Generate an output feature map based on all the output channels, the average pooling layer and the feature operation results; Wherein, the multimodal intelligent diagnosis model is constructed based on all the target feature maps, including: Based on the output feature map and the predetermined fully connected layer, a multimodal intelligent diagnosis model is constructed.

3. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 1 or 2, characterized in that: The performing of a data preprocessing operation on each of the acquired multimodal data to obtain target multimodal data includes: For each of the acquired multimodal data, performing a data target processing operation on the multimodal data to obtain preprocessed multimodal data corresponding to the multimodal data, wherein the data target processing operation includes a data normalization operation and / or a data enhancement operation; For each multimodal data, determining, based on pre-processed multimodal data corresponding to the multimodal data, whether a data format corresponding to the pre-processed multimodal data matches a preset reference data format; When it is determined that the data format corresponding to the preprocessed multimodal data does not match the preset reference data format, performing a data conversion operation on the preprocessed multimodal data to obtain target preprocessed data corresponding to the preprocessed multimodal data; when it is determined that the data format corresponding to the preprocessed multimodal data matches the preset reference data format, determining the preprocessed multimodal data as the target preprocessed data; Based on all the target preprocessed data, target multimodal data is generated.

4. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 1 or 2, wherein: The step of performing a data analysis operation on each target multimodal data included in the target multimodal database to obtain a data analysis result of the target multimodal data includes: Performing a data classification operation on all the target multimodal data included in the target multimodal database to obtain target text data and target image data; For each target text data, performing a splitting operation on the target text data to obtain target split data corresponding to the target text data, and performing a text processing operation on the target split data to obtain target vector data of the target text data; For each target image data, performing an image compression operation on the target image data to obtain image output data corresponding to the target image data; A data splicing operation is performed based on the target vector data corresponding to each target text data and the image output data corresponding to each target image data to obtain target splicing data, and the target splicing data is input into a predetermined data analysis model to obtain a data analysis result for each target multimodal data.

5. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 1 or 2, wherein: Before constructing a target data chart based on the data analysis results of all the target multimodal data, the method further includes: Performing a data feature distribution evaluation operation on all the target multimodal data to obtain a feature distribution evaluation result, and determining whether the feature distribution evaluation result meets a preset data drift processing condition; When it is determined that the feature distribution evaluation result satisfies the preset data drift processing condition, at least one to-be-processed data is determined from all the target multimodal data, and a data drift processing operation is performed on each of the to-be-processed data to update each of the to-be-processed data and all the target multimodal data; Based on all the updated target multimodal data, determining whether all the target multimodal data meets a preset modality data completeness condition; When it is determined that all the target multimodal data do not meet the preset modal data completeness condition, a feature splicing operation is performed on all the target multimodal data to obtain feature splicing data, and based on the feature splicing data, the data analysis result is updated, and the operation of constructing a target data chart based on the data analysis results of all the target multimodal data is triggered.

6. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 1 or 2, wherein: The step of extracting target feature information from the target data chart and generating a target feature map based on all the target feature information includes: Performing a chart normalization process on the target data chart to obtain a key data chart, and performing a feature extraction operation on the key data chart to obtain target feature information, wherein the target feature information includes spatial domain feature information and / or temporal domain feature information; Determining the node dimension of the extended node based on all the target feature information, and performing a feature propagation operation on the target feature information based on the extended node and the node dimension to obtain a feature propagation result; Feature aggregation information is obtained according to the target feature information and the feature propagation result, and a target feature map is generated based on the feature aggregation information.

7. The method for generating a multimodal intelligent diagnostic model for myopic maculopathy according to claim 2, wherein: The multimodal intelligent diagnosis model is constructed based on the output feature map and the predetermined fully connected layer, including: Determining a convolution unit based on the output feature map, and determining a channel attention layer corresponding to each convolution unit; Determine a target output channel based on each of the convolution units and the channel attention layer corresponding to each of the convolution units, and generate a global average pooling layer based on the target output channel; A multimodal intelligent diagnosis model is constructed based on the global average pooling layer and the predetermined fully connected layer.

8. A multimodal intelligent diagnostic model generation device for myopic maculopathy, characterized in that: The device comprises: A preprocessing module is used to obtain multimodal data and perform data preprocessing operations on each of the obtained multimodal data to obtain target multimodal data; A generating module, configured to generate a target multimodal database based on all the target multimodal data; an analysis module, configured to perform a data analysis operation on each target multimodal data included in the target multimodal database to obtain a data analysis result of the target multimodal data; A construction module, configured to construct a target data chart based on data analysis results of all the target multimodal data; The generating module is further configured to extract target feature information from the target data chart and generate a target feature map based on all the target feature information; The construction module is also used to construct a multimodal intelligent diagnosis model based on all the target feature maps.

9. A multimodal intelligent diagnostic model generation device for myopic maculopathy, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the multimodal intelligent diagnostic model generation method for myopic maculopathy as described in any one of claims 1-7.

10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the multimodal intelligent diagnostic model generation method for myopic maculopathy as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Myopic maculopathy ATN grading method and system based on multi-modal medical image

    CN122335689A