Lung adenocarcinoma subtype classification method, device, equipment and medium

Through the lung adenocarcinoma subtype classification model of cross-modal coding and attention fusion module, the problem of artificial identification dependence is solved, and efficient and high-quality automated classification of lung adenocarcinoma subtypes is achieved.

CN120296474APending Publication Date: 2025-07-11SHENZHEN LONGGANG DISTRICT PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510448574.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the classification of lung adenocarcinoma subtypes depends on manual identification, and the classification quality and efficiency are limited by the ability of individual doctors, making it difficult to achieve efficient and high-quality classification.

Method used

The lung adenocarcinoma subtype classification model consisting of a cross-modal encoding module, attention fusion module and classification module is adopted to achieve automated classification by obtaining electronic health records and computed lung tomography images.

Benefits of technology

It significantly improves the efficiency and quality of lung adenocarcinoma subtype classification, enhances feature expression, and ensures the accuracy of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296474A_ABST
    Figure CN120296474A_ABST
Patent Text Reader

Abstract

The invention discloses a lung adenocarcinoma subtype classification method and device, equipment and a medium, and the method comprises the steps: obtaining an electronic health record and a lung computed tomography image of a current patient; determining a lung adenocarcinoma reference area in the lung computed tomography image; according to the lung adenocarcinoma reference region, performing cross-modal coding on the electronic health record and the lung computed tomography image through a cross-modal coding module of a lung adenocarcinoma subtype classification model to obtain a multi-modal feature; performing attention fusion on the multi-modal features through an attention fusion module of a lung adenocarcinoma subtype classification model to obtain cross-modal fusion features; lung adenocarcinoma subtype classification is carried out on the cross-modal fusion features through a classification module of the lung adenocarcinoma subtype classification model to obtain a classification result, and the classification efficiency and the classification quality of lung adenocarcinoma subtype classification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method for classifying subtypes of lung adenocarcinoma, a device for classifying subtypes of lung adenocarcinoma, a computer device, and a computer-readable storage medium. Background Art

[0002] The mortality rate of lung cancer ranks first among cancers. Among them, about 85% of lung cancers are non-small cell lung cancers, and lung adenocarcinoma is the most common type of non-small cell lung cancer. According to histological characteristics and invasion degree, lung adenocarcinoma can be further divided into subtypes such as in-situ adenocarcinoma, minimally invasive adenocarcinoma, and invasive adenocarcinoma. In the related art, doctors usually use the method of visual recognition to manually classify the computed tomography images of the lungs of patients to determine the subtype category of lung adenocarcinoma. However, the histological differences between different subtype categories of lung adenocarcinoma are subtle, which limits the classification quality and classification efficiency of this classification method based on manual recognition to the classification ability of individual doctors. Summary of the Invention

[0003] Embodiments of this application provide a method for classifying subtypes of lung adenocarcinoma, a device for classifying subtypes of lung adenocarcinoma, a computer device, and a computer-readable storage medium, which can improve the classification efficiency and classification quality of subtypes of lung adenocarcinoma.

[0004] In a first aspect, the method for classifying subtypes of lung adenocarcinoma provided by this application includes: Obtaining the electronic health record and computed tomography image of the lungs of the current patient; Determining the reference region of lung adenocarcinoma in the computed tomography image of the lungs; According to the reference region of lung adenocarcinoma, performing cross-modal encoding on the electronic health record and the computed tomography image of the lungs through the cross-modal encoding module of the lung adenocarcinoma subtype classification model to obtain multi-modal features; Performing attention fusion on the multi-modal features through the attention fusion module of the lung adenocarcinoma subtype classification model to obtain cross-modal fusion features; Performing classification of subtypes of lung adenocarcinoma on the cross-modal fusion features through the classification module of the lung adenocarcinoma subtype classification model to obtain a classification result.

[0005] In a second aspect, the device for classifying subtypes of lung adenocarcinoma provided by this application includes: A data acquisition module, configured to obtain the electronic health record and computed tomography image of the lungs of the current patient; A region determination module, configured to determine the reference region of lung adenocarcinoma in the computed tomography image of the lungs; A feature encoding module, configured to perform cross-modal encoding on the electronic health record and the computed tomography image of the lungs through the cross-modal encoding module of the lung adenocarcinoma subtype classification model according to the reference region of lung adenocarcinoma to obtain multi-modal features; A feature fusion module, which is used to perform attention fusion on multimodal features through the attention fusion module of the lung adenocarcinoma subtype classification model to obtain cross-modal fusion features; A feature classification module, which is used to classify the cross-modal fusion features into lung adenocarcinoma subtypes through the classification module of the lung adenocarcinoma subtype classification model to obtain a classification result.

[0006] In a third aspect, the computer device provided by the present application includes a processor and a memory. A computer program that can run on the processor is stored in the memory. When the processor runs the computer program, the lung adenocarcinoma subtype classification method provided by the embodiments of the present application is implemented.

[0007] In a fourth aspect, the computer-readable storage medium provided by the present application stores a computer program. When the computer program is executed by a processor, the lung adenocarcinoma subtype classification method provided by the embodiments of the present application is implemented.

[0008] The lung adenocarcinoma subtype classification solution provided by the present application provides a lung adenocarcinoma subtype classification model composed of a cross-modal encoding module, an attention fusion module, and a classification module. When using this lung adenocarcinoma subtype classification model for lung adenocarcinoma subtype classification, first, the electronic health record and the lung computed tomography image of the current patient are obtained, and the lung adenocarcinoma reference region in the lung computed tomography image is determined. Then, according to the lung adenocarcinoma reference region, the cross-modal encoding module performs cross-modal encoding on the electronic health record and the lung computed tomography image to obtain multimodal features, and the attention fusion module performs attention fusion on the multimodal features to obtain cross-modal fusion features. Finally, the classification module classifies the cross-modal fusion features into lung adenocarcinoma subtypes to obtain a classification result. In this way, the automatic classification of lung adenocarcinoma subtypes is realized, which can significantly improve the classification efficiency compared with the classification method based on manual recognition. Moreover, by using the lung adenocarcinoma reference region as a reference for cross-modal encoding during the encoding process, the feature expression of the lung adenocarcinoma reference region can be enhanced, and at the same time, multimodal features representing the medical image dimension and the clinical information dimension can be obtained. Further, by enhancing the attention of the multimodal features, the part of the features related to the lung adenocarcinoma subtype classification can be enhanced, and the part of the features unrelated to the lung adenocarcinoma subtype classification can be weakened, thereby ensuring the classification quality of the lung adenocarcinoma subtype classification, which can significantly improve the classification quality compared with the classification method based on manual recognition. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained according to these drawings without creative efforts.

[0010] Figure 1 It is a schematic diagram of the application environment of the model training system provided by an embodiment of the present application; Figure 2 It is a schematic flow diagram of a method for classifying subtypes of lung adenocarcinoma provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the architecture of a lung adenocarcinoma subtype classification model provided by an embodiment of the present application; Figure 4 It is Figure 3 a schematic diagram of the architecture of the cross-modal encoding module in Figure 5 It is Figure 4 a schematic diagram of the architecture of the image encoding branch in Figure 6 It is Figure 4 a schematic diagram of the architecture of the text encoding branch in Figure 7 It is a schematic diagram of the architecture of the channel attention enhancement unit in the attention fusion module; Figure 8 It is a schematic diagram of the structure of a lung adenocarcinoma subtype classification device provided by an embodiment of the present application; Figure 9 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0011] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0012] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0013] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0014] As used in the specification and claims of this application, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0015] In addition, in the description of the specification and claims of this application, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0016] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0017] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or posterior. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0018] It should be noted that artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0019] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes Machine Learning (ML) technology. Among them, DeepLearning (DL) is a new research direction in machine learning, which is introduced into machine learning to make it closer to the original goal, that is, artificial intelligence. Currently, deep learning is mainly applied in fields such as machine vision and natural language processing. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Using deep learning technology and the corresponding training sets, network models with different functions can be trained. For example, taking the generative model as an example, based on different types of training sets, generative models that can generate different types of content can be trained, such as generative models that can generate images, generative models that can generate text, and generative models that can generate voices, etc.

[0020] Computer Vision Technology (Computer Vision, CV) Computer vision is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition and measurement on targets, and further perform graphics processing to make the computer process into images that are more suitable for human eyes to observe or transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0021] This application mainly relates to the field of computer vision technology in artificial intelligence technology, and provides a method for classifying subtypes of lung adenocarcinoma, a device for classifying subtypes of lung adenocarcinoma, a computer device, and a computer-readable storage medium. Among them, the method for classifying subtypes of lung adenocarcinoma can be executed by the device for classifying subtypes of lung adenocarcinoma, or by a computer device integrated with the device for classifying subtypes of lung adenocarcinoma.

[0022] To illustrate the technical solution of this application, the following will be described through specific embodiments.

[0023] Please refer to Figure 1 , this application also provides a system for classifying subtypes of lung adenocarcinoma. The system for classifying subtypes of lung adenocarcinoma includes a computer device 100, which is used to execute the method for classifying subtypes of lung adenocarcinoma provided by this application. The computer device 100 can be any device configured with a processor and having processing capabilities, such as a desktop computer, a server, and other devices. Exemplarily, the computer device 100 first obtains the electronic health record and the lung computed tomography image of the current patient, and determines the reference region of lung adenocarcinoma in the lung computed tomography image. Then, according to the reference region of lung adenocarcinoma, the cross-modal encoding module of the lung adenocarcinoma subtype classification model performs cross-modal encoding on the obtained electronic health record and lung computed tomography image to obtain multi-modal features in the same feature space. Further, the attention fusion module of the lung adenocarcinoma subtype classification model performs attention fusion on the obtained multi-modal features to obtain cross-modal fusion features. Finally, the classification module of the lung adenocarcinoma subtype classification model performs classification of the subtypes of lung adenocarcinoma on the cross-modal fusion features to obtain a classification result.

[0024] In addition, as Figure 1 shown, the system for classifying subtypes of lung adenocarcinoma may further include a memory 200, which is used to store relevant data during the process of classifying subtypes of lung adenocarcinoma, such as the obtained electronic health record, the lung computed tomography image, the determined reference region of lung adenocarcinoma, the multi-modal features obtained by cross-modal encoding, the cross-modal fusion features obtained by attention fusion, and the classification result obtained by classifying the subtypes of lung adenocarcinoma, and so on.

[0025] It should be noted that the system for classifying subtypes of lung adenocarcinoma described above is only an example, which is used to more clearly illustrate the technical solution of the embodiments of this application, and does not constitute a limitation on the technical solution provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of the system for classifying subtypes of lung adenocarcinoma and the emergence of new business scenarios, the technical solution provided by the embodiments of this application is equally applicable to similar technical problems.

[0026] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a method for classifying subtypes of lung adenocarcinoma provided by an embodiment of this application. As Figure 2As shown, the process of this lung adenocarcinoma subtype classification method can be as follows: In 110, obtain the electronic health record and lung computed tomography (CT) image of the current patient.

[0027] The current patient refers to the patient who needs to undergo lung adenocarcinoma subtype classification, rather than a specific patient.

[0028] The electronic health record refers to the collection of information related to the patient's personal health managed in an electronic manner, which may include the following types of text information: Basic information such as the patient's name, gender, age, etc.; Medical history information such as the patient's previous disease diagnoses, treatment processes, hospitalization experiences, surgical histories, allergy histories, etc. For example, information such as the patient having had pneumonia, undergoing an appendectomy, and being allergic to penicillin will be recorded; Physiological index information such as the patient's height, weight, blood pressure, blood sugar, blood lipid, etc.; Medication record information such as the names, dosages, medication times, and medication frequencies of the medications used by the patient; Examination result information of various examinations such as the patient's blood tests, urine tests, and imaging examinations.

[0029] The computed tomography (CT) image refers to the cross-sectional image of the internal structure of the human body reconstructed by a computer after performing a tomographic scan of a certain part of the patient by a computed tomography scanning device. The computed tomography (CT) image can clearly show the fine structures inside the human body, such as nodules in the lungs, blood vessels in the brain, and so on.

[0030] In the embodiment of the present application, first, the electronic health record and lung computed tomography (CT) image of the current patient are obtained. For example, when obtaining the electronic health record of the current patient, the electronic health record of the current patient can be obtained from a database that maintains the electronic health records of different patients; when obtaining the lung computed tomography (CT) image of the current patient, the lung computed tomography (CT) image of the current patient can be obtained from a database that maintains the computed tomography (CT) images of different patients, or the lung computed tomography (CT) image of the current patient can be obtained by performing a real-time scan using a computed tomography scanning device.

[0031] In 120, determine the lung adenocarcinoma reference region in the lung computed tomography (CT) image.

[0032] As described above, after obtaining the computed tomography (CT) image of the lungs of the current patient, the reference region of lung adenocarcinoma in the CT image of the lungs is determined according to the configured region determination strategy. The reference region of lung adenocarcinoma can be understood as a possible rectangular lesion region with lung adenocarcinoma in the aforementioned CT image of the lungs. It should be noted that there is no specific limitation on the configuration of the region determination strategy in the embodiments of the present application. Among them, the reference region of lung adenocarcinoma can be represented in the form of relative positions corresponding to the CT image of the lungs. For example, the reference region of lung adenocarcinoma can be represented as (X, Y, Z, W, H, D), where X represents the X-axis coordinate of the upper left vertex of the reference region of lung adenocarcinoma in the CT image of the lungs, Y represents the Y-axis coordinate of the upper left vertex of the reference region of lung adenocarcinoma in the CT image of the lungs, Z represents the Z-axis coordinate of the upper left vertex of the reference region of lung adenocarcinoma in the CT image of the lungs, W represents the width of the reference region of lung adenocarcinoma, H represents the height of the reference region of lung adenocarcinoma, and D represents the depth of the reference region of lung adenocarcinoma.

[0033] As an optional way to configure the region determination strategy, the template matching method can be used to determine the reference region of lung adenocarcinoma. Among them, the image region with lung nodules is matched from the CT image of the lungs by using the nodule template corresponding to the lung nodules, and this image region is determined as the reference region of lung adenocarcinoma in the CT image of the lungs.

[0034] As another optional way to configure the region determination strategy, the manual marking method can be used to determine the reference region of lung adenocarcinoma. Among them, the obtained CT image of the lungs can be displayed on the display interface, and the region marking operation for the displayed CT image of the lungs is received, and the reference region of lung adenocarcinoma in the CT image of the lungs is determined according to the received region marking operation. There is no limitation on the input method of the region marking operation here. For example, when the computer device is equipped with a touch screen, the input region marking operation can be received through the touch screen. When the computer device is equipped with external input devices such as a mouse or a keyboard, the input region marking operation can also be received through the mouse or keyboard and other external input devices configured by the computer device.

[0035] As yet another optional way to configure the region determination strategy, the object detection method can be used to determine the reference region of lung adenocarcinoma. Among them, the object detection model corresponding to lung adenocarcinoma can be used to perform object detection on the CT image of the lungs to obtain the reference region of lung adenocarcinoma in the CT image of the lungs.

[0036] In 130, according to the reference region of lung adenocarcinoma, the electronic health record and the CT image of the lungs are cross-modally encoded by the cross-modal encoding module to obtain multi-modal features.

[0037] Please refer to Figure 3 , in the embodiment of the present application, a lung adenocarcinoma subtype classification model for lung adenocarcinoma subtype classification is pre-trained. As Figure 3 shown, the lung adenocarcinoma subtype classification model may include a cross-modal encoding module, an attention fusion module, and a classification module.

[0038] Among them, the cross-modal encoding module is configured to perform cross-modal encoding on the input lung adenocarcinoma reference region, pulmonary computed tomography image, and electronic health record to obtain multi-modal features in the same feature space. The attention fusion module is configured to fuse the multi-modal features obtained by the cross-modal encoding module of the cross-modal encoding module by using an attention mechanism to further obtain cross-modal fusion features. The classification module is configured to further perform classification processing on the cross-modal fusion features fused by the attention fusion module, map the cross-modal fusion features to the original output of the number size of the lung adenocarcinoma subtype categories first, and then convert the original output into the probability distribution of the lung adenocarcinoma subtype categories to represent the probabilities of different lung adenocarcinoma subtype categories. Finally, according to this probability distribution, the classification result corresponding to the input electronic health record, pulmonary computed tomography image, and lung adenocarcinoma reference region can be obtained, that is, the lung adenocarcinoma subtype category with the highest probability.

[0039] Correspondingly, in the embodiment of the present application, after obtaining the electronic health record and pulmonary computed tomography image of the current patient, and determining the lung adenocarcinoma reference region in the pulmonary computed tomography image, the electronic health record, pulmonary computed tomography image, and lung adenocarcinoma reference region of the current patient are further input into the cross-modal encoding module of the lung adenocarcinoma subtype classification model, so as to perform cross-modal encoding on the electronic health record and pulmonary computed tomography image by the cross-modal encoding module according to the lung adenocarcinoma reference region to obtain multi-modal features in the same feature space.

[0040] Optionally, in one embodiment, the cross-modal encoding module includes a text encoding branch and an image encoding branch. According to the lung adenocarcinoma reference region, performing cross-modal encoding on the electronic health record and pulmonary computed tomography image by the cross-modal encoding module to obtain multi-modal features includes: According to the lung adenocarcinoma reference region, encoding the computed tomography image by the image encoding branch to obtain image modality features; Encoding the electronic health record by the text encoding branch to obtain text modality features.

[0041] In the embodiment of the present application, please refer to Figure 4, the cross-modal encoding module includes encoding branches for two modalities, namely the text encoding branch for the text modality and the image encoding branch for the image modality. Among them, the text encoding branch is configured to encode the input electronic health record to obtain the text modality features of the electronic health record; the image encoding branch is configured to encode the computed tomography image according to the lung adenocarcinoma reference region to obtain the image modality features of the computed tomography image.

[0042] Correspondingly, when cross-modal encoding the electronic health record and the lung computed tomography image through the cross-modal encoding module according to the lung adenocarcinoma reference region to obtain multi-modal features, on the one hand, the lung adenocarcinoma reference region and the lung computed tomography image can be input into the image encoding branch of the cross-modal encoding module together, so that according to the lung adenocarcinoma reference region, the lung computed tomography image can be encoded by the image encoding branch to obtain the image modality features that focus on the lung adenocarcinoma reference region and represent the medical image dimension; on the other hand, the electronic health record can be input into the text encoding branch of the cross-modal encoding module, so that the electronic health record can be encoded by the text encoding branch to obtain the text modality features that represent the features of the clinical information dimension.

[0043] Exemplarily, please refer to Figure 5 , the image encoding branch may include a first embedding unit, a second embedding unit, a first concatenation unit, a first position encoding unit, and a visual encoding unit. The first embedding unit is configured to divide the input lung computed tomography image into multiple non-overlapping image patches according to a certain size, and perform a projection transformation on each image patch, so as to embed the input lung computed tomography image into a vector representation; the second embedding unit is configured to perform a projection transformation on the lung adenocarcinoma reference region in the form of relative position, so as to embed the lung adenocarcinoma reference region into a vector representation; the first concatenation unit is configured to concatenate (Concat) the vector units output by the first embedding unit and the second embedding unit, that is, concatenate the vector representations output by the first embedding unit and the second embedding unit in the channel dimension to obtain the concatenated vector of the two; the first position encoding unit is configured to perform position encoding on the concatenated vector output by the first concatenation unit to obtain the concatenated vector with additional position information; the visual encoding unit is configured to encode the concatenated vector with additional position information input by the first position encoding unit to obtain the image modality features. Among them, the visual encoding unit can be composed of multiple vision transformer blocks. The number of vision transformer blocks can be selected by those skilled in the art according to actual needs and will not be specifically limited here. For example, 6 vision transformer blocks can be used to form the visual encoding unit.

[0044] Correspondingly, when encoding the lung computed tomography image through the image encoding branch according to the lung adenocarcinoma reference region, the lung computed tomography image can be embedded into a vector representation through the first embedding unit, the lung adenocarcinoma reference region can be embedded into a vector representation through the second embedding unit, and the vector representations of the lung computed tomography image and the lung adenocarcinoma reference region are concatenated through the first concatenation unit to obtain a concatenated vector. Then, the concatenated vector is position-encoded through the first position encoding unit to obtain a concatenated vector with position information added. Finally, the concatenated vector with position information added is encoded through the visual encoding unit to obtain the image modality feature.

[0045] Please refer to Figure 6 , the text encoding branch may include a third embedding unit, a second position encoding unit, and a text encoding unit (which may be composed of a linear layer and a rectified linear unit (ReLu) layer). The third embedding unit is configured to perform a projection transformation on the input electronic health record, so as to embed the electronic health record into a vector representation; the second position encoding unit is configured to perform position encoding on the vector representation output by the third embedding unit, so as to add position information to the vector representation; the text encoding unit is configured to encode the vector representation with position information added output by the second position encoding unit to obtain the text modality feature.

[0046] Correspondingly, when encoding the electronic health record through the text encoding branch, the electronic health record can be first embedded into a vector representation through the third embedding unit, then the vector representation of the electronic health record is position-encoded through the second position encoding unit, and finally the vector representation of the electronic health record with position information added is encoded through the text encoding unit to obtain the text modality feature.

[0047] Optionally, in one embodiment, the image encoding branch includes a first image encoding branch and a second image encoding branch. Encoding the computed tomography image through the image encoding branch according to the lung adenocarcinoma reference region to obtain the image modality feature includes: Based on the lung adenocarcinoma reference region, a first to-be-encoded image of a first size and a second to-be-encoded image of a second size are obtained from the lung computed tomography image, and the first size is greater than the second size; Encoding the first to-be-encoded image through the first image encoding branch according to the lung adenocarcinoma reference region to obtain the first image modality feature; Encoding the second to-be-encoded image through the second image encoding branch according to the lung adenocarcinoma reference region to obtain the second image modality feature.

[0048] It should be noted that in the embodiments of the present application, in order to capture richer medical image features, the image encoding branch includes two branches, namely the first image encoding branch and the second image encoding branch. The first image encoding branch is suitable for capturing global medical image features, and the second image encoding branch is suitable for capturing local medical image features. Here, both the first image encoding branch and the second image encoding branch may adopt the above Figure 5 shown architecture of the image encoding branch.

[0049] Among them, first, according to the center point of the reference region of lung adenocarcinoma, the lung computed tomography image is centrally cropped according to the first size and the second size respectively. The image cropped according to the first size is denoted as the first image to be encoded, and the image cropped according to the second size is denoted as the second image to be encoded. It should be noted that in this embodiment, the values of the first size and the second size are not specifically limited. Exemplarily, the first size can be configured as 32x32x32, and the second size can be configured as 128x128x32.

[0050] Then, according to the reference region of lung adenocarcinoma, the first image to be encoded is encoded by the first image encoding branch to obtain the first image modality feature. According to the reference region of lung adenocarcinoma, the second image to be encoded is encoded by the second image encoding branch to obtain the second image modality feature. Here, the encoding processes of the first image encoding branch and the second image encoding branch will not be elaborated, and specific reference can be made to the relevant descriptions of the encoding process of the image encoding branch in the above embodiments.

[0051] In this way, through the first image encoding branch and the second image encoding branch suitable for two different sizes, it is possible to not only focus on the local detailed features of the reference region of lung adenocarcinoma, but also understand the overall morphology of the reference region of lung adenocarcinoma and its positional relationship in the overall image, so as to obtain an effective balance between global features and local features.

[0052] In 140, the multi-modal features are subjected to attention fusion through the attention fusion module to obtain cross-modal fusion features.

[0053] As described above, after the cross-modal encoding module of the lung adenocarcinoma subtype classification model performs cross-modal encoding on a set of inputs of the current patient to obtain multi-modal features in the same feature space, further through the attention fusion module of the lung adenocarcinoma subtype classification model, the multi-modal features are subjected to attention fusion by using the configured attention mechanism to obtain cross-modal fusion features that integrate the features of the medical image dimension and the features of the clinical information dimension. Here, there is no specific limitation on the attention mechanism adopted by the attention fusion module.

[0054] Optionally, in one embodiment, subjecting the multi-modal features to attention fusion through the attention fusion module to obtain cross-modal fusion features includes: Cascade the first image modality feature, the second image modality feature, and the text modality feature through an attention fusion module to obtain a cascaded feature; Perform self-attention enhancement on the cascaded feature through an attention fusion module to obtain a self-attention enhanced feature; Perform channel attention enhancement on the self-attention enhanced feature through an attention fusion module to obtain a cross-modal fusion feature.

[0055] In the embodiment of the present application, the attention fusion module is configured to perform attention fusion on multi-modal features by using a self-attention mechanism and a channel attention mechanism, so as to obtain a cross-modal fusion feature. The attention fusion module includes three channel fusion units, a second cascade unit, a self-attention enhancement unit, and a channel attention enhancement unit. Among them, the channel fusion unit is configured to perform channel fusion on the input multi-channel feature and convert it into a single-channel feature. Here, the specific method of channel fusion is not limited. For example, channel fusion can be implemented in the form of a pooling operation; the second cascade unit is configured to cascade the single-channel features converted by the three channel fusion units to obtain a cascaded feature; the self-attention enhancement unit is configured to perform self-attention enhancement on the cascaded feature output by the second cascade unit; the channel attention enhancement unit is configured to perform channel attention enhancement on the cascaded feature after self-attention enhancement by the self-attention enhancement unit to obtain a cross-modal fusion feature.

[0056] Correspondingly, when performing attention fusion on multi-modal features through an attention fusion module to obtain a cross-modal fusion feature, first perform channel fusion on the first image modality feature, the second image modality feature, and the text modality feature through three channel fusion units respectively to obtain a single-channel first image modality feature, a single-channel second image modality feature, and a single-channel text modality feature; then, cascade the single-channel first image modality feature, the single-channel second image modality feature, and the single-channel text modality feature through the second cascade unit to obtain a cascaded feature, which includes the single-channel first image modality feature, the single-channel second image modality feature, and the single-channel text modality feature; then, perform self-attention enhancement on the cascaded feature through the self-attention enhancement unit to obtain a self-attention enhanced feature; finally, perform channel attention enhancement on the self-attention enhanced feature through the channel attention enhancement unit to obtain a cross-modal feature.

[0057] It should be noted that the self-attention enhancement unit can be composed of multiple Transformer blocks. Here, the number of Transformer blocks is not specifically limited. For example, 6 Transformer blocks can be used to form the self-attention enhancement unit.

[0058] Optionally, in one embodiment, channel attention enhancement is performed on the self-attention enhanced features through an attention fusion module to obtain cross-modal fusion features, including: For the sub-features of each channel of the self-attention enhanced features, average pooling processing and max pooling processing are performed on the sub-features through the attention fusion module to obtain the average pooling result and the max pooling result of the sub-features; and according to the average pooling result, the first channel attention weight is obtained through the attention fusion module, and according to the max pooling result, the second channel attention weight is obtained through the attention fusion module; According to the first channel attention weight and the second channel attention weight of each sub-feature, channel attention enhancement is performed on the self-attention enhanced features through the attention fusion module to obtain cross-modal fusion features.

[0059] Please refer to Figure 7 , the channel attention enhancement unit may include a max pooling layer, an average pooling layer, a weight mapping layer (sequentially composed of a linear sub-layer, a rectified linear unit sub-layer, a linear sub-layer, and an S function sub-layer), an addition layer, and a multiplication layer. Among them, for each channel sub-feature of the input features, the max pooling layer is configured to perform max pooling processing on the channel sub-feature, and the average pooling layer is configured to perform average pooling processing on the sub-feature; the weight mapping layer is configured to map the pooling results of the max pooling layer and the average pooling layer to the corresponding channel attention weights respectively; the addition layer is configured to perform an addition operation on the two channel attention weights mapped by the weight mapping layer; the multiplication layer is configured to perform a multiplication operation on the operation result of the addition layer and the channel sub-feature, so as to complete the channel attention enhancement of the input multi-channel features.

[0060] Correspondingly, when performing channel attention enhancement on the self-attention enhanced features through the attention fusion module, for each channel sub-feature of the self-attention enhanced features, max pooling processing is performed on the sub-feature through the max pooling layer to obtain the max pooling result of the sub-feature, and average pooling processing is performed on the sub-feature through the average pooling layer to obtain the average pooling result of the sub-feature; the average pooling result of the sub-feature is mapped to the corresponding channel attention weight through the weight mapping layer, denoted as the first channel attention weight, and the max pooling result of the sub-feature is mapped to the corresponding channel attention weight through the weight mapping layer, denoted as the second channel attention weight; the addition layer performs an addition operation on the first channel attention weight and the second channel attention weight of the sub-feature to obtain the target channel attention weight of the sub-feature; the multiplication layer performs a multiplication operation on the sub-feature and the target channel attention weight to obtain the channel attention enhanced feature of the sub-feature. In this way, by performing channel attention enhancement on the self-attention enhanced features, cross-modal fusion features can be obtained.

[0061] In 150, the classification module classifies the cross-modal fusion features for the lung adenocarcinoma subtypes to obtain the classification results.

[0062] As described above, after the attention fusion module performs attention fusion on the multi-modal features to obtain the cross-modal fusion features, the classification module of the lung adenocarcinoma subtype classification model further classifies the cross-modal fusion features to obtain the classification results corresponding to a set of inputs of the current patient, and the classification results describe the lung adenocarcinoma subtype category of the current patient.

[0063] Exemplarily, the classification module may include a linear layer and a softmax processing layer. The linear layer is configured to map the cross-modal fusion features to the raw output of the number size of the lung adenocarcinoma subtype categories, and the softmax processing layer is configured to convert the raw output into the probability distribution of the lung adenocarcinoma subtype categories. Assume that there are three lung adenocarcinoma subtype categories: adenocarcinoma in situ, minimally invasive adenocarcinoma, and invasive adenocarcinoma. The probability distribution output by the softmax processing layer corresponds to the three lung adenocarcinoma subtype categories of adenocarcinoma in situ, minimally invasive adenocarcinoma, and invasive adenocarcinoma in sequence. For the lung adenocarcinoma reference region, the pulmonary computed tomography image, and the electronic health record of the current patient, if the probability distribution output by the softmax layer is {'0.1', '0.6', '0.3'}, it can be determined that the classification result of the current patient is "minimally invasive adenocarcinoma".

[0064] Optionally, in one embodiment, an optional training method for the lung adenocarcinoma subtype classification model is also provided: Obtain the sample electronic health record of the sample patient and obtain the sample pulmonary computed tomography image of the sample patient; Obtain the lung adenocarcinoma subtype category label of the sample pulmonary computed tomography image and determine the sample lung adenocarcinoma reference region of the sample pulmonary computed tomography image; Based on the sample lung adenocarcinoma reference region, obtain the first sample image to be encoded with the first size and the second sample image to be encoded with the second size from the sample pulmonary computed tomography image; Based on the sample lung adenocarcinoma reference region, encode the first sample image to be encoded through the first image encoding branch to obtain the first sample image modality features; and based on the sample lung adenocarcinoma reference region, encode the second sample image to be encoded through the second image encoding branch to obtain the second sample image modality features; Encode the sample electronic health record through the text encoding branch to obtain the sample text modality features; Perform attention fusion on the sample text modality features, the first sample image modality features, and the second sample image modality features through the attention fusion module to obtain the sample cross-modal fusion features; The lung adenocarcinoma subtype classification is performed on the sample cross-modal fusion features through a classification module to obtain the sample classification result; According to the sample classification result and the lung adenocarcinoma subtype category label, the focal loss is obtained; the first similarity between the sample text modality feature and the first sample image modality feature is obtained, and the first contrast loss is obtained according to the first similarity; and the second similarity between the sample text modality feature and the second image encoding modality feature is obtained, and the second contrast loss is obtained according to the second similarity; Determine the respective fusion weights of the first contrast loss, the second contrast loss, and the focal loss, and obtain the fusion loss of the first contrast loss, the second contrast loss, and the focal loss according to the respective fusion weights of the first contrast loss, the second contrast loss, and the focal loss; Update the model parameters of at least one of the cross-modal encoding module, the attention fusion module, and the classification module according to the fusion loss.

[0065] It should be noted that the process of obtaining the sample classification result of the sample patient through the lung adenocarcinoma subtype classification model can refer to the relevant descriptions in the above embodiments, and will not be elaborated here.

[0066] In addition, the embodiments of the present application do not specifically limit which contrast loss function is used to measure the above first contrast loss and second contrast loss. For example, the information noise contrast estimation loss can be used to measure the first contrast loss and the second contrast loss.

[0067] Among them, when updating the model parameters of at least one of the cross-modal encoding module, the attention fusion module, and the classification module according to the fusion loss, the Adam optimizer can be used to update the model parameters of at least one of the cross-modal encoding module, the attention fusion module, and the classification module along the gradient descent direction until the preset stop condition is met. For example, the preset stop condition can be configured as: the number of iterative updates of the model parameters reaches the number threshold, or the fusion loss converges, etc.

[0068] In order to verify the model performance of the lung adenocarcinoma subtype classification model provided in the embodiments of the present application, the following comparative experiment was carried out: 1614 groups of lung adenocarcinoma cases of 1430 anonymous patients were collected. Each group of cases included pulmonary computed tomography images, electronic health cascades, and lung adenocarcinoma reference regions. These cases were divided into three categories: adenocarcinoma in situ, minimally invasive adenocarcinoma, and invasive adenocarcinoma, and the proportion of each subtype category was 53.5%, 24.1%, and 22.4% respectively.

[0069] Based on the above case data, the lung adenocarcinoma subtype classification model provided in this application, the pure text classification model (lung adenocarcinoma subtype classification is performed only using electronic health records, and the architecture is equivalent to the above text encoding branch + classification module), and the pure visual classification model (including three groups: using 32-size lung computed tomography images for lung adenocarcinoma subtype classification, the architecture is equivalent to the above second image encoding branch + classification module; using 128-size lung computed tomography images for lung adenocarcinoma subtype classification, the architecture is equivalent to the above first image encoding branch + classification module; using 32-size and 128-size lung computed tomography images for lung adenocarcinoma subtype classification, the architecture is equivalent to the above first image encoding branch + second image encoding branch + classification module), respectively, are used. The comparative experimental results are shown in Table 1 (Experimental Results Comparison Table): According to Table 1 above, the lung adenocarcinoma subtype classification model provided by this application achieves the best performance in the above indicators. Compared with the pure visual classification model (32 dimensions + 128 dimensions), the accuracy rate is increased by 10.83 percentage points, the precision rate is increased by 13.41 percentage points, the recall rate is increased by 14.26 percentage points, the F1 score is increased by 14.69 percentage points, and the AUC is increased by 0.0763. Obviously, the lung adenocarcinoma subtype classification model adopted by this application has excellent lung adenocarcinoma subtype classification capabilities.

[0070] Optionally, in one embodiment, determining the fusion weights of the first contrast loss, the second contrast loss, and the focus loss includes: calculating a first loss sum value of a first contrast loss and a second contrast loss, and calculating a second loss sum value of the first contrast loss, the second contrast loss, and the focus loss; Calculating a quotient between the first loss sum and the second loss sum, and determining the quotient as a fusion weight of the first contrast loss and the second contrast loss; The difference between the preset weight coefficient and the quotient is calculated, and the difference is determined as the fusion weight of the focal loss.

[0071] It should be noted that the embodiment of the present application does not impose any specific restriction on the value of the preset weight coefficient. For example, the value can be 1.

[0072] The above process of determining the fusion weights of the first contrast loss and the second contrast loss can be expressed as: ; in, represents the fusion weight of the first contrast loss and the second contrast loss, represents the first contrast loss, represents the second contrast loss, represents focal loss.

[0073] The process of obtaining the combined loss of the first contrast loss, the second contrast loss, and the focal loss according to their respective fusion weights of the first contrast loss, the second contrast loss, and the focal loss can be expressed as: ; wherein, represents the combined loss, represents a preset weight coefficient.

[0074] As can be seen from the above, the lung adenocarcinoma subtype classification scheme provided by the present application provides a lung adenocarcinoma subtype classification model composed of a cross-modal encoding module, an attention fusion module, and a classification module. When using this lung adenocarcinoma subtype classification model for lung adenocarcinoma subtype classification, first, the electronic health record and the pulmonary computed tomography image of the current patient are obtained, and the reference region of lung adenocarcinoma in the pulmonary computed tomography image is determined. Then, according to the reference region of lung adenocarcinoma, the cross-modal encoding module performs cross-modal encoding on the electronic health record and the pulmonary computed tomography image to obtain multi-modal features, and the attention fusion module performs attention fusion on the multi-modal features to obtain cross-modal fusion features. Finally, the classification module performs lung adenocarcinoma subtype classification on the cross-modal fusion features to obtain a classification result. In this way, the automatic classification of lung adenocarcinoma subtypes is realized, which can significantly improve the classification efficiency compared with the classification method based on manual recognition. Moreover, by using the reference region of lung adenocarcinoma as a reference for cross-modal encoding during the encoding process, the feature expression of the reference region of lung adenocarcinoma can be enhanced, and at the same time, multi-modal features representing the medical image dimension and the clinical information dimension can be obtained. Further, the attention enhancement of the multi-modal features can enhance some features related to lung adenocarcinoma subtype classification and weaken some features unrelated to lung adenocarcinoma subtype classification, thereby ensuring the classification quality of lung adenocarcinoma subtype classification and significantly improving the classification quality compared with the classification method based on manual recognition.

[0075] To facilitate better implementation of the above lung adenocarcinoma subtype classification method, the embodiment of the present application further provides a corresponding lung adenocarcinoma subtype classification device. The meanings of the nouns are the same as those in the above lung adenocarcinoma subtype classification method, and the specific implementation details can be referred to the description in the above method embodiments.

[0076] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the lung adenocarcinoma subtype classification device provided by the embodiment of the present application. The lung adenocarcinoma subtype classification device may include a data acquisition module 210, a region determination module 220, a feature encoding module 230, a feature fusion module 240, and a feature classification module 250. Among them, The data acquisition module 210 is configured to acquire the electronic health record and the pulmonary computed tomography image of the current patient; A region determination module 220, configured to determine a reference region of lung adenocarcinoma in a pulmonary computed tomography image; A feature encoding module 230, configured to perform cross-modal encoding on an electronic health record and a pulmonary computed tomography image through a cross-modal encoding module according to the reference region of lung adenocarcinoma to obtain multi-modal features; A feature fusion module 240, configured to perform attention fusion on the multi-modal features through an attention fusion module to obtain cross-modal fusion features; A feature classification module 250, configured to perform lung adenocarcinoma subtype classification on the cross-modal fusion features through a classification module to obtain a classification result.

[0077] Optionally, in an embodiment, the cross-modal encoding module includes a text encoding branch and an image encoding branch. The feature encoding module 230 is configured to encode the computed tomography image through the image encoding branch according to the reference region of lung adenocarcinoma to obtain image modality features; and encode the electronic health record through the text encoding branch to obtain text modality features.

[0078] Optionally, in an embodiment, the image encoding branch includes a first image encoding branch and a second image encoding branch. The feature encoding module 230 is configured to obtain a first to-be-encoded image with a first size and a second to-be-encoded image with a second size based on the pulmonary computed tomography image according to the reference region of lung adenocarcinoma, where the first size is greater than the second size; encode the first to-be-encoded image through the first image encoding branch according to the reference region of lung adenocarcinoma to obtain first image modality features; and encode the second to-be-encoded image through the second image encoding branch according to the reference region of lung adenocarcinoma to obtain second image modality features.

[0079] Optionally, in an embodiment, the feature fusion module 240 is configured to cascade the first image modality features, the second image modality features, and the text modality features through the attention fusion module to obtain cascaded features; perform self-attention enhancement on the cascaded features through the attention fusion module to obtain self-attention enhanced features; and perform channel attention enhancement on the self-attention enhanced features through the attention fusion module to obtain cross-modal fusion features.

[0080] Optionally, in one embodiment, the feature fusion module 240 is configured to, for the sub-features of each channel of the self-attention enhanced feature, perform average pooling processing and max pooling processing on the sub-feature through the attention fusion module to obtain the average pooling result and the max pooling result of the sub-feature; and according to the average pooling result, obtain the first channel attention weight through the attention fusion module, and according to the max pooling result, obtain the second channel attention weight through the attention fusion module; according to the first channel attention weight and the second channel attention weight of each sub-feature, perform channel attention enhancement on the self-attention enhanced feature through the attention fusion module to obtain the cross-modal fusion feature.

[0081] Optionally, in one embodiment, the lung adenocarcinoma subtype classification device provided in this application further includes a model training module, which is configured to obtain the sample electronic health records of the sample patients, and obtain the sample lung computed tomography images of the sample patients; obtain the lung adenocarcinoma subtype category labels of the sample lung computed tomography images, and determine the sample lung adenocarcinoma reference regions of the sample lung computed tomography images; based on the sample lung adenocarcinoma reference regions, obtain the first sample image to be encoded with the first size and the second sample image to be encoded with the second size from the sample lung computed tomography images; according to the sample lung adenocarcinoma reference regions, encode the first sample image to be encoded through the first image encoding branch to obtain the first sample image modality feature; and according to the sample lung adenocarcinoma reference regions, encode the second sample image to be encoded through the second image encoding branch to obtain the second sample image modality feature; encode the sample electronic health records through the text encoding branch to obtain the sample text modality feature; perform attention fusion on the sample text modality feature, the first sample image modality feature and the second sample image modality feature through the attention fusion module to obtain the sample cross-modal fusion feature; perform lung adenocarcinoma subtype classification on the sample cross-modal fusion feature through the classification module to obtain the sample classification result; obtain the focal loss according to the sample classification result and the lung adenocarcinoma subtype category label; obtain the first similarity between the sample text modality feature and the first sample image modality feature, and obtain the first contrast loss according to the first similarity; and obtain the second similarity between the sample text modality feature and the second image encoding modality feature, and obtain the second contrast loss according to the second similarity; determine the fusion weights of the first contrast loss, the second contrast loss and the focal loss respectively, and according to the fusion weights of the first contrast loss, the second contrast loss and the focal loss respectively, obtain the fusion loss of the first contrast loss, the second contrast loss and the focal loss; update the model parameters of at least one of the cross-modal encoding module, the attention fusion module and the classification module according to the fusion loss.

[0082] Optionally, in one embodiment, the model training module is used to calculate the first loss sum of the first contrast loss and the second contrast loss, and calculate the second loss sum of the first contrast loss, the second contrast loss and the focal loss; calculate the quotient between the first loss sum and the second loss sum, and determine the quotient as the fusion weight of the first contrast loss and the second contrast loss; calculate the difference between the preset weight coefficient and the quotient, and determine the difference as the fusion weight of the focal loss.

[0083] For the specific implementation of each module of the above lung adenocarcinoma subtype classification device, reference can be made to the previous embodiments, which will not be elaborated here.

[0084] Please refer to Figure 9 , this embodiment of the present application also provides a computer device, such as Figure 9 shown, the computer device includes: at least one processor ( Figure 9 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor, and when the processor executes the computer program, the steps in the above embodiment of the lung adenocarcinoma subtype classification method are implemented.

[0085] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 9 this is only an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.

[0086] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0087] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of a computer device, and in some other embodiments, it can also be an external storage device of a computer device, for example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, a boot loader (BootLoader), data, and other programs, etc., and the other programs such as the program code of a computer program. The memory can also be used to temporarily store the data that has been output or will be output.

[0088] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0089] All or part of the processes in the above-mentioned method embodiments of this application can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above-mentioned method embodiments.

[0090] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0091] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0092] In the embodiments provided in this application, it should be understood that the disclosed devices / computer devices and methods can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical or other form.

[0093] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0094] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A method for classifying subtypes of lung adenocarcinoma, characterized in that, The lung adenocarcinoma subtype classification model includes a cross-modal encoding module, an attention fusion module, and a classification module. The lung adenocarcinoma subtype classification method includes: Obtain the electronic health record and the lung computed tomography image of the current patient; Determine the lung adenocarcinoma reference region in the lung computed tomography image; According to the lung adenocarcinoma reference region, perform cross-modal encoding on the electronic health record and the lung computed tomography image through the cross-modal encoding module to obtain multi-modal features; Perform attention fusion on the multi-modal features through the attention fusion module to obtain cross-modal fusion features; Perform lung adenocarcinoma subtype classification on the cross-modal fusion features through the classification module to obtain a classification result.

2. The lung adenocarcinoma subtype classification method according to claim 1, characterized in that The cross-modal encoding module includes a text encoding branch and an image encoding branch. The step of performing cross-modal encoding on the electronic health record and the lung computed tomography image through the cross-modal encoding module according to the lung adenocarcinoma reference region to obtain multi-modal features includes: According to the lung adenocarcinoma reference region, encode the computed tomography image through the image encoding branch to obtain image modal features; Encode the electronic health record through the text encoding branch to obtain text modal features.

3. The lung adenocarcinoma subtype classification method according to claim 2, characterized in that, The image encoding branch includes a first image encoding branch and a second image encoding branch. The step of encoding the computed tomography image through the image encoding branch according to the lung adenocarcinoma reference region to obtain image modal features includes: According to the lung adenocarcinoma reference region, obtain a first to-be-encoded image with a first size and a second to-be-encoded image with a second size based on the lung computed tomography image, where the first size is larger than the second size; According to the lung adenocarcinoma reference region, encode the first to-be-encoded image through the first image encoding branch to obtain first image modal features; According to the lung adenocarcinoma reference region, encode the second to-be-encoded image through the second image encoding branch to obtain second image modal features.

4. The lung adenocarcinoma subtype classification method according to claim 3, wherein, The step of performing attention fusion on the multi-modal features through the attention fusion module to obtain cross-modal fusion features includes: Cascade the first image modal features, the second image modal features, and the text modal features through the attention fusion module to obtain cascaded features; Perform self-attention enhancement on the cascaded features through the attention fusion module to obtain self-attention enhanced features; Perform channel attention enhancement on the self-attention enhanced features through the attention fusion module to obtain the cross-modal fusion features.

5. The method for classifying subtypes of lung adenocarcinoma according to claim 4, characterized in that, The step of performing channel attention enhancement on the self-attention enhanced features through the attention fusion module to obtain the cross-modal fusion features includes: For each sub - feature of each channel of the self - attention enhanced feature, average pooling and max pooling are performed on the sub - feature through the attention fusion module to obtain the average pooling result and the max pooling result of the sub - feature; and according to the average pooling result, the first - channel attention weight is obtained through the attention fusion module, and according to the max pooling result, the second - channel attention weight is obtained through the attention fusion module. According to the first - channel attention weight and the second - channel attention weight of each sub - feature, the channel attention of the self - attention enhanced feature is enhanced through the attention fusion module to obtain the cross - modal fusion feature.

6. The lung adenocarcinoma subtype classification method according to claim 3, wherein The lung adenocarcinoma subtype classification model is trained as follows: Obtain the sample electronic health record of the sample patient and obtain the sample lung computed tomography image of the sample patient; Obtain the lung adenocarcinoma subtype category label of the sample lung computed tomography image and determine the sample lung adenocarcinoma reference region of the sample lung computed tomography image; Based on the sample lung adenocarcinoma reference region, obtain the first - sized first sample image to be encoded and the second - sized second sample image to be encoded from the sample lung computed tomography image; According to the sample lung adenocarcinoma reference region, encode the first sample image to be encoded through the first image encoding branch to obtain the first sample image modality feature; And according to the sample lung adenocarcinoma reference region, encode the second sample image to be encoded through the second image encoding branch to obtain the second sample image modality feature; Encode the sample electronic health record through the text encoding branch to obtain the sample text modality feature; Perform attention fusion on the sample text modality feature, the first sample image modality feature, and the second sample image modality feature through the attention fusion module to obtain the sample cross - modal fusion feature; Perform lung adenocarcinoma subtype classification on the sample cross - modal fusion feature through the classification module to obtain the sample classification result; Obtain the focal loss according to the sample classification result and the lung adenocarcinoma subtype category label; Obtain the first similarity between the sample text modality feature and the first sample image modality feature and obtain the first contrast loss according to the first similarity; and obtain the second similarity between the sample text modality feature and the second image encoding modality feature and obtain the second contrast loss according to the second similarity; Determine the fusion weights of the first contrast loss, the second contrast loss, and the focal loss respectively, and obtain the fusion loss of the first contrast loss, the second contrast loss, and the focal loss according to the fusion weights of the first contrast loss, the second contrast loss, and the focal loss respectively; Update the model parameters of at least one of the cross - modal encoding module, the attention fusion module, and the classification module according to the fusion loss.

7. The method for classifying subtypes of lung adenocarcinoma according to claim 6, wherein The determining the fusion weights of the first contrast loss, the second contrast loss, and the focal loss respectively includes: Calculate the first loss sum value of the first contrast loss and the second contrast loss, and calculate the second loss sum value of the first contrast loss, the second contrast loss, and the focal loss; Calculate the quotient value between the first loss sum value and the second loss sum value, and determine the quotient value as the fusion weight of the first contrast loss and the second contrast loss; Calculate the difference between the preset weight coefficient and the quotient value, and determine the difference as the fusion weight of the focal loss.

8. A lung adenocarcinoma subtype classification device, characterized in that, The lung adenocarcinoma subtype classification model includes a cross-modal encoding module, an attention fusion module, and a classification module. The lung adenocarcinoma subtype classification device includes: A data acquisition module, configured to acquire the electronic health record and the pulmonary computed tomography image of the current patient; A region determination module, configured to determine the lung adenocarcinoma reference region in the pulmonary computed tomography image; A feature encoding module, configured to perform cross-modal encoding on the electronic health record and the pulmonary computed tomography image through the cross-modal encoding module according to the lung adenocarcinoma reference region to obtain multi-modal features; A feature fusion module, configured to perform attention fusion on the multi-modal features through the attention fusion module to obtain cross-modal fusion features; A feature classification module, configured to perform lung adenocarcinoma subtype classification on the cross-modal fusion features through the classification module to obtain a classification result.

9. A computer device, characterized in that, The computer device includes a processor and a memory. A computer program that can run on the processor is stored in the memory. When the processor runs the computer program, the lung adenocarcinoma subtype classification method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being executed by a processor to implement the lung adenocarcinoma subtype classification method according to any one of claims 1 to 7.