Image recognition method and computer-aided diagnosis method for esophageal varicosity level

By using the Multi-Organ Coordination Network (MOON++) model and combining NCCT images to identify the volume and morphological characteristics of the esophagus and related organs, the problem of misdiagnosis caused by physician experience dependence is solved. This enables accurate assessment of the level of esophageal varices and personalized treatment plans, reducing the cost of invasive examinations and patient discomfort.

CN120951165APending Publication Date: 2025-11-14ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510985369.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, doctors rely on experience to diagnose the grade of esophageal varices, which can lead to misdiagnosis or incorrect diagnosis. Furthermore, invasive endoscopic examinations are costly and unsuitable for patients.

Method used

The MOON++ model is used to identify the volume and morphological features of the esophagus and related organs through NCCT images. Combined with multimodal learning methods, it provides a more accurate assessment of the level of esophageal varices and avoids the use of contrast agents.

Benefits of technology

It improves the diagnostic accuracy of esophageal varices grades, reduces medical costs and patient discomfort, provides personalized treatment options, and is suitable for long-term screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951165A_ABST
    Figure CN120951165A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition method and a computer-aided diagnosis method for the esophageal varicosity level, and belongs to the technical field of artificial intelligence. The method comprises the steps that a target image is acquired, the target image is obtained by shooting multiple organs of a target human body, and the multiple organs comprise a target organ and at least one related organ adjacent to the target organ; segmenting a target organ image and at least one related organ image from the target image, and generating target text information for describing the volume of the plurality of organs; and inputting the target image, the target text information, the target organ image and at least one related organ image into an image recognition model, and outputting a target recognition result of the target organ. By means of the image recognition model, the accuracy of the recognition result of the target organ can be improved, and misdiagnosis or wrong diagnosis is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an image recognition method and a computer-aided diagnostic method for esophageal varices. Background Technology

[0002] CT (Computed Tomography) is a medical imaging technique that uses X-ray beams to perform tomographic scanning of the human body and generates detailed images of the body's internal structures using a computer. In clinical practice, doctors can diagnose whether a target organ has developed a disease by identifying the CT images of that organ.

[0003] However, due to limitations in doctors' diagnostic experience, the identification of target organs may not be accurate enough, and misdiagnosis or incorrect diagnosis may occur. Summary of the Invention

[0004] This application provides an image recognition method and a computer-aided diagnostic method for esophageal varices. By utilizing an image recognition model, the accuracy of target organ identification is improved, reducing misdiagnosis or incorrect diagnosis. The technical solution is as follows:

[0005] In a first aspect, an image recognition method is provided, the method comprising:

[0006] Acquire a target image, which is obtained by photographing multiple organs of a target human body, including the target organ and at least one related organ adjacent to the target organ;

[0007] The target organ image and at least one related organ image are segmented from the target image, and target text information describing the volume of the multiple organs is generated.

[0008] The target image, the target text information, the target organ image, and at least one related organ image are input into the image recognition model, and the target recognition result of the target organ is output.

[0009] Secondly, an image recognition method is provided, the method being applied to a cloud-side device, the method comprising:

[0010] The receiving end device sends an image processing task, the image processing task including a target image, the target image being obtained by capturing multiple organs of a target human body, the multiple organs including the target organ and at least one related organ adjacent to the target organ;

[0011] The target organ image and at least one related organ image are segmented from the target image, and target text information describing the volume of the multiple organs is generated.

[0012] The target image, the target text information, the target organ image, and at least one related organ image are input into the image recognition model, and the target recognition result of the target organ is output.

[0013] The target identification result is sent to the terminal device.

[0014] Thirdly, a method for training an image recognition model is provided, the method comprising:

[0015] Acquire sample images and an image recognition model to be trained. The sample images are obtained by photographing multiple organs of the human body. The multiple organs include the target organ and at least one related organ adjacent to the target organ. The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier.

[0016] The target organ image and at least one related organ image are segmented from the sample image, and sample text information describing the volume of the multiple organs is generated.

[0017] The sample image and the sample text information are input into the multimodal feature extraction sub-model, and the first sample feature vector is output.

[0018] The target organ image and at least one related organ image are input into the multi-organ feature extraction sub-model, and the second sample feature vector and at least one third sample feature vector are output.

[0019] The first sample feature vector, the second sample feature vector, and the at least one third sample feature vector are fused to obtain a sample fusion feature vector.

[0020] The sample fusion feature vector is input into the classifier, and the predicted identification result of the target organ is output.

[0021] The image recognition model is trained based on the estimated recognition result, the actual recognition result corresponding to the sample image, the second sample feature vector, and the at least one third sample feature vector.

[0022] Fourthly, a computer-aided diagnostic method for esophageal varices is provided, the method comprising:

[0023] A target chest image is acquired by photographing multiple organs of the target chest, including the esophagus and at least one related organ adjacent to the esophagus.

[0024] The esophagus image and at least one related organ image are segmented from the target chest image, and target text information describing the volume of the multiple organs is generated.

[0025] The target chest image, the target text information, the esophageal image, and at least one related organ image are input into the image recognition model, and the recognition result of the esophageal varices level is output.

[0026] Fifthly, an image recognition device is provided, the device comprising:

[0027] An acquisition module is used to acquire a target image, which is obtained by photographing multiple organs of a target human body, including the target organ and at least one related organ adjacent to the target organ;

[0028] A segmentation module is used to segment the target organ image and at least one related organ image from the target image;

[0029] A generation module is used to generate target text information describing the volume of the plurality of organs;

[0030] The input / output module is used to input the target image, the target text information, the target organ image, and at least one related organ image into the image recognition model, and output the target recognition result of the target organ.

[0031] Sixthly, an image recognition device is provided, the device being applied to cloud-side equipment, the device comprising:

[0032] A receiving module is used to receive an image processing task sent by an end-side device. The image processing task includes a target image, which is obtained by capturing images of multiple organs of a target human body. The multiple organs include the target organ and at least one related organ adjacent to the target organ.

[0033] A segmentation module is used to segment the target organ image and at least one related organ image from the target image;

[0034] A generation module is used to generate target text information describing the volume of the plurality of organs;

[0035] The input / output module allows the user to input the target image, the target text information, the target organ image, and at least one related organ image into the image recognition model, and outputs the target recognition result of the target organ.

[0036] The sending module is used to send the target identification result to the end-side device.

[0037] In a seventh aspect, a training apparatus for an image recognition model is provided, the apparatus comprising:

[0038] The acquisition module is used to acquire sample images and an image recognition model to be trained. The sample images are obtained by photographing multiple organs of the human body. The multiple organs include the target organ and at least one related organ adjacent to the target organ. The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier.

[0039] A segmentation module is used to segment the target organ image and at least one related organ image from the sample image;

[0040] A generation module is used to generate sample text information describing the volume of the plurality of organs;

[0041] The first input / output module is used to input the sample image and the sample text information into the multimodal feature extraction sub-model and output the first sample feature vector.

[0042] The second input / output module is used to input the sample image and at least one sample-related organ image into the multi-organ feature extraction sub-model, and output a second sample feature vector and at least one third sample feature vector.

[0043] The fusion module is used to fuse the first sample feature vector, the second sample feature vector, and the at least one third sample feature vector to obtain a sample fusion feature vector.

[0044] The third input / output module is used to input the sample fusion feature vector into the classifier and output the predicted identification result of the target organ;

[0045] The training module is used to train the image recognition model based on the estimated recognition result, the actual recognition result corresponding to the sample image, the second sample feature vector, and the at least one third sample feature vector.

[0046] Eighthly, a computer-aided diagnostic method for esophageal varices is provided, the method comprising:

[0047] An acquisition module is used to acquire a target chest image, which is obtained by photographing multiple organs of the target chest, including the esophagus and at least one related organ adjacent to the esophagus.

[0048] A segmentation module is used to segment an esophageal image and at least one related organ image from the target chest image;

[0049] A generation module is used to generate target text information describing the volume of the plurality of organs;

[0050] The input / output module is used to input the target chest image, the target text information, the esophageal image, and at least one related organ image into the image recognition model, and output the recognition result of the esophageal varicose vein level.

[0051] In a ninth aspect, an electronic device is provided, including a processor and a memory; the memory stores at least one piece of program code; the at least one piece of program code is used to be called and executed by the processor to implement the image recognition method of the first aspect, or the image recognition method of the second aspect, or the image recognition model training method of the third aspect, or the computer-aided diagnosis method for esophageal varices of the fourth aspect.

[0052] In a tenth aspect, a computer-readable storage medium is provided, wherein at least one computer program is stored therein, and when executed by a processor, the at least one computer program is capable of implementing the image recognition method of the first aspect, or the image recognition method of the second aspect, or the image recognition model training method of the third aspect, or the computer-aided diagnosis method for esophageal varices of the fourth aspect.

[0053] Eleventhly, a computer program product is provided, the computer program product comprising a computer program, which, when executed by a processor, is capable of implementing the image recognition method described in the first aspect, or the image recognition method described in the second aspect, or the image recognition model training method described in the third aspect, or the computer-aided diagnostic method for esophageal varices described in the fourth aspect.

[0054] The beneficial effects of the technical solutions provided in this application are:

[0055] This application embodiment acquires target images of multiple organs of a target human body, including the target organ and at least one adjacent related organ. This adjacent related organ provides crucial information for the diagnosis of the target organ. For example, if the target organ is the esophagus, the volume, shape, and structural features of the at least one adjacent related organ differ depending on the degree of esophageal varices. After acquiring the target image, the target organ image and at least one related organ image are segmented from the target image, and textual descriptions of the volume of the multiple organs are generated. Then, the target image, target textual information, target organ image, and at least one related organ image are input into an image recognition model, outputting the target organ's recognition result. This application embodiment does not rely on physician experience; based on an image recognition model trained using a multimodal learning method, it processes the overall image and segmented images of multiple organs, including the target organ, as well as the textual information of the organ's volume, resulting in a more accurate and reliable target organ recognition result. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This refers to the implementation environment involved in the image recognition method and image recognition model training method provided in the embodiments of this application;

[0058] Figure 2 This is a flowchart of an image recognition method provided in an embodiment of this application;

[0059] Figure 3 This is a schematic diagram of the structure of a dimensionality reduction module provided in an embodiment of this application;

[0060] Figure 4 This is a schematic diagram of the structure of an organ-related function module provided in an embodiment of this application;

[0061] Figure 5 This is a schematic diagram of the structure of a classifier provided in an embodiment of this application;

[0062] Figure 6 This is a flowchart of another image recognition method provided in the embodiments of this application;

[0063] Figure 7 This is a flowchart of a training method for an image recognition model provided in an embodiment of this application;

[0064] Figure 8 This is a flowchart of another image recognition model training method provided in the embodiments of this application;

[0065] Figure 9 A flowchart illustrating a computer-aided diagnostic method for esophageal varices according to an embodiment of this application is shown.

[0066] Figure 10 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0067] Figure 11 This is a schematic diagram of another image recognition device provided in an embodiment of this application;

[0068] Figure 12 This is a schematic diagram of the structure of a training device for an image recognition model provided in an embodiment of this application;

[0069] Figure 13 This is a schematic diagram of the structure of a computer-aided diagnostic device for esophageal varices provided in an embodiment of this application;

[0070] Figure 14 A structural block diagram of an electronic device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0072] It is understood that the terms "each," "multiple," and "any" used in the embodiments of this application, etc., mean that "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the corresponding multiples. For example, multiple words include 10 words, and "each word" refers to each of the 10 words, while "any word" refers to any one of the 10 words.

[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0074] Cirrhosis poses a global health challenge, impacting the lives of millions. Of particular concern is the common complication of cirrhosis: bleeding from ruptured esophageal varices. Esophageal varices place a heavy burden on healthcare systems, requiring emergency medical intervention, intensive care, massive blood transfusions, and surgery, and even long-term rehabilitation care, imposing a significant economic burden on individuals, families, and society as a whole. For high-risk cirrhosis patients, early identification and personalized preventative measures are crucial. Accurate diagnosis and meticulous management can significantly alleviate patients' psychological stress and fear, fundamentally improve their quality of life, and potentially even extend their life expectancy.

[0075] Esophageal varices are caused by portal hypertension, leading to blood reflux into the vulnerable veins of the lower esophagus. The severity of the varices is closely related to the risk of bleeding. Currently, in clinical practice, invasive endoscopy is used to take CT images of the patient's esophagus. Doctors then identify the severity of the varices by analyzing these CT images. However, this invasive endoscopic method requires the use of contrast agents, resulting in high medical costs and potential discomfort for patients sensitive to contrast agents, increasing the risk of infection and bleeding. Furthermore, this method relies heavily on the doctor's diagnostic experience in classifying esophageal varices, and differences in diagnostic experience among doctors can lead to misdiagnosis or incorrect diagnosis.

[0076] To improve diagnostic accuracy while reducing medical costs and patient discomfort associated with invasive endoscopy, this application presents a Multi-Organ-Cohesion Network++ (MOON++), which is the network of the image recognition model described in this application. This MOON++ network leverages prior clinical knowledge, particularly the strong correlation between the liver-spleen volume ratio and the severity of liver fibrosis, effectively integrating the imaging features of the esophagus, liver, and spleen with organ volume relationships. This provides more accurate information for assessing esophageal varices, improving the accuracy and reliability of esophageal varice grade assessment and providing doctors with a precise diagnostic tool. This helps reduce medical misdiagnosis and allows for more personalized treatment plans for patients. Furthermore, this MOON++ network can recognize NCCT (Non-Contrast CT) images. After scanning multiple organs in the patient's chest, including the esophagus, with NCCT images, the MOON++ network can identify the grade of the patient's esophageal varices. Because NCCT is a standard CT scan, it has high diagnostic accuracy and generalization ability, avoids the use of contrast agents, effectively improves patient comfort, and is more suitable for long-term screening in clinical settings. Furthermore, this multi-organ coordination network can not only identify the grade of esophageal varices but also the lesion grades of other organs. For ease of description, the diseased organs identified by the multi-organ coordination network in this application embodiment are referred to as target organs.

[0077] Please refer to Figure 1 This illustration shows the implementation environment involved in the image recognition method and image recognition model training method provided in the embodiments of this application. The implementation environment includes: a client 101 and a server 102. The client 101 and the server 102 communicate through a network, which can be a wired network or a wireless network, etc.

[0078] The client 101 can be an NCCT device. The client 101 is used to scan multiple organs of the target human body to obtain target images (i.e., NCCT images), and then sends the captured target images to the server 102. The multiple organs include the target organ and at least one related organ adjacent to the target organ. The target organ can be the esophagus, and the related organ can be the liver, spleen, etc.

[0079] Server 102 can be a device providing image recognition services. Server 102 can be an edge device with strong computing power, or it can include edge devices with relatively weak computing power and cloud devices with strong computing power. When server 102 is an edge device with strong computing power, server 102 receives the target image sent by client 101, segments the target organ image and at least one related organ image from the target image, generates target text information describing the volume of multiple organs, and then inputs the target image, target text information, target organ image, and at least one related organ image into the image recognition model, outputting the target organ recognition result. When server 102 includes edge devices with relatively weak computing power and cloud devices with strong computing power, the edge device receives the target image sent by client 101 and sends an image processing task to the cloud device, the image processing task including the target image, etc. In response to the image processing task, the cloud-side device segments the target organ image and at least one related organ image from the target image, and generates target text information to describe the volume of multiple organs. Then, the target image, target text information, target organ image and at least one related organ image are input into the image recognition model, and the target recognition result of the target organ is output. Finally, the target recognition result is sent to the edge device.

[0080] The image recognition model described above can be trained by server 102. Specifically, when server 102 is an edge device with strong computing power, it can be trained by the edge device; when server 102 includes an edge device with relatively weak computing power and a cloud device with strong computing power, it can also be trained by the cloud device. The image recognition model can also be trained by other devices and loaded by server 102. When the image recognition model is trained by other devices, the above implementation environment also includes server 103. The training process of server 103 includes: acquiring sample images and the image recognition model to be trained. The sample images are obtained by photographing multiple organs of the human body. The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier, etc. The target organ image and at least one related organ image are segmented from the sample images, and sample text information describing the volume of multiple organs is generated. The sample image and sample text information are input into the multimodal feature extraction sub-model, which outputs a first sample feature vector. The sample target organ image and at least one sample related organ image are input into the multi-organ feature extraction sub-model, which outputs a second sample feature vector and at least one third sample feature vector. The first sample feature vector, the second sample feature vector, and at least one third sample feature vector are then fused to obtain a sample fusion feature vector. The sample fusion feature vector is then input into the classifier, which outputs the predicted recognition result of the target organ. Based on the predicted recognition result, the actual recognition result corresponding to the sample image, the second sample feature vector, and at least one third sample feature vector, the image recognition model is trained.

[0081] This application provides an image recognition method, which can be executed by a server 102, which is an edge device with strong computing power. (See also...) Figure 2 The method flow provided in this application embodiment includes:

[0082] 201. Obtain the target image.

[0083] Studies of target organs have revealed that the severity of lesions in a target organ is reflected not only in its morphology but also in the volume of at least one adjacent related organ (such as the liver or spleen). Analyzing the volume of the target organ and related organs can help identify the severity of the lesion. For example, the higher the severity of esophageal varices, the smaller the liver and the larger the spleen. Changes in the volume of the liver and spleen can help identify the severity of esophageal varices. To improve the accuracy of identifying the severity of target organ lesions, this application uses NCCT to image multiple organs of the target human body, obtaining target images. These target images are three-dimensional images with length, width, and depth. The multiple organs include the target organ and at least one adjacent related organ. When the target organ is the esophagus, the at least one related organ can be the liver, the spleen, or both the liver and spleen, etc.

[0084] 202. Segment the target organ image and at least one related organ image from the target image, and generate target text information to describe the volume of multiple organs.

[0085] The target image obtained by using NCCT to capture images of multiple organs of a target human body includes not only the image region corresponding to the target organ and the image region corresponding to at least one related organ, but also the image regions corresponding to other organs or human tissues. To better analyze the morphological changes of the target organ and at least one related organ, embodiments of this application can segment the target organ image and at least one related organ image from the target image and generate target text information describing the volume of multiple organs. Specifically, this can include the following steps:

[0086] 2021. Locate the target organ and at least one related organ in the target image to obtain the positions of the target organ and at least one related organ on the target image.

[0087] In this embodiment of the application, an organ localization model can be trained based on multiple sample images, and then the target image can be input into the organ localization model. The organ localization model can locate the target organ and at least one related organ in the target image, identify the position of the target organ and at least one related organ on the target image, and then output the position of the target organ and at least one related organ on the target image respectively.

[0088] 2022. Based on the positions of the target organ and at least one related organ on the target image, the target image is segmented to obtain the target organ image and at least one related organ image.

[0089] In this embodiment of the application, an image segmentation model can be pre-trained. After determining the positions of the target organ and at least one related organ on the target image, the position information of the target organ and at least one related organ on the target image and the target image are input into the image segmentation model, and the target organ image and at least one related organ image are output.

[0090] 2023. Based on the positions of the target organ and at least one related organ on the target image, volume analysis is performed on the target organ and at least one related organ to obtain target text information.

[0091] Since the target image is a three-dimensional image with length, width, and height, after determining the positions of the target organ and at least one related organ on the target image, the volume of the target organ on the target image can be calculated based on its position, and the volume of each related organ can be calculated based on its position. In this embodiment, the standard volume of the target organ and the standard volume of each related organ can be pre-stored. The standard volume of the target organ refers to its volume when it is not diseased. The standard volume of each related organ refers to its volume when the target organ is not diseased. Based on the standard volume of the target organ, after calculating the volume of the target organ on the target image, the volume change of the target organ can be analyzed. Based on the standard volume of each related organ, after calculating the volume of each related organ on the target image, the volume change of each related organ can be analyzed. Based on the volume of at least one related organ on the target image, the volume ratio of any two related organs can be analyzed. Based on the volume analysis results, target text information describing the volumes of multiple organs can be generated. For example, the generated target text information could be: CT scan shows that the esophagus, liver, and spleen are all of average size, and the liver-spleen volume ratio is 3.83 (average).

[0092] 203. Input the target image, target text information, target organ image and at least one related organ image into the image recognition model, and output the target organ recognition result.

[0093] The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier. The multimodal feature extraction sub-model includes an image encoder, a text encoder, a first dimensionality reduction module, and a second dimensionality reduction module. The image encoder can be a CLIP (Contrastive Language-Image Pre-Training) image encoder, which encodes the target image to obtain a global image feature vector. The text encoder can be a CLIP text encoder, which encodes the target text information to obtain a text feature vector. The first dimensionality reduction module corresponds to the image encoder and is used to reduce the dimensionality of the global image feature vector encoded by the image encoder. The second dimensionality reduction module corresponds to the text encoder and is used to reduce the dimensionality of the text feature vector encoded by the text encoder. The structures of the first and second dimensionality reduction modules can be the same or different. Figure 3 The structures of the first and second dimensionality reduction modules are shown; see [link / reference]. Figure 3 The first and second dimensionality reduction modules have the same structure, both including MLP (Multilayer Perceptron), ReLU (Rectified Linear Unit), Layer Norm, and MLP. MLP is a feedforward neural network, including input, hidden, and output layers, used to approximate complex functions through nonlinear transformations. ReLU is an activation function in neural networks, used to introduce nonlinear transformations into the network, enabling it to learn complex features and relationships. Layer Norm addresses the limitations of batch normalization in mini-batch training or recurrent neural networks (RNNs).

[0094] The multi-organ feature extraction sub-model includes a target organ feature extraction unit and at least one related organ feature extraction unit. The target organ feature extraction unit extracts the target organ feature vector from the target organ image. At least one related organ feature extraction unit extracts the related organ feature vector from at least one related organ image. Each related organ feature extraction unit can be a Uniformer branch. UniFormer is a spatiotemporal representation learning framework that seamlessly integrates Transformer, convolution, and self-attention mechanisms, aiming to solve the problems of local redundancy and global dependency in visual data, and performs well in tasks such as image classification and video classification. At least one related organ feature extraction unit can correspond to at least one related organ. When the at least one related organ is the liver, the at least one related organ feature extraction unit is a liver feature extraction unit, used to extract the liver feature vector from the liver image; when the at least one related organ is the spleen, the at least one related organ feature extraction unit is a spleen feature extraction unit, used to extract the spleen feature vector from the spleen image; when the at least one related organ is both the liver and spleen, the at least one related organ feature extraction unit is a liver feature extraction unit and a spleen feature extraction unit, used to extract the liver feature vector from the liver image and the spleen feature vector from the spleen image. Each target organ feature extraction unit and at least one related organ feature extraction unit includes N levels of feature extraction modules, where N is a positive integer and its value is greater than or equal to 1. An Organ Representation Interaction (ORI) module is established between each level of feature extraction modules in each related organ feature extraction unit and the corresponding level of feature extraction modules in the target organ feature extraction unit. The ORI module focuses on refining and integrating the imaging features of multiple related organs, enhancing the diagnosis of the target organ by integrating morphological and textural changes of multiple organs. Especially when related organs undergo pathological changes due to the target organ, it improves the sensitivity of target organ diagnosis by directing attention to key areas. In addressing the diagnosis of esophageal varices in portal hypertension, the ORI module enhances diagnostic accuracy by comprehensively analyzing subtle morphological and textural changes in the liver and spleen using NCCT images.

[0095] The ORI module takes as input a feature vector from the target organ and a feature vector from at least one related organ, each feature vector represented as a multidimensional tensor. For example, the esophageal feature vector is represented as... The liver feature vector is represented as The spleen feature vector is represented as Generally, the feature vectors in a neural network are in the form of feature images, H e W e D eH represents the length, width, and depth of the esophageal feature image, respectively. l W l D l H represents the length, width, and depth of the liver feature image, respectively. s W s D s These represent the length, width, and depth of the spleen feature image, respectively. C represents the number of channels in the esophageal feature image, liver feature image, and spleen feature image.

[0096] The ORI module processes the feature vectors from the target organ and at least one related organ in parallel using two paths: an attention path and a direct path. For the target organ interaction module corresponding to the M-th level (M ≥ 1 ≤ N) feature extraction module of any related organ feature extraction unit, when using the attention path, the target organ interaction module performs attention calculations on the M-th level related organ feature vector and the M-th level target organ feature vector to obtain the M-th level enhanced feature vector corresponding to the related organ feature extraction unit. This M-th level related organ feature vector is the related organ feature vector output by the M-th level feature extraction module of the related organ feature extraction unit, and the M-th level target organ feature vector is the target organ feature vector output by the M-th level feature extraction module of the target organ feature extraction unit. When using the direct path, the target organ interaction module performs linear projection on the M-th level target organ feature vector to obtain the projected M-th level target organ feature vector.

[0097] Figure 4 The diagram illustrates the processing procedure of the ORI module corresponding to the M-th level feature extraction module for the esophagus, liver (or spleen) on the M-th level feature vectors. See [link to documentation]. Figure 4 The ORI module first performs pooling on the feature vectors from the esophagus and liver (or spleen) to unify the M-th level feature vectors from the esophagus, liver (or spleen) to the same dimension. The M-th level esophageal feature vector after pooling operation The M-th level liver feature vector after pooling operation The M-th level spleen feature vector after pooling operation The ORI module performs pooling operations on the Mth-level esophageal feature vector. The M-th level liver feature vector after pooling operation (or the Mth-level spleen feature vector after pooling operation) The convolution operations are performed separately to obtain the M-th level esophageal feature vector and the M-th level liver feature vector (or the M-th level spleen feature vector) after convolution. The ORI module processes the M-th level esophageal feature vector and the M-th level liver feature vector (or the M-th level spleen feature vector) after convolution in parallel using two paths. The two paths are the attention path and the direct path, respectively.

[0098] In the attention path, the ORI module performs a linear mapping between the M-th level esophageal feature vector and the M-th level liver feature vector (or the M-th level spleen feature vector) after convolution, obtaining M-th level esophageal feature vectors and M-th level liver feature vectors (or M-th level spleen feature vectors) suitable for the attention mechanism. Then, using the M-th level esophageal feature vector suitable for the attention mechanism as the key (K) and value (V), and the M-th level liver feature vector (or M-th level spleen feature vector) suitable for the attention mechanism as the query (Q), the attention mechanism calculation formula is applied. The calculation is performed to implement scaled dot product attention, resulting in the M-th level attention feature vector T. A projection matrix W is used. P (where W) P ∈R C×C Projecting the M-th level attention feature vector T onto the M-th level attention feature vector yields the projected M-th level attention feature vector F. Att F Att =TW P , The M-th level attention feature vector F after projection Att A convolution operation is performed to obtain the M-th enhanced feature vector. An interpolation operation is then performed on the M-th enhanced feature vector and the M-th esophageal feature vector to fuse them. The fused feature vector is then used as the input to the (M+1)-th level feature extraction module of the esophageal feature extraction unit.

[0099] In the direct path, the ORI module performs a linear projection on the M-th level esophageal feature vector after convolution, followed by another convolution operation to obtain the projected M-th level esophageal feature vector. Interpolation is then performed on the projected M-th level esophageal feature vector and the M-th level liver feature vector (or the M-th level spleen feature vector) to fuse the projected M-th level esophageal feature vector and the M-th level liver feature vector (or the M-th level spleen feature vector). The fused feature vector serves as the input to the (M+1)-th level feature extraction module of the liver feature extraction unit (or the (M+1)-th level feature extraction module of the spleen feature extraction unit). Linear projection can be implemented using the following mathematical expression:

[0100] Y = XW + b

[0101] Where X represents the input feature matrix, assuming the dimension of X is m×n, indicating that there are m samples, each with n features; W represents the linear transformation matrix, and if we want to map the input data from n dimensions to a lower or higher dimension p, then the size of W should be n×p; b represents the bias vector, used to adjust the center of the projected data; Y represents the projected output feature matrix, with a dimension of m×p.

[0102] Applying linear projection to the feature vector in the direct path, without attention computation, is simple and efficient, and can further supplement the information captured by the attention mechanism.

[0103] The above Figure 4 This demonstrates the processing procedure of an ORI module within a multi-organ feature extraction sub-model. The processing principles for other ORI modules are similar. Figure 4 The processing principle of the ORI shown is the same; see details below. Figure 4 The processing flow is not detailed here. The two paths of the ORI module in the multi-organ feature extraction sub-model operate alternately. After alternating iterations, convolution and interpolation operations are performed to provide more detailed feature maps for subsequent use. Through multiple iterations, the feature integration effect can be continuously improved, thereby maximizing signal processing efficiency and helping to enhance feature fusion and recognition of multi-organ images.

[0104] The ORI module designed in this application has the following advantages:

[0105] Reduced computational complexity: The ORI module significantly reduces computational complexity by optimizing path switching and feature processing.

[0106] Enhanced feature representation: The ORI module integrates attention path and direct processing path to make feature representation more refined and discriminative.

[0107] Overall, the design of the ORI module not only improves computational efficiency but also enhances diagnostic performance in the comprehensive analysis of CT images, making it particularly suitable for comprehensive clinical evaluation in the context of portal hypertension.

[0108] The classifier is used to determine the severity level of lesions in the target organ. In this embodiment, the classifier uses ordered regression to define the severity level of lesions in the target organ. In ordered regression, the model not only needs to predict the category to which each sample belongs, but also needs to reflect the natural order between categories. Traditional classifiers ignore this order, while ordered regression captures this information through ordered encoding. In this embodiment, the esophageal variceal severity levels defined by ordered regression include four grades, G0 to G3, representing lesions from none to severe.

[0109] The following describes the impact of esophageal variceal levels defined by ordered regression on the recognition results of an image recognition model. Taking the image recognition model training process as an example, the loss function calculates the loss between predicted values ​​based on the order information defined by G0-G3. Binary encoding is used to convert the category targets (e.g., G0, G1, G2, G3) into a cumulative coding. The encoded binary vector reflects the order information. The specific steps for calculating the loss function based on the encoded binary vector are as follows:

[0110] The first step is Target Transformation: Each category, such as G0, G1, G2, and G3, is converted into a binary vector. Specifically, G0 is converted to [1, 0, 0, 0], G1 to [1, 1, 0, 0], G2 to [1, 1, 1, 0], and G3 to [1, 1, 1, 1]. This encoding method ensures that higher-level categories always contain lower-level categories.

[0111] The second step is to construct the modified target: the modified target is a tensor with the same shape as the prediction, initialized to all zeros. Assuming the target is a single value (i.e., the batch contains only one sample), taking G2 as an example, we set all positions [0, 0: 2+1] of modified_target to 1, resulting in [1, 1, 1, 0]. For batch processing, this operation is performed separately for each target within a batch of data.

[0112] The third step is to calculate the loss: using the mean squared error loss (nn.MSELoss) to compare the predicted values ​​and the modified target. The loss function calculates the mean squared error between the model's predicted values ​​and the transformed target values ​​(represented as cumulative probabilities), penalizing the deviation between the predictions and the target, and specifically taking into account the order attribute.

[0113] There are three main reasons for using this method:

[0114] The first aspect is order regularization: This encoding and decoding method reflects the order relationship between categories. For example, G3 is more severe than G2, and it also includes the characteristics of G2, G1 and G0.

[0115] The second aspect is that the error is more sensitive: if the order of the model prediction is misplaced, for example, G2[1,1,1,0] is predicted as G0[1,0,0,0], then the error will be reflected in all the order markers. This bias doubles the penalty for the order error.

[0116] The third aspect is effective progressive learning: During training, the model learns not only the linear decision boundary of point-to-point classification, but also understands the probability distribution of each sequential class, thus better handling subtle differences between sequences. This method is suitable for solving problems with clear sequential classification, such as medical grade assessment and quality scoring, enabling the model to not only make correct classifications, but also capture the sequential relationship between successive levels.

[0117] Figure 5 The classifier's recognition process is shown in the diagram. See [link / reference] Figure 5 When the fused feature vector is input into the classifier, it is processed by the MLP layer and the sigmoid function to obtain the target organ recognition result.

[0118] Based on an image recognition model, this embodiment of the application inputs a target image, target text information, a target organ image, and at least one related organ image into the image recognition model, and outputs the target organ recognition result, including the following steps:

[0119] 2031. Input the target image and target text information into the multimodal feature extraction sub-model and output the first feature vector.

[0120] The first feature vector is a multimodal feature vector that concatenates global image feature vectors and text feature vectors from multiple organs. The target image and target text information are input into the multimodal feature extraction sub-model, which outputs the first feature vector. Specifically, this includes: inputting the target image into an image encoder to output a global image feature vector; inputting the target text information into a text encoder to output a text feature vector; inputting the global image feature vector into a first dimensionality reduction module to output a dimensionality-reduced global image feature vector; inputting the text feature vector into a second dimensionality reduction module to output a dimensionality-reduced text feature vector; and concatenating the dimensionality-reduced global image feature vector and the dimensionality-reduced text feature vector to obtain the first feature vector.

[0121] 2032. Input the target organ image and at least one related organ image into the multi-organ feature extraction sub-model, and output a second feature vector and at least one third feature vector.

[0122] The second feature vector is the target organ feature vector enhanced by the feature vectors of at least one related organ. The third feature vector is the related organ feature vector fused with the target organ feature vector. The target organ image and at least one related organ image are input into the multi-organ feature extraction sub-model to output the second feature vector and at least one third feature vector. Specifically, this involves inputting the target organ image into the N-level feature extraction module of the target organ feature extraction unit, and inputting at least one related organ image into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the second feature vector and at least one third feature vector are output.

[0123] The process of the multi-organ feature extraction sub-model outputting the second feature vector includes: obtaining the M-th level target organ feature vector and the M-th level related organ feature vector; inputting the M-th level target organ feature vector and the M-th level related organ feature vector into the target organ interaction module; the target organ interaction module processes the M-th level target organ feature vector and the M-th level related organ feature vector using an attention path to output the M-th level enhanced feature vector; fusing the M-th level target organ feature vector, the M-th level enhanced feature vector, and the M-th level enhanced feature vectors corresponding to other related organ feature extraction units to obtain the M-th level target organ fused feature vector; then inputting the M-th level target organ fused feature vector into the (M+1)-th level feature extraction module of the target organ feature extraction unit to output the (M+1)-th level target organ feature vector; processing the (M+1)-th level target organ feature vector in the same way as the M-th level target organ feature vector until the N-th level target organ feature vector output by the N-th level feature extraction module of the target organ feature extraction unit is obtained, and using the N-th level target organ feature vector as the second feature vector.

[0124] The process of the multi-organ feature extraction sub-model outputting the third feature vector includes: obtaining the M-th level target organ feature vector and the M-th level related organ feature vector; inputting the M-th level target organ feature vector and the M-th level related organ feature vector into the target organ interaction module; the target organ interaction module linearly projects the M-th level target organ feature vector using a direct path, and outputs the projected M-th level target organ feature vector; fusing the projected M-th level target organ feature vector with the M-th level related organ feature vector to obtain the M-th level related organ fused feature vector; inputting the M-th level related organ fused feature vector into the M+1 level feature extraction module of the related organ feature extraction unit, and outputting the M+1 level related organ feature vector; processing the M+1 level related organ feature vector in the same way as the M-th level related organ feature vector until the N-th level related organ feature vector output by the N-th level feature extraction module of the related organ feature extraction unit is obtained, and using the N-th level related organ feature vector as the third feature vector corresponding to the related organ feature extraction unit.

[0125] For example, if the target organ is the esophagus, and at least one related organ is the liver and spleen, the multi-organ feature extraction sub-model includes an esophageal feature extraction unit, a liver feature extraction unit, and a spleen feature extraction unit. After acquiring the esophageal, liver, and spleen images, these images are input into the esophageal, liver, and spleen feature extraction units, respectively. For the M-th level esophageal feature vector output by the M-th level feature extraction module of the esophageal feature extraction unit, the M-th level liver feature vector output by the M-th level feature extraction module of the liver feature extraction unit, and the M-th level spleen feature vector output by the M-th level feature extraction module of the spleen feature extraction unit, the M-th level esophageal feature vector and the M-th level liver feature vector are input into the first organ interaction module. This first organ interaction module is the organ interaction module corresponding to the M-th level feature extraction modules of the esophageal and liver feature extraction units. The first organ interaction module processes the M-th level esophageal and liver feature vectors using an attention path, outputting an enhanced M-th level feature vector. It also processes the M-th level esophageal feature vector using a direct path, outputting a projected M-th level esophageal feature vector. Simultaneously, the M-th level esophageal and spleen feature vectors are input into the second organ interaction module, which corresponds to the M-th level feature extraction modules of the esophageal and spleen feature extraction units. The second organ interaction module processes the M-th level esophageal and spleen feature vectors using an attention path, outputting an enhanced M-th level feature vector. It also processes the M-th level esophageal feature vector using a direct path, outputting a projected M-th level esophageal feature vector. Finally, the enhanced M-th level esophageal feature vectors output from the first and second organ interaction modules are fused to obtain the fused M-th level esophageal feature vector. The M-th enhanced feature vector and the M-th liver feature vector output by the first organ interaction module are fused to obtain the M-th liver fused feature vector. The M-th enhanced feature vector and the M-th spleen feature vector output by the second organ interaction module are fused to obtain the M-th spleen fused feature vector.The M-th level esophageal fusion feature vector is input into the (M+1)-th level feature extraction module of the esophageal feature extraction unit, the M-th level liver fusion feature vector is input into the (M+1)-th level feature extraction module of the liver feature extraction unit, and the M-th level spleen fusion feature vector is input into the (M+1)-th level feature extraction module of the spleen feature extraction unit. After processing by the (M+1)-th level feature extraction modules of the esophageal feature extraction unit, the liver feature extraction unit, and the spleen feature extraction unit, as well as subsequent feature extraction modules and multiple organ interaction modules, the esophageal feature extraction unit finally outputs the second feature vector (i.e., esophageal feature vector), the liver feature extraction unit outputs the third feature vector (i.e., liver feature vector), and the spleen feature extraction unit outputs the third feature vector (i.e., spleen feature vector).

[0126] 2033. The first feature vector, the second feature vector, and at least one third feature vector are fused to obtain the fused feature vector.

[0127] After obtaining the first feature vector, the second feature vector, and at least one third feature vector, the first feature vector, the second feature vector, and at least one third feature vector can be fused to obtain a fused feature vector. This fused feature vector is used to describe the features of multiple organs from both global and local perspectives.

[0128] 2034. Input the fused feature vector into the classifier and output the target recognition result.

[0129] When the fused feature vector is input into the classifier, the classifier outputs the target organ identification result, which is a binary vector. This binary vector needs to be converted into the corresponding level. For example, the target identification result for esophageal varices is [1, 1, 0, 0], which indicates that the esophageal varices level is G1.

[0130] This application embodiment extracts the overall image features of the entire CT image using an image encoder and extracts the textual features of the relevant organ volumes using a text encoder. Simultaneously, three Uniformer branches are used to extract local image features of the organs. During this process, an attention mechanism is used in the intermediate layers for interactive learning to capture the intrinsic connections between organs. Then, the local image features of the organs are fused with the overall image and textual features, and an ordered regression loss function is used for classification. This method effectively improves the accuracy of target organ identification. The advantages of this application embodiment are specifically reflected in:

[0131] Multi-organ fusion and attention mechanism: Using the attention mechanism for multi-organ fusion and interactive learning can effectively and comprehensively analyze the features of multiple related organs in plain CT images.

[0132] CLIP model integration with clinical imaging features: Combining volumetric features from clinical images with image and text information, the CLIP model enhances the understanding of organ volume, morphology, and structural features.

[0133] Comprehensive information analysis: By comprehensively assessing the status of multiple organs, it provides important value for the diagnosis of target organs, significantly improving the accuracy and reliability of diagnosis.

[0134] Personalized treatment support: Enhanced image analysis depth provides doctors with a more comprehensive clinical perspective, helping to develop more precise and personalized treatment plans.

[0135] Furthermore, the image recognition method provided in this application is a non-invasive method. Through in-depth analysis of multi-organ imaging data, this method provides patients with a more comfortable and efficient target organ screening experience, significantly reducing the stress on patients' bodies, improving patient acceptance and compliance. At the same time, automation and efficiency help doctors quickly assess risks, formulate treatment plans, and reduce doctors' workload, improving the standardization and consistency of the overall diagnostic process, providing more accurate diagnostic tools for clinical practice, achieving higher-quality medical services in the entire clinical management, and further promoting the development of medical imaging.

[0136] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0137] This application provides an image recognition method, which can be executed by a cloud-side device with strong computing power in the server 102. See [link to relevant documentation]. Figure 6 The method flow provided in this application embodiment includes:

[0138] 601. An image processing task sent by the receiving end device, the image processing task including the target image.

[0139] The target image is obtained by capturing images of multiple organs of the target human body, including the target organ and at least one adjacent related organ. Due to the limited computing power of the edge device, it cannot recognize the target image. After acquiring the target image, an image processing task is generated and sent to the cloud-based device. This image processing task includes the target image and instructs the cloud-based device to process the target image to identify the lesion level of the target organ.

[0140] 602. Segment the target organ image and at least one related organ image from the target image, and generate target text information to describe the volume of multiple organs.

[0141] 603. Input the target image, target text information, target organ image and at least one related organ image into the image recognition model, and output the target organ recognition result.

[0142] 604. Send the target recognition result to the end-side device.

[0143] After the cloud-side device identifies the target organ by calling the image recognition model, it sends the target identification result to the end-side device so that the user on the end-side device can know the lesion level of the target organ.

[0144] This application provides a method for training an image recognition model, which can be executed by server 102 or server 103. See also Figure 7 The method flow provided in this application embodiment includes:

[0145] 701. Obtain sample images and the image recognition model to be trained.

[0146] The sample images are images of multiple organs of the human body captured by NCCT. These organs include the target organ and at least one adjacent related organ, which can be the liver, spleen, or both. The image recognition model to be trained includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier.

[0147] 702. Segment the target organ image and at least one related organ image from the sample image, and generate sample text information to describe the volume of multiple organs.

[0148] The sample images obtained by using NCCT to photograph multiple organs of the human body include not only the image region corresponding to the target organ and the image region corresponding to at least one related organ, but also the image regions corresponding to other organs or human tissues. In order to better analyze the morphological changes of the target organ and at least one related organ, this embodiment of the application can segment the sample target organ image and the sample related organ image from the sample image and generate sample text information describing the volume of multiple organs. Specifically, it can include the following steps:

[0149] 7021. Locate the target organ and at least one related organ in the sample image to obtain the positions of the target organ and at least one related organ on the sample image.

[0150] In this embodiment of the application, a sample image can be input into an organ localization model. The organ localization model can then be used to locate the target organ and at least one related organ in the sample image, identify the positions of the target organ and at least one related organ on the sample chest image, and output the positions of the target organ and at least one related organ on the sample chest image.

[0151] 7022. Based on the positions of the target organ and at least one related organ on the sample chest image, the sample chest image is segmented to obtain the sample target organ image and at least one sample related organ image.

[0152] In this embodiment of the application, an image segmentation model can be pre-trained. After determining the positions of the target organ and at least one related organ on the sample chest image, the position information of the target organ and at least one related organ on the sample image and the sample image are input into the image segmentation model, and the sample target organ image and at least one sample related organ image are output.

[0153] 7023. Based on the positions of the target organ and at least one related organ on the sample chest image, perform volume analysis on the target organ and at least one related organ to obtain sample text information.

[0154] Since the sample image is a three-dimensional image with length, width, and height, after determining the positions of the target organ and at least one related organ on the sample image, the volume of the target organ can be calculated based on its position, and the volume of each related organ can be calculated based on its position. Based on the standard volume of the target organ, after calculating its volume on the sample image, the volume variation of the target organ can be analyzed. Based on the standard volume of each related organ, after calculating the volume of each related organ on the sample image, the volume variation of each related organ can be analyzed. Based on the volume of at least one related organ on the sample image, the volume ratio of any two related organs can be analyzed. Based on the volume analysis results, sample text information describing the volumes of multiple organs can be generated.

[0155] 703. Input the sample image and sample text information into the multimodal feature extraction sub-model and output the first sample feature vector.

[0156] The multimodal feature extraction sub-model of this application embodiment includes an image encoder, a text encoder, a first dimensionality reduction module, and a second dimensionality reduction module. The sample image and sample text information are input into the multimodal feature extraction sub-model, and a first sample feature vector is output. Specifically, this includes: inputting the sample image into the image encoder to output a global image feature vector; inputting the sample text information into the text encoder to output a sample text feature vector; inputting the global image feature vector into the first dimensionality reduction module to output a dimensionality-reduced global image feature vector; inputting the sample text feature vector into the second dimensionality reduction module to output a dimensionality-reduced sample text feature vector; and concatenating the dimensionality-reduced global image feature vector and the dimensionality-reduced sample text feature vector to obtain the first sample feature vector.

[0157] 704. Input the sample image and at least one sample-related organ image into the multi-organ feature extraction sub-model, and output the second sample feature vector and at least one third sample feature vector.

[0158] The target organ image of the sample is input into the N-level feature extraction module of the target organ feature extraction unit, and at least one related organ image of the sample is input into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the second feature vector of the sample and at least one third feature vector of the sample are output.

[0159] 705. The first sample feature vector, the second sample feature vector, and at least one third sample feature vector are fused to obtain the sample fusion feature vector.

[0160] 706. Input the sample fusion feature vector into the classifier and output the predicted identification result of the target organ.

[0161] 707. The image recognition model is trained based on the predicted recognition results, the actual recognition results corresponding to the sample images, the feature vectors of the second sample, and at least one feature vector of the third sample.

[0162] This application embodiment pre-designs a total loss function, which is composed of a weighted sum of a first loss function and a second loss function. The first loss function is used to calculate the deviation between the predicted result and the true result of the sample chest image. Ordinal The expression is as follows:

[0163]

[0164] Among them, H F Y represents the prediction result. ord This indicates the actual result.

[0165] The second loss function is used to calculate the correlation loss between the feature vectors of the second and third samples. This second loss function can be the DCCA (Deep Canonical Correlation Analysis) loss function. DCCA is an algorithm used to learn the correlation feature representation between two views. The goal of DCCA is to project two related input variables into a shared representation space through a neural network, making the projected representations as correlated as possible within that space. The core steps of the DCCA loss function are explained below:

[0166] 1. Non-linear projection:

[0167] For input feature vectors X1 and X2, two nonlinear neural networks f1 and f2 are used for projection to obtain two new representations H1 = f1(X1) and H2 = f2(X2). In this way, DCCA can capture complex nonlinear relationships in the input data.

[0168] 2. Standardization:

[0169] Perform standardization operations on H1 and H2:

[0170]

[0171] The purpose of standardization is to set the mean of the data to 0 and the standard deviation to 1, preventing numerical instability and enhancing the comparability between features. ∈ is a small constant used to prevent division by zero errors.

[0172] 3. Calculate the Cross-Correlation Matrix:

[0173] Calculate the cross-correlation matrix C between H1 and H2:

[0174]

[0175] The elements of the cross-correlation matrix C represent the linear correlation between features from the two projection spaces.

[0176] 4. DCCA Loss Calculation:

[0177] Calculate DCCA loss

[0178] Among them, T r(C) is the trace of matrix C, representing the sum of its diagonal elements, which reflect the degree of aggregation correlation for each pair of related features. ∥H1∥ F and ∥H2∥ F is the Frobenius norm of H1 and H2, representing the "size" of the matrix. The Frobenius norm is used to standardize correlations.

[0179] The above formula minimizes the loss by maximizing the correlation between features, which is consistent with the goal of canonical correlation analysis.

[0180] The advantages of DCCA are as follows:

[0181] High-dimensional nonlinear correlations: By using neural networks, DCCA can capture complex nonlinear correlations that are impossible to achieve in classical linear CCA.

[0182] Versatility: It can be applied to various data types and structures, and is suitable for association learning of multimodal data such as images, text, and audio.

[0183] By optimizing the DCCA loss, the model learns a projection function that makes the projected feature representations between two views as relevant as possible in a shared low-dimensional space.

[0184] The total loss function L constructed in the embodiments of this application Overall as follows:

[0185]

[0186] Where λ represents the weighting parameter, H E H represents the second eigenvector. O H represents the third eigenvector, where O is L. O For the liver feature vector, when O is S, H O This represents the feature vector of the spleen.

[0187] Specifically, the image recognition model is trained based on the predicted recognition results, the actual recognition results corresponding to the sample images, the feature vectors of the second samples, and at least one feature vector of the third samples, including:

[0188] 7071. Input the predicted recognition result and the actual recognition result into the first loss function, and output the value of the first loss function.

[0189] 7072. Input the second sample feature vector and at least one third sample feature vector into the second loss function, and output the value of the second loss function.

[0190] 7073. The first loss function value and the second loss function value are weighted and summed to obtain the total loss function value.

[0191] 7074. Based on the total loss function value, adjust the model parameters of the image recognition model to obtain a trained image recognition model.

[0192] The image recognition model trained in this application embodiment is used to identify the lesion level of a target organ. Depending on the target organ, the sample images used for training the image recognition model differ, resulting in different functions for the trained model, but the training methods also differ. For example, when the target organ is the esophagus, the sample image may include multiple organs such as the esophagus and at least one related organ adjacent to the esophagus, such as the liver and spleen. The trained image recognition model is then used to identify the level of esophageal varices.

[0193] Figure 8 The training process of the image recognition model provided in this application embodiment is illustrated. This image recognition model is used to identify the grade of esophageal varices. See also Figure 8 An NCCCT was used to photograph multiple organs of the human body, obtaining sample images including the esophagus, liver, and spleen. The esophagus, liver, and spleen were located in the sample images. Based on the location results, sample esophageal images, sample liver images, and sample spleen images were segmented from the sample images, and sample text information describing the volume of the esophagus, liver, and spleen was generated. The sample images and sample text information were input into the image encoder and text encoder of the multimodal feature extraction sub-model, respectively. The image encoder encoded the sample images to obtain a global image feature vector. The first dimensionality reduction feature extraction module was used to reduce the dimensionality of the global image feature vector to obtain the dimensionality-reduced global image feature vector H. img Simultaneously, the text encoder encodes the sample text information to obtain the sample text feature vector. The second dimensionality reduction feature extraction module then performs dimensionality reduction processing on the sample text feature vector to obtain the dimensionality-reduced sample text feature vector H. text The dimensionality-reduced global image feature vector H is then used to... img and the dimensionality-reduced sample text feature vector H text The images are concatenated to obtain the first sample feature vector. The sample esophageal image, sample liver image, and sample spleen image are then input into the multi-organ feature extraction sub-model, which outputs the sample esophageal feature vector H. E Sample liver feature vector H L and the spleen feature vector H of the sample S Then, the first sample feature vector and the sample esophageal feature vector H are... E Sample liver feature vector H L and the spleen feature vector H of the sample SThe samples are fused to obtain a fused feature vector. This fused feature vector is then input into a classifier, which outputs a prediction result for the level of esophageal varices. Based on this prediction result, the actual result corresponding to the sample chest image, and the sample esophageal feature vector H... E Sample liver feature vector H L and the spleen feature vector H of the sample S A model for identifying the level of esophageal varices was trained. Specifically, the value of the first loss function was calculated based on the predicted and actual results. This was based on the sample esophageal feature vector H. E Sample liver feature vector H L and the spleen feature vector H of the sample S Calculate the value of the second loss function. Then, weightedly sum the values ​​of the first and second loss functions to obtain the total loss function value. Based on the total loss function value, adjust the model parameters of the image recognition model to obtain the trained image recognition model.

[0194] This application provides a computer-aided diagnostic method for esophageal varices, using an end-side device with strong computing power as an example. See also... Figure 9 The method flow provided in this application embodiment includes:

[0195] 901. Obtain the target chest image.

[0196] The target chest image is obtained by photographing multiple organs of the target chest, including the esophagus and at least one related organ adjacent to the esophagus.

[0197] 902. Segment the esophagus image and at least one related organ image from the target chest image, and generate target text information to describe the volume of multiple organs.

[0198] 903. Input the target chest image, target text information, esophageal image, and at least one related organ image into the image recognition model, and output the recognition result of the esophageal varices level.

[0199] When the computer-aided diagnostic method for esophageal varices is executed by a cloud-based device with strong computing power, the cloud-based device first receives an image processing task sent by the end-side device. This image processing task includes a target chest image. After obtaining the identification result of the esophageal varices level through the above steps 902 and 903, the identification result is sent to the end-side device.

[0200] Please refer to Figure 10 The diagram illustrates a structural schematic of an image recognition device provided in an embodiment of this application. This device can be implemented through software, hardware, or a combination of both, and can be all or part of an electronic device. The device includes:

[0201] The acquisition module 1001 is used to acquire a target image, which is obtained by photographing multiple organs of the target human body, including the target organ and at least one related organ adjacent to the target organ.

[0202] The segmentation module 1002 is used to segment the target organ image and at least one related organ image from the target image;

[0203] The generation module 1003 is used to generate target text information describing the volume of the plurality of organs;

[0204] The input / output module 1004 is used to input the target image, the target text information, the target organ image and at least one related organ image into the image recognition model, and output the target recognition result of the target organ.

[0205] In another embodiment of this application, the segmentation module 1002 is used to locate the target organ and the at least one related organ in the target image to obtain the positions of the target organ and the at least one related organ on the target image respectively; based on the positions of the target organ and the at least one related organ on the target image respectively, the target image is segmented to obtain the target organ image and the at least one related organ image;

[0206] The generation module 1003 is used to perform volume analysis on the target organ and the at least one related organ based on their respective positions on the target image to obtain the target text information.

[0207] In another embodiment of this application, the image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier. An input / output module 1004 is used to input the target image and the target text information into the multimodal feature extraction sub-model and output a first feature vector, which is a multimodal feature vector concatenated with global image feature vectors and text feature vectors of the multiple organs; input the target organ image and at least one related organ image into the multi-organ feature extraction sub-model and output a second feature vector and at least one third feature vector, where the second feature vector is a target organ feature vector enhanced by the related organ feature vectors of at least one related organ, and the third feature vector is a related organ feature vector fused with the target organ feature vector; fuse the first feature vector, the second feature vector, and the at least one third feature vector to obtain a fused feature vector, which is used to describe the features of the multiple organs from both global and local perspectives; and input the fused feature vector into the classifier to output the target recognition result.

[0208] In another embodiment of this application, the multimodal feature extraction sub-model includes an image encoder, a text encoder, a first dimensionality reduction module, and a second dimensionality reduction module. An input-output module 1004 is used to input the target image into the image encoder and output a global image feature vector; input the target text information into the text encoder and output a text feature vector; input the global image feature vector into the first dimensionality reduction module and output a dimensionality-reduced global image feature vector; input the text feature vector into the second dimensionality reduction module and output a dimensionality-reduced text feature vector; and concatenate the dimensionality-reduced global image feature vector and the dimensionality-reduced text feature vector to obtain the first feature vector.

[0209] In another embodiment of this application, the multi-organ feature extraction sub-model includes a target organ feature extraction unit and at least one related organ feature extraction unit. The target organ feature extraction unit and at least one related organ feature extraction unit each include N-level feature extraction modules. An organ interaction module is provided between each level feature extraction module of each related organ feature extraction unit and the same level feature extraction module of the target organ feature extraction unit. N is a positive integer.

[0210] For any M-th level feature extraction module of a related organ feature extraction unit, the target organ interaction module is greater than or equal to 1 and less than or equal to N. The target organ interaction module is used to perform attention calculation on the M-th level related organ feature vector and the M-th level target organ feature vector to obtain the M-th level enhanced feature vector corresponding to the related organ feature extraction unit. The M-th level related organ feature vector is the related organ feature vector output by the M-th level feature extraction module of the related organ feature extraction unit, and the M-th level target organ feature vector is the target organ feature vector output by the M-th level feature extraction module of the target organ feature extraction unit.

[0211] The target organ interaction module is also used to perform linear projection on the feature vector of the Mth level target organ to obtain the projected feature vector of the Mth level target organ.

[0212] In another embodiment of this application, the input / output module 1004 is used to input the target organ image into the N-level feature extraction module of the target organ feature extraction unit, and input the at least one related organ image into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the second feature vector and the at least one third feature vector are output.

[0213] In another embodiment of this application, the input / output module 804 is used to acquire the M-th level target organ feature vector and the M-th level related organ feature vector; input the M-th level target organ feature vector and the M-th level related organ feature vector into the target organ interaction module, and output the M-th level enhanced feature vector; fuse the M-th level target organ feature vector, the M-th level enhanced feature vector, and the M-th level enhanced feature vector corresponding to other related organ feature extraction units to obtain the M-th level target organ fused feature vector; input the M-th level target organ fused feature vector into the M+1 level feature extraction module of the target organ feature extraction unit, and output the M+1 level target organ feature vector; process the M+1 level target organ feature vector according to the processing method of the M-th level target organ feature vector until the N-th level target organ feature vector output by the N-th level feature extraction module of the target organ feature extraction unit is obtained, and use the N-th level target organ feature vector as the second feature vector.

[0214] In another embodiment of this application, the input / output module 804 is used to acquire the M-th level target organ feature vector and the M-th level related organ feature vector; input the M-th level target organ feature vector and the M-th level related organ feature vector into the target organ interaction module, and output the projected M-th level target organ feature vector; fuse the projected M-th level target organ feature vector with the M-th level related organ feature vector to obtain the M-th level related organ fused feature vector; input the M-th level related organ fused feature vector into the M+1 level feature extraction module of the related organ feature extraction unit, and output the M+1 level related organ feature vector; process the M+1 level related organ feature vector according to the processing method of the M-th level related organ feature vector until the N-th level related organ feature vector output by the N-th level feature extraction module of the related organ feature extraction unit is obtained, and the N-th level related organ feature vector is used as the third feature vector corresponding to the related organ feature extraction unit.

[0215] Please refer to Figure 11 The diagram illustrates a structural schematic of an image recognition device provided in an embodiment of this application. This device can be implemented through software, hardware, or a combination of both, and can be all or part of an electronic device. The device includes:

[0216] The receiving module 1101 is used to receive an image processing task sent by the end-side device. The image processing task includes a target image, which is obtained by capturing multiple organs of a target human body. The multiple organs include the target organ and at least one related organ adjacent to the target organ.

[0217] Segmentation module 1102 is used to segment a target organ image and at least one related organ image from the target image;

[0218] The generation module 1103 is used to generate target text information describing the volume of the plurality of organs;

[0219] The input / output module 1104 is used to input the target image, the target text information, the target organ image and at least one related organ image into the image recognition model, and output the target recognition result of the target organ;

[0220] The sending module 1105 is used to send the target identification result to the end-side device.

[0221] Please refer to Figure 12 The diagram illustrates a structural schematic of a training device for an image recognition model provided in this application embodiment. This device can be implemented through software, hardware, or a combination of both, and can be all or part of an electronic device. The device includes:

[0222] The acquisition module 1201 is used to acquire sample images and an image recognition model to be trained. The sample images are obtained by taking pictures of multiple organs of the human body. The multiple organs include a target organ and at least one related organ adjacent to the target organ. The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier.

[0223] The segmentation module 1202 is used to segment the target organ image and at least one related organ image from the sample image;

[0224] The generation module 1203 is used to generate sample text information describing the volume of the plurality of organs;

[0225] The first input / output module 1204 is used to input the sample image and the sample text information into the multimodal feature extraction sub-model and output the first sample feature vector.

[0226] The second input / output module 1205 is used to input the sample target organ image and at least one sample related organ image into the multi-organ feature extraction sub-model, and output a second sample feature vector and at least one third sample feature vector.

[0227] The fusion module 1206 is used to fuse the first sample feature vector, the second sample feature vector, and the at least one third sample feature vector to obtain a sample fusion feature vector.

[0228] The third input / output module 1207 is used to input the sample fusion feature vector into the classifier and output the estimated recognition result of the target organ.

[0229] The training module 1208 is used to train the image recognition model based on the estimated recognition result, the actual recognition result corresponding to the sample image, the second sample feature vector, and the at least one third sample feature vector.

[0230] In another embodiment of this application, the training module 1208 is used to input the estimated recognition result and the actual recognition result into a first loss function and output a first loss function value; input the second sample feature vector and the at least one third sample feature vector into a second loss function and output a second loss function value; perform a weighted summation of the first loss function value and the second loss function value to obtain a total loss function value; and adjust the model parameters of the image recognition model based on the total loss function value to obtain the trained image recognition model.

[0231] Please refer to Figure 13This application illustrates an embodiment of a computer-aided diagnostic device for esophageal varices, which can be implemented through software, hardware, or a combination of both, and can be all or part of an electronic device. The device includes:

[0232] The acquisition module 1301 is used to acquire a target chest image, which is obtained by taking pictures of multiple organs of the target chest, including the esophagus and at least one related organ adjacent to the esophagus.

[0233] The segmentation module 1302 is used to segment an esophageal image and at least one related organ image from the target chest image;

[0234] The generation module 1303 is used to generate target text information describing the volume of the plurality of organs;

[0235] The input / output module 1304 is used to input the target chest image, the target text information, the esophageal image and at least one related organ image into the image recognition model, and output the recognition result of the esophageal varicose vein level.

[0236] Figure 14 This illustration shows a structural block diagram of an electronic device 1400 provided in an exemplary embodiment of this application. Typically, the electronic device 1400 includes a processor 1401 and a memory 1402.

[0237] Processor 1401 can be implemented in at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1401 may also include a main processor and a coprocessor; the main processor is a processor for processing data in the wake-up state, and the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, processor 1401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1401 may also include an artificial intelligence processor for handling computational operations related to machine learning.

[0238] The memory 1402 may include one or more computer-readable storage media, which may be non-transitory computer-readable storage media, such as CD-ROM (Compact Disc Read-Only Memory), ROM, RAM (Random Access Memory), magnetic tape, floppy disk, and optical data storage devices. The computer-readable storage medium stores at least one computer program, which, when executed, can implement an image recognition method, an image recognition model training method, or a computer-aided diagnostic method for esophageal varices.

[0239] Of course, the aforementioned electronic device may also include other components, such as input / output interfaces and communication components. Input / output interfaces provide an interface between the processor and peripheral interface modules, which can be output devices, input devices, etc. Communication components are configured to facilitate wired or wireless communication between the electronic device and other devices.

[0240] Those skilled in the art will understand that Figure 14 The structure shown does not constitute a limitation on the electronic device 1400, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0241] This application provides a computer-readable storage medium storing at least one computer program. When executed by a processor, the at least one computer program can implement the above-mentioned image recognition method, or the image recognition model training method, or the computer-aided diagnosis method for identifying the level of esophageal varices.

[0242] This application provides a computer program product, which includes a computer program that, when executed by a processor, can implement the above-mentioned image recognition method, or image recognition model training method, or computer-aided diagnostic method for identifying the level of esophageal varices.

[0243] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0244] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An image recognition method, characterized in that, The method includes: Acquire a target image, which is obtained by photographing multiple organs of a target human body, including the target organ and at least one related organ adjacent to the target organ; The target organ image and at least one related organ image are segmented from the target image, and target text information describing the volume of the multiple organs is generated. The target image, the target text information, the target organ image, and at least one related organ image are input into the image recognition model, and the target recognition result of the target organ is output.

2. The method according to claim 1, characterized in that, The step of segmenting the target organ image and at least one related organ image from the target image, and generating target text information describing the volume of the plurality of organs, includes: The target organ and the at least one related organ in the target image are located to obtain the positions of the target organ and the at least one related organ on the target image, respectively; Based on the positions of the target organ and the at least one related organ on the target image, the target image is segmented to obtain the target organ image and the at least one related organ image; Based on the target image and the positions of the at least one related organ on the target image, volume analysis is performed on the target organ and the at least one related organ to obtain the target text information.

3. The method according to claim 1, characterized in that, The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier. The step of inputting the target image, the target text information, the target organ image, and at least one related organ image into the image recognition model, and outputting the target organ recognition result, includes: The target image and the target text information are input into the multimodal feature extraction sub-model, and a first feature vector is output. The first feature vector is a multimodal feature vector that is a combination of global image feature vectors and text feature vectors of the multiple organs. The target organ image and the image of at least one related organ are input into the multi-organ feature extraction sub-model, and a second feature vector and at least one third feature vector are output. The second feature vector is the target organ feature vector enhanced by the feature vector of the related organ of at least one related organ, and the third feature vector is the related organ feature vector fused with the target organ feature vector. The first feature vector, the second feature vector, and the at least one third feature vector are fused to obtain a fused feature vector, which is used to describe the features of the multiple organs from both global and local perspectives. The fused feature vector is input into the classifier, and the target recognition result is output.

4. The method according to claim 3, characterized in that, The multimodal feature extraction sub-model includes an image encoder, a text encoder, a first dimensionality reduction module, and a second dimensionality reduction module. The step of inputting the target image and the target text information into the multimodal feature extraction sub-model and outputting a first feature vector includes: The target image is input into the image encoder, and the global image feature vector is output. The target text information is input into a text encoder, which outputs the text feature vector. The global image feature vector is input into the first dimensionality reduction module, and the dimensionality-reduced global image feature vector is output. The text feature vector is input into the second dimensionality reduction module, and the dimensionality-reduced text feature vector is output. The first feature vector is obtained by concatenating the dimensionality-reduced global image features and the dimensionality-reduced text feature vector.

5. The method according to claim 3, characterized in that, The multi-organ feature extraction sub-model includes a target organ feature extraction unit and at least one related organ feature extraction unit. Each target organ feature extraction unit and at least one related organ feature extraction unit includes N-level feature extraction modules. An organ interaction module is provided between each level feature extraction module of each related organ feature extraction unit and the same level feature extraction module of the target organ feature extraction unit. N is a positive integer. For any M-th level feature extraction module of a related organ feature extraction unit, the target organ interaction module is greater than or equal to 1 and less than or equal to N. The target organ interaction module is used to perform attention calculation on the M-th level related organ feature vector and the M-th level target organ feature vector to obtain the M-th level enhanced feature vector corresponding to the related organ feature extraction unit. The M-th level related organ feature vector is the related organ feature vector output by the M-th level feature extraction module of the related organ feature extraction unit, and the M-th level target organ feature vector is the target organ feature vector output by the M-th level feature extraction module of the target organ feature extraction unit. The target organ interaction module is also used to perform linear projection on the feature vector of the Mth level target organ to obtain the projected feature vector of the Mth level target organ.

6. The method according to claim 5, characterized in that, The step of inputting the target organ image and the image of at least one related organ into the multi-organ feature extraction sub-model, and outputting a second feature vector and at least one third feature vector, includes: The target organ image is input into the N-level feature extraction module of the target organ feature extraction unit, and the at least one related organ image is input into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the second feature vector and the at least one third feature vector are output.

7. The method according to claim 6, characterized in that, The process involves inputting the target organ image into the N-level feature extraction module of the target organ feature extraction unit, and inputting the at least one related organ image into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the second feature vector is output, including: Obtain the feature vector of the Mth level target organ and the feature vector of the Mth level related organs; The M-th level target organ feature vector and the M-th level related organ feature vector are input into the target organ interaction module, and the M-th level enhanced feature vector is output. The Mth level target organ feature vector, the Mth level enhanced feature vector, and the Mth level enhanced feature vectors corresponding to other relevant organ feature extraction units are fused to obtain the Mth level target organ fused feature vector. The Mth level target organ fusion feature vector is input into the M+1 level feature extraction module of the target organ feature extraction unit, and the M+1 level target organ feature vector is output. The M+1 level target organ feature vector is processed according to the processing method of the M level target organ feature vector until the N level target organ feature vector output by the N level feature extraction module of the target organ feature extraction unit is obtained, and the N level target organ feature vector is used as the second feature vector.

8. The method according to claim 6, characterized in that, The process involves inputting the target organ image into the N-level feature extraction module of the target organ feature extraction unit, and inputting the at least one related organ image into the N-level feature extraction module of the corresponding related organ feature extraction unit. After processing by the N-level feature extraction module of the target organ feature extraction unit, the N-level feature extraction module of at least one related organ feature extraction unit, and the corresponding organ interaction module, the at least one third feature vector is output, including: Obtain the feature vector of the Mth level target organ and the feature vector of the Mth level related organs; The M-th level target organ feature vector and the M-th level related organ feature vector are input into the target organ interaction module, and the projected M-th level target organ feature vector is output. The projected M-th level target organ feature vector is fused with the M-th level related organ feature vector to obtain the M-th level related organ fused feature vector; The Mth-level related organ fusion feature vector is input into the M+1-level feature extraction module of the related organ feature extraction unit, and the M+1-level related organ feature vector is output. Following the processing method for the M-th level related organ feature vector, the (M+1)-th level related organ feature vector is processed until the N-th level related organ feature vector output by the N-th level feature extraction module of the related organ feature extraction unit is obtained, and the N-th level related organ feature vector is used as the third feature vector corresponding to the related organ feature extraction unit.

9. An image recognition method, characterized in that, The method is applied to cloud-side devices, and the method includes: The receiving end device sends an image processing task, the image processing task including a target image, the target image being obtained by capturing multiple organs of a target human body, the multiple organs including the target organ and at least one related organ adjacent to the target organ; The target organ image and at least one related organ image are segmented from the target image, and target text information describing the volume of the multiple organs is generated. The target image, the target text information, the target organ image, and at least one related organ image are input into the image recognition model, and the target recognition result of the target organ is output. The target identification result is sent to the terminal device.

10. A method for training an image recognition model, characterized in that, The method includes: Acquire sample images and an image recognition model to be trained. The sample images are obtained by photographing multiple organs of the human body. The multiple organs include the target organ and at least one related organ adjacent to the target organ. The image recognition model includes a multimodal feature extraction sub-model, a multi-organ feature extraction sub-model, and a classifier. The target organ image and at least one related organ image are segmented from the sample image, and sample text information describing the volume of the multiple organs is generated. The sample image and the sample text information are input into the multimodal feature extraction sub-model, and the first sample feature vector is output. The target organ image and at least one related organ image are input into the multi-organ feature extraction sub-model, and the second sample feature vector and at least one third sample feature vector are output. The first sample feature vector, the second sample feature vector, and the at least one third sample feature vector are fused to obtain a sample fusion feature vector. The sample fusion feature vector is input into the classifier, and the predicted identification result of the target organ is output. The image recognition model is trained based on the estimated recognition result, the actual recognition result corresponding to the sample image, the second sample feature vector, and the at least one third sample feature vector.

11. A computer-aided diagnostic method for esophageal varices, characterized in that, The method includes: A target chest image is acquired by photographing multiple organs of the target chest, including the esophagus and at least one related organ adjacent to the esophagus. The esophagus image and at least one related organ image are segmented from the target chest image, and target text information describing the volume of the multiple organs is generated. The target chest image, the target text information, the esophageal image, and at least one related organ image are input into the image recognition model, and the recognition result of the esophageal varices level is output.

12. An electronic device, characterized in that, It includes a processor and a memory; the memory stores at least one piece of program code; the at least one piece of program code is used to be called and executed by the processor to implement the image recognition method as described in any one of claims 1 to 8, or the image recognition method as described in claim 9, or the image recognition model training method as described in claim 10, or the computer-aided diagnosis method for esophageal varices as described in claim 11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which, when executed by a processor, is capable of implementing the image recognition method as described in any one of claims 1 to 8, or the image recognition method as described in claim 9, or the image recognition model training method as described in claim 10, or the computer-aided diagnostic method for esophageal varices as described in claim 11.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, can implement the image recognition method as described in any one of claims 1 to 8, or the image recognition method as described in claim 9, or the image recognition model training method as described in claim 10, or the computer-aided diagnostic method for esophageal varices as described in claim 11.