Method and device for predicting prognosis of cerebral hemorrhage, electronic equipment and storage medium

By using multi-layer downsampling layers and the SAM-CLIP cross-modal interaction module to generate effective masks in cerebral hemorrhage image prediction, combined with a multi-task feature fusion module, the problem of inaccurate prediction caused by segmentation errors in traditional methods is solved, and the prediction accuracy and reliability of the treatment plan are improved.

CN118866403BActive Publication Date: 2025-10-10LONGGANG DISTRICT CENT HOSPITAL OF SHENZHEN +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410802957.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-10-10
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

In traditional cerebral hemorrhage prediction methods, there are slight errors in the deep learning network segmentation of cerebral hemorrhage areas, resulting in large errors in the prediction results, affecting the accuracy of the treatment plan.

Method used

By inputting cerebral hemorrhage images into multi-layer downsampling layers to generate image features of multiple scales, and using the SAM-CLIP cross-modal interaction module to generate effective masks, combined with the multi-task feature fusion module to generate segmentation output information and classification feature output information, and finally generate prediction results.

Benefits of technology

It improves the accuracy of prognosis prediction for cerebral hemorrhage, assists medical staff in formulating feasible treatment plans, and reduces the risk of misdiagnosis and missed diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118866403B_ABST
    Figure CN118866403B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing. The present application discloses a kind of cerebral hemorrhage prognosis prediction method, device, electronic equipment and storage medium cerebral hemorrhage prognosis prediction method, it can improve the accuracy of cerebral hemorrhage prognosis prediction. The cerebral hemorrhage prognosis prediction method cerebral hemorrhage prognosis prediction method includes obtaining cerebral hemorrhage image;The cerebral hemorrhage image is input into the multilayer down-sampling layer in deep learning network, generates multiple different scale image features;All the image features are input into SAM-CLIP cross-modal interaction module in deep learning network, generate the effective mask corresponding to each image feature;All the image features and all the effective mask are input into the multi-task feature fusion module in deep learning network for feature fusion, generate segmentation output information and classification feature output information;The segmentation output information and the classification feature output information are input into the corresponding downstream task module in deep learning network, generate prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology. More specifically, the embodiments of the present application relate to a method, device, electronic device, and storage medium for predicting the prognosis of cerebral hemorrhage. Background Art

[0002] Traditional methods for predicting the prognosis of intracerebral hemorrhage (ICH) involve segmenting ICH images using a deep learning network to obtain the ICH region within the image. The N-dimensional image features within the ICH region are then extracted using the deep learning network. Finally, the N-dimensional image features are fused with the M-dimensional image features within the entire ICH image. These features are then input into a pre-trained ICH prediction model to obtain the prediction results. This method integrates the N-dimensional image features of the local image within the ICH image with the M-dimensional image features of the overall image within the ICH image as necessary factors for fusion to obtain the prediction results. However, due to the complexity of the lesions, the ICH region (i.e., the local image) obtained through deep learning network segmentation may contain subtle errors. Consequently, the prediction results derived from analyzing the local and overall images are inaccurate, leading medical personnel to formulate incorrect treatment plans based on these inaccurate prediction results, endangering the patient's life. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a method, device, electronic device, and storage medium for predicting the prognosis of cerebral hemorrhage, which can improve the accuracy of the prognosis prediction of cerebral hemorrhage. The embodiments of the present application are mainly achieved through the following technical solutions:

[0004] A first aspect of the embodiments of the present application provides a method for predicting the prognosis of cerebral hemorrhage, comprising:

[0005] Obtain images of intracerebral hemorrhage;

[0006] Inputting the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales;

[0007] Input all the image features into the SAM-CLIP (Segment Anything Model-Contrastive Language-Image Pre-training) cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each of the image features;

[0008] Inputting all the image features and all the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information;

[0009] The segmentation output information and the classification feature output information are input into the corresponding downstream task module in the deep learning network to generate a prediction result.

[0010] According to one embodiment of the present application, the step of acquiring an image of cerebral hemorrhage includes:

[0011] Get the image to be processed;

[0012] The image to be processed is filtered using a window selection algorithm to obtain the cerebral hemorrhage image.

[0013] According to one embodiment of the present application, the step of acquiring an image to be processed includes:

[0014] Obtaining a DICOM (Digital Imaging and Communications in Medicine) format data table, wherein the DICOM format data table includes a pixel sequence having cerebral hemorrhage image information, patient information, examination type, and equipment parameter information;

[0015] A pixel sequence is read from the DICOM format data table as the image to be processed.

[0016] According to one embodiment of the present application, the step of filtering the image to be processed using a window selection algorithm to obtain the cerebral hemorrhage image includes:

[0017] Set window position information;

[0018] Set window width information;

[0019] The image to be processed is filtered according to the window level information and the window width information to generate the cerebral hemorrhage image.

[0020] According to one embodiment of the present application, after the step of acquiring the cerebral hemorrhage image, the method for predicting the prognosis of cerebral hemorrhage further includes:

[0021] Obtain clinical text information, segmentation prompt box information and rough segmentation mask;

[0022] Determining data to be processed based on the clinical text information, the segmentation prompt box information and the rough segmentation mask;

[0023] The data to be processed is used to be input into the SAM-CLIP cross-modal interaction module together with all the image features to obtain a valid mask corresponding to each of the image features.

[0024] According to one embodiment of the present application, the step of determining the data to be processed based on the clinical text information, the segmentation prompt box information and the rough segmentation mask includes:

[0025] The clinical text information, the segmentation prompt box information and the rough segmentation mask are subjected to data format adjustment to generate the to-be-processed data.

[0026] According to an embodiment of the present application, after the step of inputting the brain hemorrhage image into a plurality of down-sampling layers in a deep learning network to generate a plurality of image features of different scales, the brain hemorrhage prognosis prediction method further comprises:

[0027] The plurality of image features are subjected to bilinear interpolation processing to input the SAM-CLIP cross-modal interaction module to generate an effective mask corresponding to each image feature.

[0028] According to an embodiment of the present application, the step of inputting all the image features into a SAM-CLIP cross-modal interaction module in a deep learning network to generate an effective mask corresponding to each image feature specifically comprises:

[0029] All the image features are processed by using a predetermined algorithm to generate a CLIP mask corresponding to each image feature.

[0030] All the CLIP masks are subjected to refinement processing by using a SAM model to generate the effective mask.

[0031] According to an embodiment of the present application, the calculation formula of the predetermined algorithm is:

[0032]

[0033] wherein, mask clip represents a CLIP mask, Norm represents a residual error output by a normalization layer, Atention() represents an attention mechanism, Q i is formed by applying a maximum pooling and a convolution layer to an image feature, K t and V t are obtained by processing a text feature by using two different convolution layers, represents an image encoder of CLIP processing an image feature.

[0034] According to an embodiment of the present application, the calculation formula of the SAM model performing refinement processing on all the CLIP masks is:

[0035]

[0036] wherein, mask sam-clip represents an effective mask, represents a SAM mask decoder, Conv represents a convolution layer processing function, mask cliprepresents the CLIP mask, i represents the image features, b is a bounding box composed of the segmentation prompt box and the rough segmentation mask; p is a set of randomly selected points composed of the segmentation prompt box and the rough segmentation mask, It is the SAM prompt word encoder.

[0037] A second aspect of the embodiments of the present application provides a device for predicting the prognosis of cerebral hemorrhage, comprising:

[0038] A first acquisition module is used to acquire cerebral hemorrhage images;

[0039] An image feature generation module, configured to input the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales;

[0040] An effective mask generation module is used to input all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each image feature;

[0041] A segmentation output information and classification feature output information generation module is used to input all the image features and all the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information;

[0042] The prediction result generation module is used to input the segmentation output information and the classification feature output information into the corresponding downstream task module in the deep learning network to generate a prediction result.

[0043] In a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the steps of the above-mentioned method for predicting the prognosis of cerebral hemorrhage.

[0044] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the steps of the above-mentioned method for predicting the prognosis of cerebral hemorrhage.

[0045] The beneficial effects of the embodiments of the present application include:

[0046] The present invention provides a method for predicting the prognosis of cerebral hemorrhage by inputting a cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features of different scales. These multiple image features of different scales are then input into the SAM-CLIP cross-modal interaction module in the deep learning network to generate corresponding valid masks. Then, these multiple image features of different scales and the valid masks corresponding to each image feature are used as necessary factors for fusion. The multi-task feature fusion module in the deep learning network is used to perform feature fusion on these two necessary factors to generate segmentation output information and classification feature output information. Finally, the segmentation output information and classification feature output information are input into the corresponding downstream task module in the deep learning network to generate a prediction result. Compared with the existing technology that uses the N-dimensional image features of the local image in the cerebral hemorrhage image and the M-dimensional image features of the overall image in the cerebral hemorrhage image as necessary factors for fusion, the present invention can avoid the errors caused by local image processing, improve the accuracy of cerebral hemorrhage prognosis prediction, thereby assisting medical personnel in formulating feasible treatment plans and increasing the probability of successful treatment of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 This is a diagram of application scenarios of the method for predicting the prognosis of cerebral hemorrhage of the present invention in some embodiments;

[0049] Figure 2 A flowchart of a method for predicting the prognosis of cerebral hemorrhage according to some embodiments of the present invention;

[0050] Figure 3 is a schematic diagram of cerebral hemorrhage images in some embodiments of the present invention;

[0051] Figure 4 A flowchart of a method for predicting the prognosis of cerebral hemorrhage according to some embodiments of the present invention;

[0052] Figure 5 Schematic diagram of segmentation prompt box information in some embodiments of the present invention;

[0053] Figure 6 Schematic diagram of a coarse segmentation mask in some embodiments of the present invention;

[0054] Figure 7 A flowchart of another embodiment of the method for predicting the prognosis of cerebral hemorrhage according to the present invention;

[0055] Figure 8 Flow chart of the method for predicting the prognosis of cerebral hemorrhage of the present invention in some further embodiments;

[0056] Figure 9 This is a flowchart of the SAM model in the present invention performing refinement processing on all CLIP masks;

[0057] Figure 10 This is a flowchart of the multi-task feature fusion module in the present invention fusing all image features and all valid masks;

[0058] Figure 11 A block diagram of a device for predicting the prognosis of cerebral hemorrhage according to some embodiments of the present invention;

[0059] Figure 12 FIG. 5 is a block diagram of an electronic device according to some embodiments of the present invention. DETAILED DESCRIPTION

[0060] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0061] In the description of the present application, “multi-layer” means at least two layers, such as two layers, three layers, etc., unless otherwise specifically defined.

[0062] Unless otherwise defined, all technical and scientific terms used in the specification of this application have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in the specification of this application includes any and all combinations of one or more of the relevant listed items.

[0063] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0064] In this document, the terms“include,”“includes” or“including” are used to mean“including but not limited to,”“comprising” or“comprising but not limited to,” and connotes the phrase“and / or,” so as to affirm the inclusion of some but not all phenomena, such as a process, method, procedure, system, product or apparatus that includes a list of steps or elements, but not only those steps or elements concretely listed.

[0065] Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. These functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks, processor devices or microcontroller devices.

[0066] Figure 1 An application scenario diagram for the prediction method of cerebral hemorrhage prognosis provided in an embodiment is shown in FIG. 1, which includes a terminal 1 and a server 2. Figure 1

[0067] In some embodiments, the server 2 can be used to train a deep learning network. After the server 2 obtains the trained deep learning network, it can be deployed in a cerebral hemorrhage prediction application. The terminal 1 can install the cerebral hemorrhage prediction application. When the terminal 1 obtains a cerebral hemorrhage image, the user can issue a cerebral hemorrhage prediction instruction through corresponding operation, and the terminal 1 can receive the cerebral hemorrhage prediction instruction, perform cerebral hemorrhage prediction on the cerebral hemorrhage image, and obtain a cerebral hemorrhage prediction result.

[0068] The terminal 1 includes, but is not limited to, a desktop computer, a notebook computer, a tablet computer, a mobile phone, a smart watch and the like.

[0069] The server 2 is a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0070] The cerebral hemorrhage prediction application is a disease treatment or prevention application.

[0071] ​The deep learning network is a general term for a class of pattern analysis methods, which mainly include three types of methods: convolutional neural networks (CNNs), autoencoder neural networks based on multi-layer neurons, and deep belief networks (DBNs). This application uses the deep learning network to train the multi-layer downsampling layers, SAM-CLIP cross-modal interaction module, multi-task feature fusion module, and downstream task modules required by this application.

[0072] In other implementations, the deep learning network can be trained by the terminal 1.

[0073] It is understandable that the above application scenario is only an example and does not constitute a limitation on the prediction method for cerebral hemorrhage prognosis and the deep learning network training method provided in the embodiments of the present application.

[0074] The specific implementation process of the embodiment of the present application is described in detail below.

[0075] like Figure 2 The figure shows a flow chart of the method for predicting the prognosis of cerebral hemorrhage provided in the embodiment of the present application. Figure 2 The method for predicting the prognosis of cerebral hemorrhage includes:

[0076] S10. Obtain images of cerebral hemorrhage.

[0077] The cerebral hemorrhage images can be referred to Figure 3 shown.

[0078] S20: Input the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features of different scales.

[0079] Specifically, refer to Figure 4 As shown, in the embodiment of the present application, the multiple downsampling layers include a first downsampling layer, a second downsampling layer, and a third downsampling layer. In other embodiments, the number of the downsampling layers can be determined by those skilled in the art according to actual needs.

[0080] The multiple image features of different scales include a first image feature, a second image feature, a third image feature, and a fourth image feature. The first image feature has a resolution of 512 dpi, the second image feature has a resolution of 256 dpi, the third image feature has a resolution of 128 dpi, and the fourth image feature has a resolution of 64 dpi. Of course, in other embodiments, the resolution of each image feature can be determined by those skilled in the art based on actual needs.

[0081] In the examples of this application, reference is made to Figure 4 The implementation of step S20 can also be to use the cerebral hemorrhage image as the first image feature I0 ; Input the first image feature into the first downsampling layer to generate the second image feature I 1 ; Input the second image feature into the second downsampling layer to generate the third image feature I 2 ; Input the third image feature into the third downsampling layer to generate the fourth image feature I 3 .

[0082] In other embodiments, step S20 is implemented by using the cerebral hemorrhage image as the first image feature; inputting the cerebral hemorrhage image into the first downsampling layer to generate the second image feature; inputting the cerebral hemorrhage image into the second downsampling layer to generate the third image feature; and inputting the cerebral hemorrhage image into the third downsampling layer to generate the fourth image feature.

[0083] The embodiment of the present application uses image features of multiple different scales in order to extract lesion features of different scales, so that the present application can achieve a better balance between the prediction accuracy of the prediction results and the network efficiency of the prediction method.

[0084] S30. Input all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each of the image features.

[0085] For details, see Figure 4 As shown, the first image feature I 0 Input the SAM-CLIP cross-modal interaction module to generate the first valid mask VM 0 , the second image feature I 1 Input the SAM-CLIP cross-modal interaction module to generate the second effective mask VM 1 , the third image feature I 2 Input the SAM-CLIP cross-modal interaction module to generate the third effective mask VM 2 , the fourth image feature I 3 Input the SAM-CLIP cross-modal interaction module to generate the fourth effective mask VM 3 It should be understood that the Figure 4 The four SAM-CLIP cross-modal interaction modules in are modules of the same structure.

[0086] S40: Input all the image features and all the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion, and generate segmentation output information and classification feature output information.

[0087] For example, Figure 4 As shown, the first valid mask VM0 , the first image feature I 0 , the second effective mask VM 1 , the second image feature I 1 , the third effective mask VM 2 , the third image feature I 2 , the fourth effective mask VM 3 and the fourth image feature I 3 All are input into the multi-task feature fusion module for feature fusion to generate segmentation output information S and classification feature output information F. The segmentation output information S and classification feature output information F are both image information.

[0088] S50: Input the segmentation output information and the classification feature output information into a corresponding downstream task module in a deep learning network to generate a prediction result.

[0089] refer to Figure 4 After the segmentation output information S enters the downstream task module, the downstream task module adds the segmentation output information S and the cerebral hemorrhage image to obtain marked image information.

[0090] refer to Figure 4 , after the classification feature output information F enters the downstream task module, it is input into the DenseNet-121 classifier (i.e. Figure 4 ), and obtain a binary classification label. The binary classification label includes a good prognosis and a poor prognosis. In other embodiments, the binary classification label can be determined by those skilled in the art based on actual needs and is not specifically limited to the two contents of good prognosis and poor prognosis mentioned above.

[0091] It can be understood that the prediction result includes the labeled image information and the binary classification label.

[0092] It should also be understood that the classification feature output information F is output by the multi-task feature fusion module and is used for the deep learning feature vector of the classification task. If the classification feature output information F is simply processed by the fully connected layer, it may not be possible to fully utilize this feature. Therefore, the embodiment of the present application adopts the DenseNet-121 architecture to fully learn this feature, thereby obtaining a more accurate classification result, while also avoiding overfitting of the downstream task model.

[0093] In the above embodiment, the embodiment of the present application generates multiple image features of different scales by inputting the cerebral hemorrhage image into the multi-layer downsampling layer in the deep learning network, and then uses these multiple image features of different scales to input the SAM-CLIP cross-modal interaction module in the deep learning network to generate corresponding valid masks. Then, using these multiple image features of different scales and the valid masks corresponding to each image feature as necessary factors for fusion, the multi-task feature fusion module in the deep learning network is used to perform feature fusion on these two necessary factors to generate segmentation output information and classification feature output information. Finally, the segmentation output information and classification feature output information are input into the corresponding downstream task module in the deep learning network to generate a prediction result. Compared with the existing technology that uses the N-dimensional image features of the local image in the cerebral hemorrhage image and the M-dimensional image features of the overall image in the cerebral hemorrhage image as necessary factors for fusion, the embodiment of the present application can avoid the errors caused by local image processing, improve the accuracy of cerebral hemorrhage prognosis prediction, thereby assisting medical personnel in formulating feasible treatment plans and increasing the probability of successful treatment of patients.

[0094] In some embodiments, step S10 includes:

[0095] S11. Obtain an image to be processed.

[0096] S12. Filter the image to be processed using a window selection algorithm to obtain the cerebral hemorrhage image.

[0097] Furthermore, step S11 includes:

[0098] S111 . Obtain a DICOM format data table, where the DICOM format data table includes a pixel sequence having cerebral hemorrhage image information, patient information, examination type, and equipment parameter information.

[0099] The DICOM format data table is obtained from CT hardware equipment.

[0100] The purpose of adopting the DICOM format data table in the embodiment of the present application is that it can improve the accuracy of the prediction of the present application and avoid the occurrence of misdiagnosis and missed diagnosis by medical staff. The traditional film review method cannot display images of all levels and cannot clearly display tiny lesions, while the original images in the DICOM format are clearer and contain more comprehensive information. In addition, the images in the DICOM format contain all the information of the examination. According to different accuracies or positions, the images are classified into different sequences. Medical staff can clearly see the size, position, morphology of the lesions in the image, and the relationship with the surrounding important organs, blood vessels, nerves and other information. Therefore, the use of the DICOM format data table is conducive to improving the prediction accuracy of the prediction method for the prognosis of cerebral hemorrhage.

[0101] The patient information includes the patient's name, age, gender, medical history and other information.

[0102] The examination types include CT examination, CR / DR examination, MR examination, nuclear medicine and ultrasound examination. In this application, the examination type is CT examination.

[0103] The device parameter information includes device information of the CT hardware device.

[0104] S112: Read a pixel sequence from the DICOM format data table as the image to be processed.

[0105] Furthermore, step S12 includes:

[0106] S121. Set window level information.

[0107] The window level information is the median grayscale value. The median is the value in the middle of all grayscale levels in a digital image. When the number of grayscale levels is even, the average of the two middle grayscale values ​​is taken. For example, if the grayscale levels are 188, 176, 171, 166, and 160, the median is 171; if the grayscale levels are 188, 176, 166, and 160, the median is 171.

[0108] S122: Set window width information.

[0109] The window width information is a grayscale range. The range is from 0 to 255. Grayscale refers to the color depth of a point in a black and white image, where white is 255 and black is 0. The grayscale value or grayscale level of a grayscale image is used to represent the brightness or intensity of each pixel in the image. Of course, those skilled in the art can also set the range as needed.

[0110] S123: Filter the image to be processed according to the window level information and the window width information to generate the cerebral hemorrhage image.

[0111] In some technical solutions, after step S10, the method for predicting the prognosis of cerebral hemorrhage further includes:

[0112] S60: Acquire clinical text information, segmentation prompt box information, and a rough segmentation mask.

[0113] The clinical text information includes the patient's name, gender, age, occupation and medical history, etc.

[0114] The segmentation prompt box information is a pixel range. The pixel range is composed of a pixel start point and a pixel end point. The pixel range is within the overall pixel range of the cerebral hemorrhage image and is located in the approximate area of ​​the bleeding location in the cerebral hemorrhage image. The segmentation prompt box information can be referenced Figure 5 Shown in red box.

[0115] The rough segmentation mask refers to an image that accurately separates the cerebral hemorrhage area from the background of the cerebral hemorrhage image, assigns a label to each pixel in the image, and the two-dimensional matrix formed by the labels of all pixels is the rough segmentation mask. The rough segmentation mask can be referred to Figure 6 shown.

[0116] S70. Determine the data to be processed based on the clinical text information, the segmentation prompt box information, and the rough segmentation mask; wherein the data to be processed is used to input the SAM-CLIP cross-modal interaction module together with all the image features to obtain a valid mask corresponding to each of the image features.

[0117] Exemplary, reference Figure 7 As shown, the first image feature I 0 , the second image feature I 1 , the third image feature I 2 , the fourth image feature I 3 , the clinical text information, the segmentation prompt box information and the rough segmentation mask are input into the SAM-CLIP cross-modal interaction module to obtain the first image feature I 0 The corresponding first valid mask VM 0 , and the second image feature I 1 The corresponding second effective mask VM 1 , and the third image feature I 2 The corresponding third valid mask VM 2 and the fourth image feature I 3 The corresponding fourth valid mask VM 3 .

[0118] In the embodiment of the present application, the data to be processed and all the image features are input into the SAM-CLIP cross-modal interaction module to obtain an effective mask, which can be more precise and further improves the prediction accuracy of the prediction method.

[0119] Furthermore, step S70 includes:

[0120] S71 , adjusting the data format of the clinical text information, the segmentation prompt box information, and the rough segmentation mask to generate the data to be processed.

[0121] The data format adjustment refers to adjusting the clinical text information, the segmentation prompt box information, and the rough segmentation mask into a data format suitable for the input of the SAM-CLIP cross-modal interaction module. The specific data format can be determined by those skilled in the art according to actual needs.

[0122] In some embodiments, after step S20, the method for predicting the prognosis of cerebral hemorrhage further includes:

[0123] Perform bilinear interpolation processing on the plurality of image features, and input them into the SAM-CLIP cross-modal interaction module to generate a valid mask corresponding to each of the image features. Figure 8 , respectively for the second image feature I 1 , the third image feature I 2 and the fourth image feature I 3 Perform bilinear interpolation processing. After this operation is completed, the first image feature I 0 , the second image feature I 1 The result after bilinear interpolation processing, the third image feature I 2 The result after bilinear interpolation processing and the fourth image feature I 3 The results after bilinear interpolation are input into the SAM-CLIP cross-modal interaction module. 0 The first valid mask VM is generated by the SAM-CLIP cross-modal interaction module 0 The second image feature I 1 The result after bilinear interpolation is processed by the SAM-CLIP cross-modal interaction module to generate the second effective mask VM 1 The third image feature I 2 The result after bilinear interpolation is processed by the SAM-CLIP cross-modal interaction module to generate the third effective mask VM 2 The fourth image feature I 3 The result after bilinear interpolation is processed by the SAM-CLIP cross-modal interaction module to generate the fourth effective mask VM 3 .

[0124] The purpose of adopting the bilinear interpolation processing method in the embodiment of the present application is that the image encoder of the SAM model in the SAM-CLIP cross-modal interaction module uses a standard visual decoder (Vision Transformer, ViT). Therefore, in the embodiment of the present application, each of the image features needs to be appropriately adjusted to the size of the visual decoder to adapt to the image encoder.

[0125] Bilinear interpolation is a mathematical interpolation method that performs linear interpolation in two directions. The core idea of ​​this method is to perform linear interpolation in two directions, first using linear interpolation in one direction and then using linear interpolation in the other direction to perform bilinear interpolation.

[0126] The basic steps of bilinear interpolation are: 1) linearly interpolating the corresponding image feature in the X direction to obtain the values ​​of the two intermediate points corresponding to the image feature; 2) linearly interpolating the image feature in the Y direction using the values ​​of the intermediate points obtained in the first step to finally obtain the value of the target point.

[0127] In some implementations, step S30 specifically includes:

[0128] S31 . Process all the image features using a predetermined algorithm to generate a CLIP mask corresponding to each image feature.

[0129] Furthermore, the calculation formula of the predetermined algorithm is:

[0130]

[0131] Among them, mask clip represents the CLIP mask, Norm represents the residual output of the normalized layer, Attention() represents the attention mechanism, Q i Formed by applying max pooling and convolution layers to image features, K t and V t Obtained by processing text features by two different convolutional layers, The image encoder representing CLIP processes image features.

[0132] Specifically, refer to Figure 9 As shown, it is an image encoder using CLIP (i.e. Figure 9 The CLIP ImageEncoder in the image encoder) processes four image features (i) at different scales and uses the text encoder of CLIP (i.e. Figure 9 Then, a special adaptive weighted attention mechanism is introduced to apply maximum pooling and convolutional layers (i.e. Figure 9 Conv in Q i , two different convolutional layers are used for clinical text information (i.e. Figure 9 Conv) in the process to get K t and V t Then, use the attention mechanism Attention and add Q i Connect the residual Norm output of the normalization layer to enhance cross-modal interaction and output a CLIP mask clip Then, under the guidance of the segmentation hint box information and the rough segmentation mask, the CLIP mask is refined using the SAM model to generate a valid mask.sam-clip .

[0133] S32: Use the SAM model to refine all the CLIP masks to generate the effective mask.

[0134] Furthermore, the calculation formula for the SAM model to refine all the CLIP masks is:

[0135]

[0136] Among them, mask sam-clip Represents the valid mask, Represents SAM mask decoder, Conv represents convolutional layer processing function, mask clip represents the CLIP mask, i represents the image features, b is a bounding box composed of the segmentation prompt box and the rough segmentation mask; p is a set of randomly selected points composed of the segmentation prompt box and the rough segmentation mask, It is the SAM prompt word encoder.

[0137] Specifically, the specific steps of the SAM model to refine the CLIP mask are as follows: first, the segmentation prompt box information and the rough segmentation mask are combined into a bounding box (that is, Figure 9 Bounding box in) and a set of randomly selected points (i.e. Figure 9 Randomly selected points in ), then, the SAM prompt word encoder (i.e. Figure 9 PromptEncoder in) becomes a constraint for bounding box and point processing, SAM mask decoder (i.e. Figure 9 The MaskDecoder in the convolutional layer converts the result of adding the rough segmentation mask processed by the convolutional layer to the cerebral hemorrhage image according to the constraints to generate the final effective mask mask. sam-clip (i.e. Figure 9 in the VM).

[0138] In some implementations, step S40 specifically includes:

[0139] S41. Use a group aggregation bridge (GAB) structure to merge the image features of each scale, its corresponding effective mask and its corresponding underlying features to obtain a fused feature corresponding to the graphic features of each scale, and merge all the fused features to form the classification feature output information.

[0140] For example, see Figure 10 As shown, the fourth effective mask VM3 and two fourth image features I 3 Input GAB to generate the fourth fusion feature R 3 , at this time, one of the fourth image features I 3 , which is the fourth image feature I 3 Corresponding underlying features; the third effective mask VM 2 , the third image feature I 2 and the fourth fusion feature R 3 Input GAB to generate the third fusion feature R 2 , at this time, the fourth fusion feature R 3 , which is the third image feature I 2 Corresponding underlying features; the second effective mask VM 1 , the second image feature I 1 and the third fusion feature R 2 Input GAB to generate the second fusion feature R 1 , at this time, the third fusion feature R 2 , which is the second image feature I 1 Corresponding underlying features; the first valid mask VM 0 , the first image feature I 0 and the second fusion feature R 1 Input GAB to generate the first fusion feature R 0 , at this time, the second fusion feature R 1 , which is the first image feature I 0 The corresponding underlying features. The first fusion feature R 0 The second fusion feature R 1 , the third fusion feature R 2 and the fourth fusion feature R 3 After merging, the classification feature output information F is formed.

[0141] Furthermore, the second fusion feature R 1 , the third fusion feature R 2 and the fourth fusion feature R 3 After bilinear interpolation reverse processing, it is combined with the first fusion feature R 0 The classification feature output information F is formed by merging.

[0142] The inverse bilinear interpolation process refers to the reverse operation of the bilinear interpolation. This reverse thinking can be understood as: the first value is obtained after a certain operation, and the second value is obtained after a certain operation is reversed.

[0143] The group aggregation bridge structure is an architecture for processing data aggregation and transmission, which can mix multiple data inputs and generate a mixed representation.

[0144] S42, pixel-by-pixel addition is performed on each fusion feature, and a convolution layer is applied for processing to generate a prediction mask, and one of the prediction masks is taken as the segmentation output information.

[0145] For example, referring to Figure 10 , the first fusion feature R 0 , the second fusion feature R 1 , the third fusion feature R 2 , and the fourth fusion feature R 3 perform pixel-by-pixel addition processing, and a convolution layer (i.e. Figure 10 Conv in the figure) is applied to generate a prediction mask. Among them, only the prediction mask corresponding to the first fusion feature R 0 can be taken as the segmentation output information S.

[0146] Further, the downstream task module is provided with a joint loss function for establishing the correlation between tasks and promoting the optimization of the model. Referring to Figure 10 , each prediction mask is optimized by the joint loss function to obtain a true value corresponding to each prediction mask.

[0147] The calculation formula of the joint loss function is:

[0148] Loss=L mta +αL seg +βL cla ;

[0149] Wherein, Loss is the result of the joint loss function; L mta is the consistency loss function of the segmentation task and the classification task; L seg is the overall segmentation loss function, which is composed of the weighted sum of four scale image features, the weight of the first image feature is 1, the weight of the second image feature is 0.75, the weight of the third image feature is 0.5, and the weight of the fourth image feature is 0.25; L cla is the classification loss function, which uses a weighted cross-entropy loss function to quantify the distance between the prediction result and the true label; α and β represent the weights of the overall segmentation loss function and the classification loss function respectively, the value of α is 0.2, and the value of β is 0.8. α and β are used to ensure the proper balance between segmentation and classification goals.

[0150] L mta The classification feature output information is processed through a softmax layer to obtain intermediate information, and then the Jensen-Shannon divergence is used to quantify the difference between the intermediate information and the segmentation output information.

[0151] L seg For each scale image feature, its Dice Similarity Coefficient (DSC) and Jaccard index are calculated to evaluate the consistency between the predicted masks.

[0152] The method embodiments of the present application are described in detail above, and the device embodiments of the present application are described in detail below Figure 11 , it should be understood that the device embodiments correspond to the method embodiments, and similar descriptions can be referred to the method embodiments.

[0153] Figure 11 A block diagram of a cerebral hemorrhage prognosis prediction device according to an embodiment of the present application is schematically shown. The cerebral hemorrhage prognosis prediction device can adopt a software unit or a hardware unit, or a combination of the two as part of a computer device. As Figure 11 shown, the cerebral hemorrhage prognosis prediction device 11 provided by the embodiment of the present application comprises:

[0154] The first acquisition module 81 is configured to acquire a cerebral hemorrhage image.

[0155] The image feature generation module 82 is configured to input the cerebral hemorrhage image into a plurality of multi-layer down-sampling layers in a deep learning network to generate a plurality of image features of different scales.

[0156] The effective mask generation module 83 is configured to input all the image features into a SAM-CLIP cross-modal interaction module in the deep learning network to generate an effective mask corresponding to each of the image features.

[0157] The segmentation output information and classification feature output information generation module 84 is configured to input all the image features and all the effective masks into a multi-task feature fusion module in the deep learning network for feature fusion to generate segmentation output information and classification feature output information.

[0158] The prediction result generation module 85 is configured to input the segmentation output information and the classification feature output information into a corresponding downstream task module in the deep learning network to generate a prediction result.

[0159] In some embodiments, the first acquisition module 81 comprises a to-be-processed image acquisition module and a cerebral hemorrhage image acquisition module.

[0160] The to-be-processed image acquisition module is configured to acquire a to-be-processed image.

[0161] The cerebral hemorrhage image acquisition module is configured to filter the to-be-processed image by using a window selection algorithm to obtain the cerebral hemorrhage image.

[0162] In some embodiments, the module for acquiring images to be processed includes:

[0163] an acquiring unit, configured to acquire a DICOM format data table, wherein the DICOM format data table includes a pixel sequence having cerebral hemorrhage image information, patient information, examination type, and equipment parameter information;

[0164] A reading unit is used to read a pixel sequence from the DICOM format data table as the image to be processed.

[0165] In some embodiments, the cerebral hemorrhage image acquisition module includes:

[0166] A first setting unit, used for setting window level information;

[0167] A second setting unit is used to set window width information;

[0168] A filtering unit is used to filter the image to be processed according to the window level information and the window width information to generate the cerebral hemorrhage image.

[0169] In some embodiments, the device for predicting the prognosis of cerebral hemorrhage 8 further comprises:

[0170] The second acquisition module is used to obtain clinical text information, segmentation prompt box information and rough segmentation mask;

[0171] a module for determining data to be processed, configured to determine data to be processed based on the clinical text information, the segmentation prompt box information, and the rough segmentation mask;

[0172] The data to be processed is used to be input into the SAM-CLIP cross-modal interaction module together with all the image features to obtain a valid mask corresponding to each of the image features.

[0173] In some embodiments, the module for determining data to be processed includes:

[0174] An adjustment unit is used to adjust the data format of the clinical text information, the segmentation prompt box information and the rough segmentation mask to generate the data to be processed.

[0175] In some embodiments, the device for predicting the prognosis of cerebral hemorrhage 8 further comprises:

[0176] A bilinear interpolation processing module is used to perform bilinear interpolation processing on the plurality of image features, so as to input the bilinear interpolation processing into the SAM-CLIP cross-modal interaction module to generate a valid mask corresponding to each of the image features.

[0177] In some embodiments, the effective mask generation module 83 includes:

[0178] A CLIP mask generating unit, configured to process all the image features using a predetermined algorithm to generate a CLIP mask corresponding to each of the image features;

[0179] The effective mask unit is used to refine all the CLIP masks using a SAM model to generate the effective mask.

[0180] The embodiment of the present application further provides an electronic device, the principle block diagram of the electronic device can be as follows Figure 12 As shown. The electronic device includes a processor, a memory, a network interface, a display screen and a temperature sensor connected via a system bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for predicting the prognosis of cerebral hemorrhage is implemented. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor is pre-set inside the electronic device to detect the operating temperature of the internal device.

[0181] Those skilled in the art will understand that Figure 12 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0182] In some implementations, an embodiment of the present application provides an electronic device including a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute instructions of the computer program stored in the memory to perform the following operations:

[0183] Obtain images of intracerebral hemorrhage;

[0184] Inputting the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales;

[0185] Inputting all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each of the image features;

[0186] Inputting all the image features and all the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information;

[0187] The segmentation output information and the classification feature output information are input into the corresponding downstream task module in the deep learning network to generate a prediction result.

[0188] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, wherein the computer program enables a computer to execute the steps of the above-mentioned method for predicting the prognosis of cerebral hemorrhage.

[0189] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described embodiments. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0190] The technical features of the above embodiments can be combined without changing the basic principles of the present invention. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0191] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the patent scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.

Claims

1. A method for predicting the prognosis of cerebral hemorrhage, characterized in that: include: Obtain images of intracerebral hemorrhage; Inputting the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales; Inputting all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each of the image features; Inputting all the image features and all the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information; all the image features include the fourth image feature, the third image feature, the second image feature, and the first image feature, and all the valid masks include the fourth valid mask, the third valid mask, the second valid mask, and the first valid mask; Input the segmentation output information and the classification feature output information into the corresponding downstream task module in the deep learning network to generate a prediction result; Inputting all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each of the image features specifically includes: processing all the image features using a predetermined algorithm to generate a CLIP mask corresponding to each of the image features; refining all the CLIP masks using a SAM model to generate the valid mask; The SAM model refines the CLIP mask as follows: first, the segmentation hint box information and the rough segmentation mask are combined into a bounding box and a set of randomly selected points. Then, the SAM hint word encoder processes the bounding box and points into constraints. The SAM mask decoder transforms the result of adding the convolutional CLIP mask to the image features according to the constraints to generate the final effective mask. The segmentation prompt box information is a pixel range, and the pixel range is composed of a pixel start point and a pixel end point; The rough segmentation mask is a two-dimensional matrix, and the process of forming the two-dimensional matrix is ​​to separate the cerebral hemorrhage area from the background of the cerebral hemorrhage image, assign a label to each pixel in the image, and the labels of all pixels form the two-dimensional matrix; Inputting all the image features and all the effective masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information includes: inputting a fourth effective mask and two fourth image features into a group aggregation bridge structure to generate a fourth fusion feature, wherein one fourth image feature is an underlying feature corresponding to the fourth image feature; inputting a third effective mask, a third image feature and a fourth fusion feature into a group aggregation bridge structure to generate a third fusion feature, wherein the fourth fusion feature is an underlying feature corresponding to the third image feature; inputting a second effective mask, a second image feature and a third fusion feature into a group aggregation bridge structure to generate a second fusion feature, wherein the third fusion feature is an underlying feature corresponding to the second image feature; inputting a first effective mask, a first image feature and a second fusion feature into a group aggregation bridge structure to generate a first fusion feature, wherein the second fusion feature is an underlying feature corresponding to the first image feature; performing bilinear interpolation inverse processing on the second fusion feature, the third fusion feature and the fourth fusion feature, and then merging them with the first fusion feature to form the classification feature output information; The downstream task module is provided with a joint loss function for establishing the correlation between tasks and promoting the optimization of the model; the calculation formula of the joint loss function is: ;in, is the result of the joint loss function; It is the consistency loss function for segmentation and classification tasks; is the overall segmentation loss function, which is composed of the weighted sum of the four scale image features. The weight of the first image feature is 1, the weight of the second image feature is 0.75, the weight of the third image feature is 0.5, and the weight of the fourth image feature is 0.

25. It is a classification loss function that uses a weighted cross entropy loss function to quantify the distance between the predicted result and the true label; and Represent the weights of the overall segmentation loss function and the classification loss function, respectively. The value of is 0.2, The value of is 0.8; and Used to ensure a balance between segmentation and classification objectives; The classification feature output information is processed through a softmax layer to obtain intermediate information. Subsequently, the Jensen-Shannon divergence is used to quantify the difference between the intermediate information and the segmentation output information; For each scale image feature, the Dice similarity coefficient and the Jaccard similarity coefficient are calculated to evaluate the consistency between the predicted masks.

2. The method for predicting the prognosis of cerebral hemorrhage according to claim 1, wherein: The steps to obtain an image of an intracerebral hemorrhage include: Acquire a DICOM format data table, wherein the DICOM format data table includes a pixel sequence having cerebral hemorrhage image information, patient information, examination type, and equipment parameter information; Reading a pixel sequence from the DICOM format data table as an image to be processed; Set window position information; Set window width information; The image to be processed is filtered according to the window level information and the window width information to generate the cerebral hemorrhage image.

3. The method for predicting the prognosis of cerebral hemorrhage according to claim 1, wherein: After the step of obtaining the cerebral hemorrhage image, the method for predicting the prognosis of cerebral hemorrhage further includes: Obtain clinical text information, segmentation prompt box information and rough segmentation mask; Adjusting the data format of the clinical text information, the segmentation prompt box information, and the rough segmentation mask to generate data to be processed; The data to be processed is used to be input into the SAM-CLIP cross-modal interaction module together with all the image features to obtain a valid mask corresponding to each of the image features.

4. The method for predicting the prognosis of cerebral hemorrhage according to claim 1, wherein: After the step of inputting the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales, the method for predicting the prognosis of cerebral hemorrhage further includes: Bilinear interpolation processing is performed on the plurality of image features, and the bilinear interpolation processing is used to input the SAM-CLIP cross-modal interaction module to generate a valid mask corresponding to each of the image features.

5. The method for predicting the prognosis of cerebral hemorrhage according to claim 1, wherein: The calculation formula of the predetermined algorithm is: ; in, Represents the CLIP mask, represents the residual of the normalization layer output, represents the attention mechanism, It is formed by applying max pooling and convolution layers to image features. and Obtained by processing text features by two different convolutional layers, The image encoder representing CLIP processes image features.

6. The method for predicting the prognosis of cerebral hemorrhage according to claim 1, wherein: The calculation formula for the SAM model to refine all the CLIP masks is: ; in, Represents the valid mask, stands for SAM mask decoder, Represents the convolutional layer processing function, Represents the CLIP mask, Represents image features, It is a bounding box synthesized by the segmentation hint box and the rough segmentation mask; is a set of randomly selected points synthesized by the segmentation hint box and the rough segmentation mask, It is the SAM prompt word encoder.

7. A device for predicting the prognosis of cerebral hemorrhage, characterized in that: include: A first acquisition module is used to acquire cerebral hemorrhage images; An image feature generation module, configured to input the cerebral hemorrhage image into a multi-layer downsampling layer in a deep learning network to generate multiple image features at different scales; An effective mask generation module is used to input all the image features into the SAM-CLIP cross-modal interaction module in the deep learning network to generate a valid mask corresponding to each image feature; a segmentation output information and classification feature output information generation module, configured to input all of the image features and all of the valid masks into a multi-task feature fusion module in a deep learning network for feature fusion to generate segmentation output information and classification feature output information; all of the image features include the fourth image feature, the third image feature, the second image feature, and the first image feature, and all of the valid masks include the fourth valid mask, the third valid mask, the second valid mask, and the first valid mask; A prediction result generation module, configured to input the segmentation output information and the classification feature output information into a corresponding downstream task module in a deep learning network to generate a prediction result; The device for predicting the prognosis of cerebral hemorrhage is further configured to process all the image features using a predetermined algorithm to generate a CLIP mask corresponding to each of the image features; and refine all the CLIP masks using a SAM model to generate the effective mask; The cerebral hemorrhage prognosis prediction device is further configured to first combine the segmentation prompt box information and the rough segmentation mask into a bounding box and a set of randomly selected points. Then, the bounding box and the points are processed into constraints by the SAM prompt word encoder. The SAM mask decoder converts the result of adding the CLIP mask processed by the convolution layer to the image features according to the constraints to generate a final effective mask. The segmentation prompt box information is a pixel range, and the pixel range is composed of a pixel start point and a pixel end point; The rough segmentation mask is a two-dimensional matrix, and the process of forming the two-dimensional matrix is ​​to separate the cerebral hemorrhage area from the background of the cerebral hemorrhage image, assign a label to each pixel in the image, and the labels of all pixels form the two-dimensional matrix; The segmentation output information and classification feature output information generation module is further used to implement a fourth effective mask and two fourth image feature input group aggregation bridge structure to generate a fourth fusion feature, wherein one fourth image feature is the underlying feature corresponding to the fourth image feature; a third effective mask, a third image feature and a fourth fusion feature input group aggregation bridge structure to generate a third fusion feature, and the fourth fusion feature is the underlying feature corresponding to the third image feature; a second effective mask, a second image feature and a third fusion feature input group aggregation bridge structure to generate a second fusion feature, and the third fusion feature is the underlying feature corresponding to the second image feature; a first effective mask, a first image feature and a second fusion feature input group aggregation bridge structure to generate a first fusion feature, and the second fusion feature is the underlying feature corresponding to the first image feature; after performing bilinear interpolation inverse processing on the second fusion feature, the third fusion feature and the fourth fusion feature, they are merged with the first fusion feature to form the classification feature output information; The downstream task module is provided with a joint loss function for establishing the correlation between tasks and promoting the optimization of the model; the calculation formula of the joint loss function is: ;in, is the result of the joint loss function; It is the consistency loss function for segmentation and classification tasks; is the overall segmentation loss function, which is composed of the weighted sum of the four scale image features. The weight of the first image feature is 1, the weight of the second image feature is 0.75, the weight of the third image feature is 0.5, and the weight of the fourth image feature is 0.

25. It is a classification loss function that uses a weighted cross entropy loss function to quantify the distance between the predicted result and the true label; and Represent the weights of the overall segmentation loss function and the classification loss function, respectively. The value of is 0.2, The value of is 0.8; and Used to ensure a balance between segmentation and classification objectives; The classification feature output information is processed through a softmax layer to obtain intermediate information. Subsequently, the Jensen-Shannon divergence is used to quantify the difference between the intermediate information and the segmentation output information; For each scale image feature, the Dice similarity coefficient and the Jaccard similarity coefficient are calculated to evaluate the consistency between the predicted masks.

8. An electronic device, characterized in that include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the steps of the method for predicting the prognosis of cerebral hemorrhage as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that Used to store a computer program, wherein the computer program enables a computer to execute the steps of the method for predicting the prognosis of cerebral hemorrhage according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Cerebral hemorrhage prognosis prediction method and device, electronic equipment and storage medium

    CN113724184A

  • Medical image segmentation method based on multiple attention fusion

    CN116152492A

  • Image indication segmentation method based on pre-training model migration and prompt learning

    CN117808819A