A medical image semantic segmentation method, device, electronic device and storage medium

Through the method of feature cross-alignment and optimization, the semantic segmentation accuracy of medical images of rare diseases is improved, the problem of insufficient training data is solved, and the diagnosis and treatment of rare diseases is provided.

CN119625325BActive Publication Date: 2025-07-22PEKING UNION MEDICAL COLLEGE HOSPITAL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163039.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-22
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing image semantic segmentation technology is not effective in the diagnosis of rare diseases, mainly due to the limited number of cases, resulting in insufficient training data.

Method used

Features of the patient's current and historical images are extracted through a feature extraction network, feature cross-alignment and optimization are performed, and features are fusion-optimized by using decoders to generate masked results, improving the segmentation accuracy of medical images for rare diseases.

Benefits of technology

It enhances the accuracy of semantic segmentation of medical images of rare diseases and provides assistance for the diagnosis and treatment of patients with rare diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625325B_ABST
    Figure CN119625325B_ABST
Patent Text Reader

Abstract

The present application provides a medical image semantic segmentation method, apparatus, electronic device, and storage medium, relating to the technical field of image processing. The method includes respectively extracting a first image feature corresponding to a first image and a plurality of second image features corresponding to at least one second image through a feature extraction network; performing feature cross-alignment based on the mean of the first image feature and the plurality of second image features to obtain a plurality of cross features; respectively optimizing the plurality of cross features based on the similarity between the first image feature and the plurality of second image features to obtain a plurality of optimized features; and fusing the first image feature, the plurality of optimized features, and the mask matched with each optimized feature through a decoder to obtain a mask result corresponding to the first image, so as to improve the accuracy of semantic segmentation of rare disease medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular, to a medical image semantic segmentation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous progress of medical imaging technology, image semantic segmentation technology is increasingly widely used in the field of medical diagnosis. This technology can automatically identify and segment different tissues and organs in medical images, thereby helping doctors make more accurate disease diagnoses and treatment plans. In the prior art, image semantic segmentation algorithms usually rely on a large amount of labeled data to train models to achieve high-precision segmentation effects. However, in the case of rare diseases, due to the limited number of cases, the available medical image samples are also relatively few, which makes it difficult for existing image semantic segmentation technologies to obtain sufficient training data, thereby affecting their performance in rare disease diagnoses. Therefore, under the existing technical conditions, the image semantic segmentation effect is often unsatisfactory when dealing with medical images related to rare diseases. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a medical image semantic segmentation method, apparatus, electronic device, and storage medium to improve the accuracy of semantic segmentation of rare disease medical images.

[0004] In a first aspect, the present invention provides a medical image semantic segmentation method, the method including respectively extracting a first image feature corresponding to a first image and a plurality of second image features corresponding to at least one second image through a feature extraction network, where the second image is an image taken during the historical medical treatment process of a target patient and semantically segmented for a medical semantic label, and the first image is an image of a pixel region with a medical semantic label currently taken by the target patient; performing feature cross-alignment based on the mean of the first image feature and the plurality of second image features to obtain a plurality of cross features; optimizing the plurality of cross features respectively based on the similarity between the first image feature and the plurality of second image features to obtain a plurality of optimized features; fusing the first image feature, the plurality of optimized features, and a mask matched with each optimized feature through a decoder to obtain a mask result corresponding to the first image, and the mask result is used to indicate a target pixel region belonging to the medical semantic label and a non-target pixel region not belonging to the medical semantic label in the first image; where, for the first image feature and each second image feature, the corresponding cross feature is obtained through the following method:

[0005] For each first element value at each position in the first image feature, calculate the product of the first element value and the second element value at the same position in the second image feature as the third element value at the corresponding position of the corresponding cross feature.

[0006] In an alternative embodiment, the feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block connected in sequence, where each convolutional block includes a first convolutional layer, an asymmetric convolutional layer, and a second convolutional layer connected in sequence.

[0007] In an alternative embodiment, for the first image feature and each second image feature, the corresponding cross feature is optimized in the following manner to obtain the corresponding optimized feature:

[0008] Based on the first image feature and the current second image feature, a first element value distance feature is determined; based on the first image feature and the remaining second image features, a second element value distance feature is determined; for each third element value at each position in the current cross feature, based on the first distance value at this position in the first element value distance feature and the second distance value at this position in the second element value distance feature, it is determined whether the third element value at this position needs to be optimized; if so, the third element value is modified based on the second element value at this position in the current second image feature to be used as the fourth element value at this position in the corresponding optimized feature.

[0009] In an alternative embodiment, the following method is used to determine whether the third element value at each position needs to be optimized:

[0010] Determine whether the first distance value corresponding to this position is greater than a first target value and whether the second distance value corresponding to this position is less than a second target value; if the first distance value is greater than the first target value and the second distance value is less than the second target value, it is determined that the third element value needs to be optimized.

[0011] In an alternative embodiment, the following method is used to calculate the fourth element value:

[0012] The product of multiplying the third element value at the corresponding position in the current cross feature by the amplification factor is used as the fourth element value.

[0013] In an alternative embodiment, the first element value distance feature is determined in the following manner:

[0014] For the first element value at each position of the first image feature and the second element value at the corresponding position of the current second image feature, they are decomposed to respectively obtain the first integer element value and the first decimal element value corresponding to the first element value, and the second integer element value and the second decimal element value corresponding to the second element value; calculate the first difference between the first integer element value and the second integer element value; calculate the second difference between the first decimal element value and the second decimal element value; calculate the sum of the absolute value of the first difference and the absolute value of the second difference to be used as the first distance value at the corresponding position in the first element value distance feature.

[0015] In an alternative embodiment, the second element distance feature is determined as follows:

[0016] For the second element values at the corresponding positions of the remaining second image features, calculate the average value of the second elements at this position; for the first element value at each position of the first image feature and the average value of the second elements at the corresponding position, calculate the cosine similarity as the second distance value at the corresponding position in the second element distance feature.

[0017] In a second aspect, the present invention provides a medical image semantic segmentation device, which includes:

[0018] An extraction module for respectively extracting the first image features corresponding to the first image and a plurality of second image features corresponding to at least one second image through a feature extraction network, wherein the second image is an image taken during the historical medical treatment of the target patient and semantically segmented for medical semantic labels, and the first image is an image of the target patient taken currently with a pixel region having medical semantic labels;

[0019] An alignment module for performing feature cross-alignment based on the average values of the first image features and the plurality of second image features to obtain a plurality of cross features; wherein, for the first image feature and each second image feature, the corresponding cross feature is obtained through the following method:

[0020] For each first element value at each position in the first image feature, calculate the product between the first element value and the second element value at this position of the second image feature as the third element value at the corresponding position of the corresponding cross feature;

[0021] An optimization module for respectively optimizing the plurality of cross features based on the similarity between the first image features and the plurality of second image features to obtain a plurality of optimized features;

[0022] A decoding module for fusing the first image features, the plurality of optimized features, and the masks matched with each optimized feature through a decoder to obtain a mask result corresponding to the first image, and the mask result is used to indicate the target pixel region belonging to the medical semantic label and the non-target pixel region not belonging to the medical semantic label in the first image.

[0023] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of any of the foregoing medical image semantic segmentation methods.

[0024] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of any of the medical image semantic segmentation methods in the foregoing embodiments.

[0025] A medical image semantic segmentation method, device, electronic device and storage medium provided by the present application take the segmented images and unsegmented images of a patient with the same medical semantic labels as inputs at the same time, extract features, align and optimize the extracted features, further enhance the description effect of some features on the medical semantic labels, and finally decode to obtain the segmentation result of the patient's medical image, improving the accuracy of medical semantic segmentation in the diagnosis and treatment of rare diseases and providing assistance for the diagnosis and treatment of rare disease patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a flowchart of a medical image semantic segmentation method provided by an embodiment of the present application;

[0028] Figure 2 It is a flowchart of the steps of a feature optimization provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic structural diagram of a medical image semantic segmentation device provided by an embodiment of the present application;

[0030] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] First, the application scenarios of the present application are described. The technical solutions of the present application can be applied to image semantic segmentation in medical diagnosis and treatment.

[0032] The following will describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0033] Embodiment 1

[0034] In the diagnosis and treatment process of rare disease patients, due to the complexity and individual differences of rare diseases, and the scarcity of case samples, if the model is trained and constructed in a conventional way, the segmentation effect is often not good. Therefore, in an embodiment of the present application, medical images of the same patient are used, specifically, medical images such as MRI and CT can be used.

[0035] Taking the medical semantic label of tumor as an example, first, the MRI image taken during the patient's historical medical treatment process is processed to generate a corresponding mask.

[0036] Figure 1 The flowchart of a medical image semantic segmentation method provided by an embodiment of the present application is as follows. As Figure 1 shown, an embodiment of the present application provides a medical image semantic segmentation method, which can be executed by a computer device. The method includes:

[0037] S1. Respectively extract the first image feature corresponding to the first image and multiple second image features corresponding to at least one second image through a feature extraction network.

[0038] Among them, the second image can be the brain MRI image taken during the target patient's historical medical treatment process and semantically segmented for the medical semantic label, and the first image is the brain MRI image of the target patient currently taken and having a pixel region with a medical semantic label.

[0039] Here, a feature extraction network based on Resnet is used to extract the features of the first image and the second image.

[0040] Specifically, the feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block connected in sequence. Among them, each convolutional block includes a first convolutional layer, an asymmetric convolutional layer, and a second convolutional layer connected in sequence. The convolutional kernel of the first convolutional block is 3×3×3, the convolutional kernel of the second convolutional block is 3×1×1, the convolutional kernel of the third convolutional block is 1×3×1, the convolutional kernel of the fourth convolutional block is 1×1×3, and the convolutional kernel of the fifth convolutional block is 1×1×1. Among them, the sizes of the first convolutional layer and the second convolutional layer are 3×3×3 and 1×1×1 respectively.

[0041] The first image feature and the second image feature here can be the features output by the fifth convolutional block. Taking two second images as an example, the first image feature can be obtained here , and the second image feature and .

[0042] Furthermore, in order to improve the accuracy of semantic segmentation, the features output by the third convolutional block, the fourth convolutional block, and the fifth convolutional block can all be extracted as the corresponding first image feature or second image feature.

[0043] S2. Perform feature cross - alignment based on the mean of the first image feature and multiple second image features to obtain multiple cross - features.

[0044] Among them, for the first image feature and each second image feature, the corresponding cross - feature is obtained through the following method:

[0045] For each first element value at each position in the first image feature, calculate the product between the first element value and the second element value at the same position in this second image feature, and use it as the third element value at the corresponding position of the corresponding cross - feature.

[0046] Taking the first image feature and the second image feature as an example, in step S2, feature cross - multiplication can be performed in the form of Hadamard product.

[0047] Calculate the product between the first element value at (1, 1, 1) in the first image feature and the second element value at (1, 1, 1) in the second image feature as the third element value at (1, 1, 1) in the cross - feature . And so on, finally obtaining the cross - feature .

[0048] By performing feature cross - multiplication between the features of the medical image to be semantically segmented and the medical image that has been semantically segmented, the description effect of some features on the medical semantic labels can be enhanced.

[0049] S3. Optimize multiple cross - features respectively based on the similarity between the first image feature and multiple second image features to obtain multiple optimized features.

[0050] As Figure 2 shown, taking the first image feature and the second image feature as an example, the corresponding cross - feature can be optimized through the following method to obtain the corresponding optimized feature:

[0051] S100. Determine the first - element - value distance feature based on the first image feature and the current second image feature.

[0052] In step S100, taking the first image feature and the second image feature as an example, the corresponding first - element - value distance feature can be determined through the following method:

[0053] The first element value at each position of the first image feature and the second element value at the corresponding position of the current second image feature are decomposed to obtain the first integer element value and the first decimal element value corresponding to the first element value, and the second integer element value and the second decimal element value corresponding to the second element value, respectively.

[0054] The first image feature The first element value at (1, 1, 1) Decomposes into the integer part as the first integer element value 1 and the fractional part as the first fractional element value 0.36.

[0055] The second image feature The second element value at (1, 1, 1) Decomposed into the integer part as the second integer element value 0 and the decimal part as the second decimal element value 0.51.

[0056] A first difference between the first integer element value and the second integer element value is calculated. A second difference between the first decimal element value and the second decimal element value is calculated. A sum of the absolute value of the first difference and the absolute value of the second difference is calculated as a first distance value at a corresponding position in the first element value distance feature.

[0057] Calculate the first distance value at (1, 1, 1) in the first element value distance feature By analogy, we get .

[0058] Here, Manhattan distance is used to calculate the pixel-level similarity between the first image feature and the second image feature, which can reduce the sensitivity to outliers and better capture the similarity between pixels.

[0059] S101. Determine a second element value distance feature based on the first image feature and the remaining second image features.

[0060] First image feature and the second image feature , For example, the corresponding second element distance feature can be determined in the following way:

[0061] For the second element values at the positions corresponding to the remaining second image features, the second element mean at the positions is calculated.

[0062] calculate , The second element value at (1, 1, 1) , The mean between is taken as the mean of the second element at (1, 1, 1) .

[0063] For each first element value at each position of the first image feature and the mean value of the second elements at the corresponding positions, calculate the cosine similarity as the second distance value at the corresponding position in the second element distance feature.

[0064] Calculate the first image feature The first element value at (1, 1, 1) in And the mean value of the second elements, and use the cosine similarity to calculate the second distance value at (1, 1, 1) in the second element distance feature :

[0065] .

[0066] And so on. After traversing all positions, finally, .

[0067] Calculating the second element value distance feature using the cosine similarity can reflect the directional consistency between pixels in multi-dimensional space and pay more attention to semantic similarity.

[0068] S102. For each third element value at each position in the current cross feature, based on the first distance value at this position in the first element value distance feature and the second distance value at this position in the second element value distance feature, determine whether the third element value at this position needs to be optimized. Specifically, the following method can be used to determine whether the third element value at each position needs to be optimized:

[0069] Determine whether the first distance value corresponding to this position is greater than the first target value and whether the second distance value corresponding to this position is less than the second target value. If the first distance value is greater than the first target value and the second distance value is less than the second target value, then determine that the third element value needs to be optimized.

[0070] For the pixel feature at a certain position, if the corresponding first distance value is greater than the first target value and the second distance value is less than the second target value, then optimize the third element value to better capture the similarity of the image feature.

[0071] S103. If necessary, modify the third element value based on the second element value at this position in the current second image feature as the fourth element value at the corresponding position in the corresponding optimized feature.

[0072] Multiply the third element value at the corresponding position in the current cross feature by the magnification factor as the fourth element value.

[0073] The magnification factor here can be preset, and the value range of the magnification factor is between 0 and 1.

[0074] In this way, the element values in the optimized features can enhance the pointing of the cross features to the semantic segmentation elements, so as to obtain a more accurate semantic segmentation result.

[0075] S4. The decoder fuses the first image feature, multiple optimized features, and the masks matched with each optimized feature to obtain a mask result corresponding to the first image, and the mask result is used to indicate the target pixel region belonging to the medical semantic label and the non-target pixel region not belonging to the medical semantic label in the first image.

[0076] In step S4, the TransUnet decoder can be used. It can not only provide the ability to retain spatial information, but also incorporates the Transformer mechanism to enhance the global context understanding. The mask result corresponding to the first image output by the decoder can indicate the pixel part belonging to the tumor and the pixel part not belonging to the tumor in the image.

[0077] A medical image semantic segmentation method provided by this application takes the segmented images and unsegmented images of the same patient with the same medical semantic label as inputs at the same time, extracts features, aligns and optimizes the extracted features, further enhances the description effect of some features on the medical semantic label, and finally decodes to obtain the segmentation result of the patient's medical image, improving the accuracy of medical semantic segmentation in the diagnosis and treatment of rare diseases and providing assistance for the diagnosis and treatment of rare disease patients.

[0078] Embodiment 2

[0079] In an embodiment of this application, a training step of a medical image semantic segmentation model is provided, which specifically includes:

[0080] Obtain the brain MRI images taken at different medical treatment times of the same patient, and use the latest taken image as the first image, and the remaining images as the second images. And obtain the masks corresponding to each first image and second image. In the mask of the second image, the tumor part can be shown as white and other parts as black.

[0081] Extract features from the first image and the second images through a feature extraction network to obtain a first image feature and a corresponding number of second image features.

[0082] Based on the mean value of the first image feature and multiple second image features, perform feature cross alignment to obtain multiple cross features. Based on the similarity between the first image feature and multiple second image features, optimize the multiple cross features respectively to obtain multiple optimized features.

[0083] The decoder fuses the first image feature, multiple optimized features, and the masks matched to each optimized feature, so that the decoder outputs a mask result corresponding to the first image, and the mask result is used to indicate the target pixel regions belonging to the medical semantic labels and the non-target pixel regions not belonging to the medical semantic labels in the first image.

[0084] Based on the deviation between the predicted mask result and the actual mask result of the first image, a loss function is calculated to adjust the parameters in the model, and finally a trained medical image semantic segmentation model is generated.

[0085] Embodiment 3

[0086] As Figure 3 shown, based on the same inventive concept, an embodiment of the present application further provides a medical image semantic segmentation device. The device 30 includes:

[0087] An extraction module 310, configured to extract, through a feature extraction network, a first image feature corresponding to a first image and multiple second image features corresponding to at least one second image, where the second image is an image taken during the historical medical treatment of the target patient and semantically segmented for the medical semantic labels, and the first image is an image of the target patient taken currently and having pixel regions with medical semantic labels;

[0088] An alignment module 320, configured to perform feature cross-alignment based on the means of the first image feature and the multiple second image features to obtain multiple cross features; wherein, for the first image feature and each second image feature, the corresponding cross feature is obtained by the following method:

[0089] For each first element value at each position in the first image feature, calculate the product of the first element value and the second element value at the same position in the second image feature as the third element value at the corresponding position in the corresponding cross feature;

[0090] An optimization module 330, configured to optimize the multiple cross features respectively based on the similarity between the first image feature and the multiple second image features to obtain multiple optimized features;

[0091] A decoding module 340, configured to fuse the first image feature, the multiple optimized features, and the masks matched to each optimized feature through a decoder to obtain a mask result corresponding to the first image, and the mask result is used to indicate the target pixel regions belonging to the medical semantic labels and the non-target pixel regions not belonging to the medical semantic labels in the first image.

[0092] In a preferred embodiment, the feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block connected in sequence, where each convolutional block includes a first convolutional layer, an asymmetric convolutional layer, and a second convolutional layer connected in sequence.

[0093] In a preferred embodiment, for the first image feature and each second image feature, the optimization module 330 optimizes the corresponding cross feature in the following manner to obtain the corresponding optimized feature:

[0094] Based on the first image feature and the current second image feature, determine the first element value distance feature; based on the first image feature and the remaining second image features, determine the second element value distance feature; for each third element value at each position in the current cross feature, based on the first distance value at this position in the first element value distance feature and the second distance value at this position in the second element value distance feature, determine whether the third element value at this position needs to be optimized; if so, modify the third element value based on the second element value at this position in the current second image feature to be the fourth element value at the corresponding position of the corresponding optimized feature.

[0095] In a preferred embodiment, the optimization module 330 determines whether each third element value needs to be optimized in the following manner:

[0096] Determine whether the first distance value corresponding to this position is greater than the first target value and whether the second distance value corresponding to this position is less than the second target value; if the first distance value is greater than the first target value and the second distance value is less than the second target value, determine that the third element value needs to be optimized.

[0097] In a preferred embodiment, the optimization module 330 calculates the fourth element value in the following manner:

[0098] Take the product of multiplying the third element value at the corresponding position in the current cross feature by the amplification factor as the fourth element value.

[0099] In a preferred embodiment, the optimization module 330 determines the first element value distance feature in the following manner:

[0100] Decompose the first element value at each position of the first image feature and the second element value at the corresponding position of the current second image feature, and respectively obtain the first integer element value and the first decimal element value corresponding to the first element value, and the second integer element value and the second decimal element value corresponding to the second element value; calculate the first difference between the first integer element value and the second integer element value; calculate the second difference between the first decimal element value and the second decimal element value; calculate the sum of the absolute value of the first difference and the absolute value of the second difference as the first distance value at the corresponding position in the first element value distance feature.

[0101] In a preferred embodiment, the optimization module 330 determines the second element distance feature in the following manner:

[0102] For the second element values at the corresponding positions of the remaining second image features, calculate the average value of the second elements at this position; for the first element value at each position of the first image feature and the average value of the second elements at the corresponding position, calculate the cosine similarity as the second distance value at the corresponding position in the second element distance feature.

[0103] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown in

[0104] the electronic device 400 includes a processor 410, a memory 420, and a bus 430. Figure 1 The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, it can execute the steps of a medical image semantic segmentation method in the method embodiment as described above. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here.

[0105] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, it can execute the steps of a medical image semantic segmentation method in the method embodiment as described above. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here. Figure 1 shown above

[0106] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0107] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0108] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] Furthermore, in each embodiment of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0110] It should be noted that if a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0111] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0112] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A medical image semantic segmentation method, characterized in that, The method includes: Extracting first image features corresponding to a first image and second image features corresponding to a plurality of second images respectively through a feature extraction network, where the second images are images taken during the historical medical visits of the target patient and semantically segmented for medical semantic labels, and the first image is an image of the target patient taken currently and having a pixel region with the medical semantic label; Performing feature cross-alignment based on the first image features and the plurality of second image features to obtain a plurality of cross features; For the first image feature and each second image feature, optimizing the corresponding cross feature in the following manner to obtain the corresponding optimized feature: determining a first element value distance feature based on the first image feature and the current second image feature; determining a second element value distance feature based on the first image feature and the remaining second image features; for each third element value at each position in the current cross feature, determining whether the third element value at this position needs to be optimized based on the first distance value at this position in the first element value distance feature and the second distance value at this position in the second element value distance feature; if so, modifying the third element value based on the second element value at this position in the current second image feature to be the fourth element value at this position in the corresponding optimized feature; Fusing the first image features, the plurality of optimized features, and the masks matched with each optimized feature through a decoder to obtain a mask result corresponding to the first image, where the mask result is used to indicate the target pixel regions belonging to the medical semantic label and the non-target pixel regions not belonging to the medical semantic label in the first image; Among them, for the first image feature and each second image feature, obtaining the corresponding cross feature in the following manner: Calculating the product of the first element value at each position in the first image feature and the second element value at this position in the second image feature as the third element value at this position in the corresponding cross feature.

2. The method according to claim 1, characterized in that, The feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, and a fifth convolutional block connected in sequence, where each convolutional block includes a first convolutional layer, an asymmetric convolutional layer, and a second convolutional layer connected in sequence.

3. The method according to claim 1, characterized in that, Determining whether the third element value at each position needs to be optimized in the following manner: Determining whether the first distance value corresponding to this position is greater than a first target value and whether the second distance value corresponding to this position is less than a second target value; If the first distance value is greater than the first target value and the second distance value is less than the second target value; Then it is determined that the third element value needs to be optimized.

4. The method according to claim 1, characterized in that, Calculating the fourth element value in the following manner: Taking the product of the third element value at the corresponding position in the current cross feature and the amplification factor as the fourth element value.

5. The method according to claim 1, characterized in that Determining the first element value distance feature in the following manner: Decompose the first element value at each position of the first image feature and the second element value at the corresponding position of the current second image feature, and respectively obtain the first integer element value and the first decimal element value corresponding to the first element value, and the second integer element value and the second decimal element value corresponding to the second element value; Calculate the first difference between the first integer element value and the second integer element value; Calculate the second difference between the first decimal element value and the second decimal element value; Calculate the sum between the absolute value of the first difference and the absolute value of the second difference, and use it as the first distance value at the corresponding position in the first element value distance feature.

6. The method according to claim 1, characterized in that Determine the second element distance feature by the following method: For the second element value at the corresponding position of the remaining second image features, calculate the second element mean value at this position; Calculate the cosine similarity between the first element value at each position of the first image feature and the second element mean value at the corresponding position, and use it as the second distance value at the corresponding position in the second element distance feature.

7. A medical image semantic segmentation device, characterized in that, The device includes: An extraction module, configured to respectively extract the first image feature corresponding to the first image and the second image features corresponding to multiple second images through a feature extraction network, where the second image is an image taken by the target patient during a historical medical visit and semantically segmented for medical semantic labels, and the first image is an image of the target patient taken currently and having a pixel region with the medical semantic label; An alignment module, configured to perform feature cross-alignment based on the first image feature and multiple second image features to obtain multiple cross features; where, for the first image feature and each second image feature, obtain the corresponding cross feature through the following method: For each first element value at each position in the first image feature, calculate the product of the first element value and the second element value at this position in the second image feature, and use it as the third element value at the corresponding position of the corresponding cross feature; An optimization module, configured to optimize the corresponding cross feature for the first image feature and each second image feature through the following method to obtain the corresponding optimized feature: determine the first element value distance feature based on the first image feature and the current second image feature; determine the second element value distance feature based on the first image feature and the remaining second image features; for each third element value at each position in the current cross feature, determine whether the third element value at this position needs to be optimized based on the first distance value at this position in the first element value distance feature and the second distance value at this position in the second element value distance feature; if so, modify the third element value based on the second element value at this position in the current second image feature to be used as the fourth element value at the corresponding position of the corresponding optimized feature; A decoding module, configured to fuse the first image feature, multiple optimized features, and the mask matched with each optimized feature through a decoder to obtain a mask result corresponding to the first image, and the mask result is used to indicate the target pixel region belonging to the medical semantic label and the non-target pixel region not belonging to the medical semantic label in the first image.

8. An electronic device, characterized in that, Including: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the medical image semantic segmentation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it performs the steps of the medical image semantic segmentation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image semantic segmentation method and device, equipment, storage medium and program product

    CN114612902A

  • Semantic segmentation method and system

    US20240054652A1