Medical Image Segmentation Model Training Method, Device, Medium and Equipment

By constructing a semi-supervised medical image segmentation model, combining the interactive fusion of visual and text features, and using kernel functions and multiple loss functions for training, the problem of insufficient image segmentation accuracy in traditional Chinese medicine in the existing technology is solved, and segmentation accuracy and recognition ability of complex organs are improved.

CN119810623BActive Publication Date: 2025-05-30HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510294930.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-30
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

When dealing with multi-organ segmentation of CT images, it is difficult for the prior art to accurately distinguish different organs, especially when facing complex situations or limited data volume, resulting in insufficient segmentation accuracy and blurred boundaries.

Method used

A medical image segmentation model training method is adopted to build a semi-supervised segmentation framework by obtaining labeled and unlabeled medical image samples, including multiple parallel visual coding branches and text coding branches. The kernel function and multiple loss functions (supervised loss, comparing pseudo-label supervision loss) are trained, combined with visual and text features for interactive fusion, feature maps are generated, and the loss is calculated through the slice matrix in the orthogonal direction.

Benefits of technology

The segmentation accuracy and reliability of medical image segmentation models are improved, especially when dealing with organs of small sizes and complex shapes, and the model's ability to recognize different organs and understanding of the orthogonal geometric relationships of 3D medical images is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810623B_ABST
    Figure CN119810623B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, medium and equipment for training a medical image segmentation model. The method includes: obtaining medical image samples; constructing a semi-supervised segmentation framework of an initial segmentation model; performing feature interaction and fusion processing on the medical image samples through the initial segmentation model to generate feature maps of the medical image samples; calculating a first loss of slice matrices of the feature maps of multiple visual coding branches in an orthogonal direction based on a kernel function; calculating a second loss of the feature maps corresponding to the first image samples under multiple visual coding branches based on a supervised loss function; calculating a third loss of the feature maps corresponding to the second image samples under multiple visual coding branches based on a contrastive pseudo-label supervised loss function; training the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss and the third loss, and outputting a target segmentation model after meeting the iteration termination condition. Thereby, the accuracy of multi-organ segmentation of medical images is improved, and the overall segmentation accuracy and reliability are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical image analysis, and particularly to a method, device, medium and equipment for training a medical image segmentation model. Background Art

[0002] When dealing with multi-organ segmentation of CT images, it mainly extracts visual features at the image pixel level and trains the model. As a result, when the model faces complex situations or limited data volume, it is difficult to accurately distinguish different organs, affecting the segmentation accuracy. Moreover, due to the complex layout and morphological relationships of human organs in space, relying solely on general image feature learning will ignore the constraint effect of the orthogonal geometric relationship between the cross-section and the longitudinal plane in 3D medical images on the model, resulting in segmentation errors or blurred boundaries caused by inaccurate boundary definition, especially for the segmentation of small-sized and complex-shaped organs, which cannot meet the clinical requirements for high-precision medical image segmentation. Summary of the Invention

[0003] In view of this, the present application provides a method, device, medium and equipment for training a medical image segmentation model, which solves the problem of insufficient accuracy in medical image segmentation.

[0004] According to one aspect of the present application, there is provided a method for training a medical image segmentation model, including:

[0005] Obtaining a medical image sample, where the medical image sample includes a first image sample labeled with organ categories and a second image sample not labeled with organ categories;

[0006] Constructing a semi-supervised segmentation framework of an initial segmentation model, where the semi-supervised segmentation framework includes a plurality of parallel visual coding branches and a text coding branch;

[0007] Performing feature interaction and fusion processing on the medical image sample through the initial segmentation model to generate a feature map of the medical image sample;

[0008] Calculating a first loss of the slice matrix of the feature maps of the plurality of visual coding branches in the orthogonal direction based on a kernel function;

[0009] Calculating a second loss of the feature map corresponding to the first image sample under the plurality of visual coding branches based on a supervised loss function;

[0010] Calculating a third loss of the feature map corresponding to the second image sample under the plurality of visual coding branches based on a contrastive pseudo-label supervised loss function;

[0011] Training the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and outputting the target segmentation model after meeting the iteration termination condition.

[0012] Optionally, the feature interaction and fusion processing of the medical image sample by the initial segmentation model to generate the feature map of the medical image sample includes:

[0013] Encoding the text features and visual features in the medical image sample through the visual encoding branch and the text encoding branch respectively to obtain text feature encoding and visual feature encoding with the same number of channels, where the text features are used to describe the organ category;

[0014] Performing multi-modal semantic fusion processing on the text feature encoding and the visual feature encoding to generate text-visual fusion encoding;

[0015] Connecting the text-visual fusion encoding and the visual feature encoding through skip connection to form the feature map.

[0016] Optionally, the performing multi-modal semantic fusion processing on the text feature encoding and the visual feature encoding to generate text-visual fusion encoding includes:

[0017] Projecting the text feature encoding and the visual feature encoding into the query space respectively to generate text query encoding and visual query encoding;

[0018] Projecting the text feature encoding and the visual feature encoding into the key-value pair space respectively to generate text key-value pair encoding and visual key-value pair encoding;

[0019] Performing an alignment operation on the text query encoding based on the visual key-value pair encoding to generate target text feature encoding;

[0020] Performing an alignment operation on the visual query encoding based on the text key-value pair encoding to generate target visual feature encoding;

[0021] Performing a fusion operation on the target text feature encoding and the target visual feature encoding to determine the text-visual fusion encoding.

[0022] Optionally, the encoding the text features and visual features in the medical image sample through the visual encoding branch and the text encoding branch respectively to obtain text feature encoding and visual feature encoding with the same number of channels includes:

[0023] Identifying the text features in the medical image sample;

[0024] Input the text feature into the text encoder of the text encoding branch to obtain the text feature encoding;

[0025] Input the medical image sample into the image segmentation network of the visual encoding branch to obtain the image feature encoding of the medical image sample;

[0026] Based on the channel dimension to which the text feature encoding belongs, perform a shape reshaping operation on the image feature encoding to determine the visual feature encoding.

[0027] Optionally, the image segmentation network of the visual encoding branch includes multiple encoder layers; the multi-modal semantic fusion processing of the text feature encoding and the visual feature encoding to determine the text-visual fusion encoding includes:

[0028] Perform a multi-modal semantic fusion operation on the text feature encoding and the visual feature encoding corresponding to the last encoder layer to determine the text-visual fusion encoding of the target encoder layer, where the target encoder layer is the penultimate encoder layer of the image segmentation network.

[0029] Optionally, the connecting operation of the text-visual fusion encoding and the visual feature encoding through skip connection to form the feature map includes:

[0030] Perform a connecting operation on the text-visual fusion encoding and the visual feature encoding belonging to the same encoder layer through skip connection.

[0031] Optionally, the orthogonal direction includes a first direction and a second direction that are perpendicular to each other. The calculating of the first loss of the slice matrices of the feature maps of multiple visual encoding branches in the orthogonal direction based on the kernel function includes:

[0032] Perform slice operations on the feature map along the first direction and the second direction respectively to obtain the first direction slice matrix and the second direction slice matrix;

[0033] Based on the kernel function, calculate the kernel losses of the first direction slice matrix and the second direction slice matrix under different visual encoding branches respectively;

[0034] Perform a summation operation on the kernel losses corresponding to multiple visual encoding branches and the total kernel loss corresponding to multiple visual encoding branches to obtain the first loss.

[0035] Optionally, the calculating of the kernel losses of the first direction slice matrix and the second direction slice matrix under different visual encoding branches based on the kernel function includes:

[0036] Determine the Euclidean distance matrices corresponding to the first-direction slice matrix and the second-direction slice matrix respectively under the target visual coding branch, where the target visual coding branch is any one of multiple visual coding branches;

[0037] Extract the triangular matrix from the Euclidean distance matrix;

[0038] Calculate the gamma value of the first-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the first-direction slice matrix, and calculate the gamma value of the second-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the second-direction slice matrix;

[0039] Calculate the kernel matrix based on the kernel function, the gamma value of the first-direction slice matrix, and the gamma value of the second-direction slice matrix;

[0040] Determine the kernel loss under the target visual coding branch for the kernel matrix.

[0041] According to another aspect of the present application, there is provided a medical image segmentation model training device, including:

[0042] A data acquisition module for acquiring medical image samples, where the medical image samples include first image samples labeled with organ categories and second image samples not labeled with organ categories;

[0043] A construction module for constructing a semi-supervised segmentation framework of an initial segmentation model, where the semi-supervised segmentation framework includes multiple parallel visual coding branches and a text coding branch;

[0044] A feature extraction module for performing feature interaction and fusion processing on the medical image samples through the initial segmentation model to generate feature maps of the medical image samples;

[0045] A training module for calculating the first loss of the feature maps of multiple visual coding branches in the orthogonal direction for the slice matrices based on the kernel function; and calculating the second loss of the feature maps corresponding to the first image samples under multiple visual coding branches based on the supervised loss function; and calculating the third loss of the feature maps corresponding to the second image samples under multiple visual coding branches based on the contrast pseudo-label supervised loss function; and training the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and outputting a target segmentation model after meeting the iteration termination condition.

[0046] Optionally, the feature extraction module includes:

[0047] An encoding module for encoding the text features and visual features in the medical image sample through the visual encoding branch and the text encoding branch respectively to obtain text feature encoding and visual feature encoding with the same number of channels, where the text features are used to describe the organ category;

[0048] A multi-modal fusion module for performing multi-modal semantic fusion processing on the text feature encoding and the visual feature encoding to generate text-visual fusion encoding;

[0049] The encoding module is further configured to perform a connection operation on the text-visual fusion encoding and the visual feature encoding through a skip connection to form the feature map.

[0050] Optionally, the multi-modal fusion module is specifically configured to project the text feature encoding and the visual feature encoding into a query space respectively to generate a text query encoding and a visual query encoding; project the text feature encoding and the visual feature encoding into a key-value pair space respectively to generate a text key-value pair encoding and a visual key-value pair encoding; perform an alignment operation on the text query encoding based on the visual key-value pair encoding to generate a target text feature encoding; perform an alignment operation on the visual query encoding based on the text key-value pair encoding to generate a target visual feature encoding; perform a fusion operation on the target text feature encoding and the target visual feature encoding to determine the text-visual fusion encoding.

[0051] Optionally, the encoding module is specifically configured to identify the text features in the medical image sample; input the text features into a text encoder of the text encoding branch to obtain the text feature encoding; input the medical image sample into an image segmentation network of the visual encoding branch to obtain an image feature encoding of the medical image sample; perform a shape reshaping operation on the image feature encoding based on the channel dimension to which the text feature encoding belongs to determine the visual feature encoding.

[0052] Optionally, the image segmentation network of the visual encoding branch includes multiple encoder layers; the multi-modal fusion module is specifically configured to perform a multi-modal semantic fusion operation on the text feature encoding and the visual feature encoding corresponding to the last encoder layer to determine the text-visual fusion encoding of the target encoder layer, where the target encoder layer is the penultimate encoder layer of the image segmentation network.

[0053] Optionally, the encoding module is specifically configured to perform a connection operation on the text-visual fusion encoding and the visual feature encoding belonging to the same encoder layer through a skip connection.

[0054] Optionally, the orthogonal direction includes a first direction and a second direction that are perpendicular to each other, and the medical image segmentation model training device further includes:

[0055] A segmentation module, configured to perform slicing operations on the feature map along the first direction and the second direction respectively, to obtain a first-direction slice matrix and a second-direction slice matrix;

[0056] The training module is specifically configured to calculate kernel losses of the first-direction slice matrix and the second-direction slice matrix under different visual coding branches respectively based on a kernel function; and perform a summation operation on the kernel losses corresponding to the multiple visual coding branches and the total kernel loss corresponding to the multiple visual coding branches, to obtain the first loss.

[0057] Optionally, the training module is specifically configured to determine Euclidean distance matrices corresponding to the first-direction slice matrix and the second-direction slice matrix respectively under a target visual coding branch, where the target visual coding branch is any one of the multiple visual coding branches; extract a triangular matrix from the Euclidean distance matrix; calculate a gamma value of the first-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the first-direction slice matrix, and calculate a gamma value of the second-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the second-direction slice matrix; calculate a kernel matrix based on the kernel function, the gamma value of the first-direction slice matrix, and the gamma value of the second-direction slice matrix; and determine the kernel loss of the target visual coding branch for the kernel matrix.

[0058] According to another aspect of the present application, there is provided a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the above-mentioned medical image segmentation model training method are implemented.

[0059] According to yet another aspect of the present application, there is provided a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the program, the steps of the above-mentioned medical image segmentation model training method are implemented.

[0060] With the above technical solutions, in the first aspect, the category text is encoded and then interactively fused with the image visual feature encoding, so that the features recognized by the model can combine the visual-level features and the text descriptions of organ category knowledge. Thus, the text information is fully utilized to assist the model in recognizing organs, enhancing the recognition ability of the segmentation model for different organs, improving the accuracy of multi-organ segmentation of medical images, especially improving the segmentation effect of small-sized and complex-shaped organs. In the second aspect, the orthogonal geometric characteristics and independent relationships of 3D medical images are fully considered, and the kernel function is used to quantify the dependence relationship between feature sets. Thus, pixel geometry and feature learning are closely combined, solving the problem that the prior art ignores the image geometric structure information, further enhancing the understanding and processing ability of the segmentation model for the organ spatial layout and morphological relationships, reducing segmentation errors and boundary blurring problems, and thus improving the overall segmentation accuracy and reliability. At the same time, through semi-supervised learning, the limited labeled data and a large amount of unlabeled data are fully utilized to improve the performance of the image segmentation model. In the third aspect, training is carried out through the cross pseudo-supervision of multiple visual branches, so that the training results of multiple visual branches can learn from each other, improving the generalization ability of the segmentation model. At the same time, multiple loss functions (kernel function, supervised loss function, contrast pseudo-label supervised loss function) are used to perform joint optimization of two visual branches in parallel. Under the condition of ensuring consistency constraints, the learning process of the model can be constrained from different angles, improving the performance and robustness of the model.

[0061] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific embodiments of the present application. Brief Description of the Drawings

[0062] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0063] Figure 1 A flowchart showing the training method of the medical image segmentation model provided by the embodiment of the present application;

[0064] Figure 2 A logical diagram showing the training method of the medical image segmentation model provided by the embodiment of the present application;

[0065] Figure 3 A schematic diagram showing the encoding mapping relationship provided by the embodiment of the present application;

[0066] Figure 4 A schematic diagram showing the calculation logic of the first loss provided by the embodiment of the present application;

[0067] Figure 5 The structural block diagram of the medical image segmentation model training device provided by the embodiment of the present application is shown. Detailed implementation manners

[0068] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0069] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.

[0070] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say an element is "connected" or "joined" to another element, it can be directly connected or joined to other elements, or there may also be intermediate elements. In addition, the "connection" or "joining" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0071] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. It should be understood that these embodiments are provided so that the disclosure of the present application is thorough and complete, and the concepts of these exemplary embodiments are fully conveyed to those of ordinary skill in the art.

[0072] In this embodiment, a method for training a medical image segmentation model is provided. As Figure 1 shown, the method includes:

[0073] Step 110, obtaining a medical image sample.

[0074] Among them, the medical image sample includes a first image sample labeled with organ categories and a second image sample not labeled with organ categories.

[0075] Specifically, the medical image samples can be CT (Computed Tomography) images, MRI (Magnetic Resonance) images, ultrasound images, digital X-ray images, PET (Positron Emission Tomography) images, etc., which are used to capture images of internal organs or bones of organisms, etc.

[0076] In this embodiment, it is possible to combine the use of labeled and unlabeled data for model training, thereby reducing the dependence on labeled data, which is beneficial to reducing the overall cost of data collection. At the same time, on the basis of ensuring that the model learns accurate features and patterns, it helps the model to better generalize and handle unseen data.

[0077] The medical image segmentation model training method provided by the embodiments of this application can be applied to terminals, can also be applied to servers, or can be software running on terminals or servers. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms;

[0078] Step 120, construct a semi-supervised segmentation framework for the initial segmentation model.

[0079] Among them, the semi-supervised segmentation framework includes multiple parallel visual coding branches and a text coding branch. The multiple parallel visual branches use the same image segmentation network (such as U-Net, DeepLabResNet, VGG, etc.), but are initialized with different values and the parameters are updated independently. In this way, the multiple visual coding branches can process the same input sample through different feature extraction methods, thereby reducing the risk of overfitting of a single branch and improving the generalization ability of the model on new samples. To increase the diversity and generalization ability of the model. The text coding branch usually uses a pre-trained language model (such as BERT, CLIP, GPT, etc.) to extract text features.

[0080] Step 130, perform feature interaction and fusion processing on the medical image sample through the visual coding branch and the text coding branch of the initial segmentation model to generate a feature map of the medical image sample.

[0081] Among them, the feature map is the direct output form of feature encoding, that is, the multi-dimensional tensor output by the encoder. The feature map usually contains spatial information and channel information. Through operations such as convolution, downsampling, and non-linear activation, the feature map gradually transitions from low-level features to high-level features.

[0082] In this embodiment, the feature description of the medical image sample regarding the organ category can be identified through the text encoding branch and combined with the visual features identified by the visual encoding branch. The trained segmentation model can thus make full use of the text information to assist the model in identifying organs, enhance the recognition ability of the segmentation model for different organs, improve the accuracy of multi-organ segmentation of medical images, and especially improve the segmentation effect of small-sized and complex-shaped organs. In addition, the prediction result (pseudo-label) of one visual branch can be used as the supervision signal for other visual branches, enabling multiple visual branches to learn from each other, thereby enhancing the generalization ability of the target segmentation model.

[0083] In one embodiment, step 130, that is, the feature interaction and fusion processing of the medical image sample is performed through the visual encoding branch and the text encoding branch of the initial segmentation model to generate the feature map of the medical image sample, which specifically includes the following steps:

[0084] Step 131, encode the text features and visual features in the medical image sample through the visual encoding branch and the text encoding branch respectively to obtain the text feature encoding and visual feature encoding with the same number of channels.

[0085] Among them, the text features are used to describe the organ category.

[0086] In this embodiment, the text feature encoding and visual feature encoding existing in the medical image sample are respectively extracted through the visual encoding branch and the text encoding branch, and the visual and text encodings are mapped to the same space so that the text feature encoding and the visual feature encoding are in the same dimension. This enables information of different modalities to be more conveniently fused, making the model more focused on the regional segmentation of organ category-related characteristics, enhancing the model's understanding of medical images, and thus improving the performance of the model's classification and segmentation tasks.

[0087] In a specific application scenario, step 131 specifically includes: identifying the text features in the medical image sample; inputting the text features into the text encoder of the text encoding branch to obtain the text feature encoding; inputting the medical image sample into the image segmentation network of the visual encoding branch to obtain the image feature encoding of the medical image sample; and performing a shape reshaping operation on the image feature encoding based on the channel dimension to which the text feature encoding belongs to determine the visual feature encoding.

[0088] In this embodiment, the features existing in the medical image sample are extracted through the encoders of the visual encoding branch and the text encoding branch, and then the channel dimension of the text encoding is aligned with the visual encoding through a shape reshaping operation to facilitate cross-modal feature fusion and make the visual features more focused on the target area described by the text. In this way, the trained segmentation model can more accurately distinguish the target organ from the background in a complex multi-organ scenario.

[0089] Exemplarily, as Figure 2 shown, the text prompt related to the organ category is input into the CLIP text encoder, and after the encoding operation, the text feature encoding is obtained , where K represents the number of organ categories, and C represents the number of channels of the text feature encoding. At the same time, the image sample first undergoes a convolution operation through a 3D convolutional layer Conv3d to obtain the image feature encoding , where C 1 represents the number of channels of the image feature encoding, and D, H, and W respectively represent the depth, height, and width of the image; then the reshape function is used for shape reshaping to obtain the visual feature encoding with the same channel dimension as the text feature encoding , where .

[0090] Step 132: Perform multimodal semantic fusion processing on the text feature encoding and the visual feature encoding to generate a text-visual fusion encoding.

[0091] In this embodiment, the category text (such as "liver", "spleen", etc.) is encoded and interactively fused with the visual feature encoding, so that the features recognized by the model can combine the visual-level features and the text descriptions of organ category knowledge. The model can obtain a more comprehensive feature representation, thereby improving the performance of classification, segmentation, or other tasks, enhancing the recognition ability of the segmentation model for different organs, and improving the accuracy of multi-organ segmentation of medical images.

[0092] It should be noted that since the semi-supervised segmentation framework includes multiple parallel visual coding branches, the visual feature encoding extracted by each visual coding branch is fused with the text feature encoding to ensure the consistency of the subsequent supervised training of multiple parallel visual coding branches.

[0093] In a specific application scenario, step 132 specifically includes: projecting the text feature encoding and the visual feature encoding into the query space respectively to generate a text query encoding and a visual query encoding; projecting the text feature encoding and the visual feature encoding into the key-value pair space respectively to generate a text key-value pair encoding and a visual key-value pair encoding; performing an alignment operation on the text query encoding based on the visual key-value pair encoding to generate a target text feature encoding; performing an alignment operation on the visual query encoding based on the text key-value pair encoding to generate a target visual feature encoding; and performing a fusion operation on the target text feature encoding and the target visual feature encoding to determine the text-visual fusion encoding.

[0094] In this embodiment, the text and visual features are first projected into a specific space. The cross-attention mechanism is used to encode the text features as Query and the visual features as Key / Value to generate the visually enhanced target text feature encoding. Similarly, the cross-attention mechanism is used to encode the visual features as Query and the text features as Key / Value to generate the semantically enhanced target visual feature encoding. Finally, the text-visual fusion encoding is obtained through dot product and summation. Thus, through the cross-attention operations from text to vision and from vision to text, the alignment of the two in the feature space is promoted, and the domain gap is reduced. Furthermore, the text information is fully utilized to assist the model in identifying organs, achieving the effective fusion of text and visual information, providing richer semantic information for the subsequent segmentation task, and solving the problem of insufficient utilization of semantic information in the prior art.

[0095] Exemplarily, as Figure 3 shown, in the text-visual cross-attention mechanism, the text feature encoding is projected into space Q as the query , while the visual feature encoding is respectively projected into spaces K and V as the key and value . In the vision-text cross-attention mechanism, the visual feature encoding is projected into space Q as the query , and the text feature encoding is projected into spaces K and V as the key and value . Through the residual connection method, the transpose of space Q is multiplied by space K (QK T ), and after softmax normalization (labeled as S in the figure), the aligned text feature encoding and visual feature encoding are obtained. The specific calculation formula is:

[0096] ;

[0097] ;

[0098] In the formula, C represents the number of channels of the text feature encoding. α and β are learnable parameters used to balance the influence brought by the cross-fusion part.

[0099] On the C dimension of the text feature encoding, dot product and summation operations are performed on and , and the enisum string form is used to represent the text-visual fusion encoding . The calculation formula of the text-visual fusion encoding is:

[0100] ;

[0101] In the formula, represents the dot product operation.

[0102] Step 133: Connect the text-visual fusion encoding and the visual feature encoding through skip connections to form a feature map.

[0103] In this embodiment, the low-level features of the encoder (downsampling path) are combined with the high-level features of the decoder (upsampling path) through skip connections, thereby restoring spatial details, enabling the model to better understand and represent the details and context information in the image, avoiding the loss of information in the deep network, and enhancing the segmentation and generalization capabilities of the model for complex medical images.

[0104] It should be noted that for the case where the image segmentation network of the visual encoding branch includes multiple encoder layers, the feature map is the direct representation of the feature encoding output by the last layer of the encoder, which is the final result of the encoder after multi-layer feature extraction of the input image.

[0105] In one embodiment, for the case where the image segmentation network of the visual encoding branch includes multiple encoder layers. Step 133 specifically includes: connecting the text-visual fusion encoding and the visual feature encoding belonging to the same encoder layer through skip connections.

[0106] Furthermore, for the case where the image segmentation network of the visual encoding branch includes multiple encoder layers. Step 132 specifically includes: performing multi-modal semantic fusion operation on the text feature encoding and the visual feature encoding corresponding to the last encoder layer to determine the text-visual fusion encoding of the target encoder layer.

[0107] Wherein, the target encoder layer is the penultimate encoder layer of the image segmentation network.

[0108] In this embodiment, considering that the output of the last layer of the encoder is already the final high-level feature map and does not require further fusion of low-level features. The visual features corresponding to the last encoder layer are used to replace the visual features corresponding to the penultimate encoder layer, and interactively fuse with the text feature encoding to generate the text-visual fusion encoding of the penultimate encoder layer. In this way, during the subsequent encoder-decoder skip connection process, according to the characteristics of semantic encoding at different network scales, more detailed semantic priors are obtained at specific layers of the visual branch and injected into the corresponding decoder layers according to specific rules. Thus, the model can utilize multi-scale information simultaneously at different network scales, improving the learning ability of the segmentation model for complex patterns and enhancing the segmentation effect of the model for different organs.

[0109] Exemplarily, the visual branch adopts a 5-layer U-Net encoder, and its encoder layers are respectively denoted as L 1 -L 5 , and the corresponding layer outputs (before performing downsampling or upsampling operations) are denoted as E v1 -E v5 . For the first three layers (L 1 -L 3 ), their outputs are respectively interactively fused with the text feature encoding to obtain the text-visual fusion encoding of each layer ; for the fourth layer L 4 , the output of the fifth layer is fused with the text feature encoding through an interactive fusion operation, and the obtained text-visual fusion encoding is used as . The visual feature encoding of the encoder layer is connected to the text-visual fusion encoding after fusion processing through a skip connection. The specific connection method is as follows:

[0110]

[0111] In the formula, represents the connection operation.

[0112] Step 140: Calculate the first loss of the slice matrix of the feature maps of multiple visual coding branches in the orthogonal direction based on the kernel function.

[0113] Among them, the orthogonal direction includes a first direction and a second direction that are perpendicular to each other.

[0114] Specifically, the kernel function is used to map data to a high-dimensional space. The kernel function includes but is not limited to Gaussian kernel function, Sigmoid kernel function, Laplace kernel function, exponential kernel function, chi-square kernel function.

[0115] In this embodiment, on the one hand, fully considering the orthogonal geometric characteristics and independent relationships of multiple planes of 3D medical images, closely combining pixel geometry with feature learning can ensure that the model can not only focus on the global information of features, but also capture the similarity of local features. Thereby improving the adaptability of the model to spatial transformation, solving the problem that the prior art ignores the image geometric structure information, further enhancing the understanding and processing ability of the segmentation model for the organ spatial layout and morphological relationship, reducing segmentation errors and boundary blurring problems, and thus improving the overall segmentation accuracy and reliability. On the other hand, by calculating the first loss of different branches, multi-view features can be effectively fused, which helps to enhance the model's understanding and expression ability of different view information, enabling the model to better capture complex patterns and structures in the image.

[0116] In one embodiment, step 140, that is, based on the kernel function, calculate the first loss of the sliced matrices of the feature maps of multiple visual encoding branches in the orthogonal directions, specifically including the following steps:

[0117] Step 141, perform slicing operations on the feature map along the first direction and the second direction respectively to obtain a first-direction sliced matrix and a second-direction sliced matrix.

[0118] It can be understood that the interval distance between slices can be reasonably set according to the orthogonal geometric accuracy, and the embodiments of the present application do not make specific limitations.

[0119] Step 142, based on the kernel function, calculate the kernel losses of the first-direction sliced matrix and the second-direction sliced matrix respectively under multiple visual encoding branches.

[0120] In a specific application scenario, step 142 specifically includes: determining Euclidean distance matrices corresponding to the first-direction sliced matrix and the second-direction sliced matrix respectively under the target visual encoding branch, where the target visual encoding branch is any one of the multiple visual encoding branches; extracting the triangular matrix from the Euclidean distance matrix; calculating the gamma value of the first-direction sliced matrix based on the median of the elements in the triangular matrix corresponding to the first-direction sliced matrix, and calculating the gamma value of the second-direction sliced matrix based on the median of the elements in the triangular matrix corresponding to the second-direction sliced matrix; calculating the kernel matrix based on the kernel function, the gamma value of the first-direction sliced matrix, and the gamma value of the second-direction sliced matrix; determining the kernel loss under the target visual encoding branch for the kernel matrix.

[0121] Step 143, perform a summation operation on the kernel losses corresponding to multiple visual encoding branches and the total kernel loss corresponding to multiple visual encoding branches to obtain the first loss.

[0122] In this embodiment, the slicing in the orthogonal directions provides feature representations of the feature map in different directions. For any visual encoding branch, the kernel losses of the first-direction sliced matrix and the second-direction sliced matrix are calculated respectively using the kernel function, so that the relationship between features in the first direction and the second direction can be captured more precisely, especially when dealing with high-dimensional data. Finally, by summing the kernel losses of the slices under multiple visual encoding branches, the kernel loss under a single visual encoding branch is integrated into the total loss function of the model for optimization, so as to closely combine pixel geometry with feature learning, solve the problem that the prior art ignores the image geometric structure information, and improve the segmentation accuracy of complex abdominal organs.

[0123] Exemplarily, taking the feature map as an example, first extract the horizontal central slice along the width dimension W from the feature map and reshape it into a horizontal sliced matrix 。Meanwhile, extract the vertical central slice along the height dimension H and reshape it into 。Calculate the pairwise Euclidean distances of the samples in matrices X and Y. The formula for Euclidean distance is:

[0124] ;

[0125] ;

[0126] In the formula, n represents the number of elements in the reshaped matrix. represents two n-dimensional vectors in matrix X, represents two n-dimensional vectors in matrix Y, and d represents the Euclidean distance.

[0127] Obtain the Euclidean distance matrix corresponding to the horizontal slice matrix through Euclidean distance calculation and the Euclidean distance matrix corresponding to the vertical slice matrix 。The size of the Euclidean distance matrix is , where m = B × C.

[0128] Then, taking as an example, first extract its upper triangular matrix (excluding the diagonal), and then calculate the median of the matrix elements, denoted as 。The formula for the gamma value is:

[0129]

[0130] And calculate in the same way. Obtain the gamma values and and then calculate the kernel matrices K and L. The calculation formulas are:

[0131] ;

[0132] ;

[0133] Finally, calculate the kernel losses of the feature map in the horizontal and vertical directions respectively based on the kernel matrix and integrate them.

[0134] The formula for calculating the kernel loss in the orthogonal direction can be simplified to:

[0135] ;

[0136] In the formula, , is the identity matrix, is the all-ones matrix.

[0137] For example, Figure 1 andFigure 4 As shown in the figure, taking the semi-supervised segmentation framework including two parallel visual branches (A and G) as an example. For multiple visual branches A and G, through orthogonal and independent constraint conditions, the respective HSIC (Hilbert-Schmidt Independence Criterion) losses are calculated as follows:

[0138] ;

[0139] ;

[0140] In the formula, and represent the orthogonal central visual feature slices of visual branch A, represents the kernel loss of the output feature map of visual branch A in the orthogonal direction, and represent the orthogonal central visual feature slices of visual branch G, represents the kernel loss of the output feature map of visual branch G in the orthogonal direction. At the same time, the HSIC loss between branches A and G is also calculated, and the calculation formula is:

[0141] ;

[0142] Integrate these HSIC losses into the total loss function of the model, that is:

[0143] .

[0144] Step 150, calculate the second loss of the feature map corresponding to the first image sample under multiple visual coding branches based on the supervised loss function.

[0145] In this embodiment, by calculating the supervised loss, that is, the second loss, the model can be guided to suppress irrelevant or redundant features, so that the segmentation prediction result of the model on the labeled data is as close as possible to the true mask, thereby improving the segmentation effect of the model. And by sharing the feature representation among different branches with the same input sample, the existing labeled information can be effectively utilized to improve the learning efficiency of the model.

[0146] Exemplarily, for the samples with the labeled organ type, the following supervised loss function is used to calculate the second loss:

[0147]

[0148] ;

[0149] In the formula, represents the second loss, represents the batch size of the labeled data, and respectively represent the image segmentation loss and the cross-entropy loss, represents the probability map output by the neural network.

[0150] Step 160, calculate the third loss of the feature map corresponding to the second image sample under multiple visual coding branches based on the contrastive pseudo-label supervision loss function.

[0151] In this embodiment, on the unlabeled data, pseudo-labels are generated through a visual coding branch and used as the supervision signal for other visual branches. Then, through the contrastive pseudo-label supervision loss function, the cross pseudo-supervision (CPS) loss is calculated to constrain the consistency between the prediction results of multiple branches and the pseudo-labels. Thus, the difference between the model prediction results and the pseudo-labels can be quantified, enabling the model to still learn useful features even with fewer true labels, thereby reducing the dependence on labeled data and the cost and time of data annotation.

[0152] Exemplarily, for samples with labeled organ types, the following supervision loss function is used to calculate the second loss:

[0153]

[0154] ;

[0155] In the formula, represents the batch size of the unlabeled data, represents the predicted probability map of the i-th unlabeled sample. By penalizing the prediction difference, the prediction of the model on the unlabeled data becomes more accurate and stable, thereby improving the generalization ability of the model.

[0156] Step 170, train the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and output the target segmentation model after meeting the iteration termination condition.

[0157] In this embodiment, multiple loss functions (kernel function, supervision loss function, contrastive pseudo-label supervision loss function) are used to jointly optimize two visual branches in parallel. Under the condition of ensuring consistency constraints, the learning process of the model can be constrained from different angles, improving the performance and robustness of the target segmentation model.

[0158] Exemplarily, such as Figure 2As shown in the figure, taking the semi-supervised segmentation framework that includes two parallel visual branches (A and G) as an example, first collect medical image samples to form a training data set. The training data set includes N labeled data and M unlabeled data, and satisfies . The labeled data is denoted as , where represents the input volume data, represents the corresponding ground truth mask; the unlabeled data is denoted as . During the training process, each training step samples from the labeled data and the unlabeled data respectively, and sends them into networks A and G for processing. During training, the text feature encoding in the medical image samples is extracted through the text encoder, and the visual feature encoding in the medical image samples is extracted through the image encoder. The text feature encoding is interactively fused with the visual feature encodings obtained from the two visual branches, and the fused feature encoding and the visual feature encoding are input into the image decoder together. Orthogonal direction center visual feature slices are performed based on the visual feature encoding finally output by the image encoder, and the HSIC loss of the branch is calculated at the same time. For the labeled data, the SUP loss is calculated using the supervised loss function based on the visual feature encoding finally output by the image encoder ( Figure 2 the feature map segmented by different colors in). For the unlabeled data, the CPS loss is calculated based on the visual feature encoding finally output by the image encoder ( Figure 2 the feature map segmented by different colors in).

[0159] Apply the HSIC loss to all sample data, apply the CPS loss to the unlabeled data, and apply the SUP loss to the labeled data. The total loss function is: .

[0160] In practical applications, empirically set to 0.1, to 0.01, and use the CPS weight growth function to gradually increase the proportion of and in the total loss, and use the loss function to guide the model training, so as to balance the influence of different loss terms on the model training, and enable the model to make full use of the labeled data and the unlabeled data for learning and optimization.

[0161] In one embodiment, after obtaining the trained target segmentation model, the medical image to be processed can be input into the target segmentation model, and the multi-organ segmentation result in the medical image can be obtained through the inference of the target segmentation model.

[0162] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0163] Further, as Figure 5 shown, as a specific implementation of the above-mentioned medical image segmentation model training method, an embodiment of the present application provides a medical image segmentation model training apparatus 500, and the medical image segmentation model training apparatus 500 includes: a data acquisition module 501, a construction module 502, a feature extraction module 503, and a training module 504.

[0164] Among them, the data acquisition module 501 is used to acquire medical image samples, where the medical image samples include first image samples labeled with organ categories and second image samples not labeled with organ categories;

[0165] The construction module 502 is used to construct a semi-supervised segmentation framework of the initial segmentation model, where the semi-supervised segmentation framework includes a plurality of parallel visual coding branches and text coding branches;

[0166] The feature extraction module 503 is used to perform feature interaction and fusion processing on the medical image samples through the visual coding branch and the text coding branch of the initial segmentation model to generate a feature map of the medical image samples;

[0167] The training module 504 is used to calculate a first loss of the sliced matrix of the feature maps of multiple visual coding branches in the orthogonal direction based on the kernel function; and calculate a second loss of the feature maps corresponding to the first image samples under multiple visual coding branches based on the supervised loss function; and calculate a third loss of the feature maps corresponding to the second image samples under multiple visual coding branches based on the contrastive pseudo-label supervised loss function; and train the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and output the target segmentation model after meeting the iteration termination condition.

[0168] Optionally, the feature extraction module 503 includes:

[0169] An encoding module (not shown in the figure) is used to encode the text features and visual features in the medical image samples through the visual coding branch and the text coding branch respectively to obtain text feature encodings and visual feature encodings with the same number of channels, where the text features are used to describe the organ categories;

[0170] A multi-modal fusion module (not shown in the figure) is used to perform multi-modal semantic fusion processing on the text feature encodings and the visual feature encodings to generate text-visual fusion encodings;

[0171] The encoding module is further used to perform a connection operation on the text-visual fusion encoding and the visual feature encoding through skip connection to form a feature map.

[0172] Optionally, the multimodal fusion module is specifically configured to project the text feature encoding and the visual feature encoding into the query space respectively to generate a text query encoding and a visual query encoding; project the text feature encoding and the visual feature encoding into the key-value pair space respectively to generate a text key-value pair encoding and a visual key-value pair encoding; perform an alignment operation on the text query encoding based on the visual key-value pair encoding to generate a target text feature encoding; perform an alignment operation on the visual query encoding based on the text key-value pair encoding to generate a target visual feature encoding; and perform a fusion operation on the target text feature encoding and the target visual feature encoding to determine a text-visual fusion encoding.

[0173] Optionally, the encoding module is specifically configured to identify the text features in the medical image sample; input the text features into the text encoder of the text encoding branch to obtain a text feature encoding; input the medical image sample into the image segmentation network of the visual encoding branch to obtain an image feature encoding of the medical image sample; and perform a shape reshaping operation on the image feature encoding based on the channel dimension to which the text feature encoding belongs to determine a visual feature encoding.

[0174] Optionally, the image segmentation network of the visual encoding branch includes multiple encoder layers; the multimodal fusion module is specifically configured to perform a multimodal semantic fusion operation on the text feature encoding and the visual feature encoding corresponding to the last encoder layer to determine the text-visual fusion encoding of the target encoder layer, where the target encoder layer is the penultimate encoder layer of the image segmentation network.

[0175] Optionally, the encoding module is specifically configured to perform a connection operation on the text-visual fusion encoding and the visual feature encoding belonging to the same encoder layer through a skip connection.

[0176] Optionally, the orthogonal directions include a first direction and a second direction that are perpendicular to each other, and the medical image segmentation model training device 500 further includes:

[0177] A segmentation module (not shown in the figure) for slicing the feature map along the first direction and the second direction respectively to obtain a first direction slice matrix and a second direction slice matrix;

[0178] The training module 504 is specifically configured to calculate the kernel losses of the first direction slice matrix and the second direction slice matrix under different visual encoding branches respectively based on a kernel function; and perform a summation operation on the kernel losses corresponding to multiple visual encoding branches and the total kernel loss corresponding to multiple visual encoding branches to obtain a first loss.

[0179] Optionally, the training module 504 is specifically configured to determine Euclidean distance matrices corresponding to the first-direction slice matrix and the second-direction slice matrix respectively under the target visual encoding branch, where the target visual encoding branch is any one of multiple visual encoding branches; extract the triangular matrix from the Euclidean distance matrix; calculate the gamma value of the first-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the first-direction slice matrix, and calculate the gamma value of the second-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the second-direction slice matrix; calculate the kernel matrix based on the kernel function, the gamma value of the first-direction slice matrix, and the gamma value of the second-direction slice matrix; and determine the kernel loss under the target visual encoding branch for the kernel matrix.

[0180] For the specific limitations of the medical image segmentation model training device, reference can be made to the limitations on the medical image segmentation model training method in the foregoing text, which will not be elaborated herein. Each module in the above medical image segmentation model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0181] Based on the above as Figure 1 shown in the method, correspondingly, an embodiment of the present application further provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the medical image segmentation model training method as Figure 1 shown above.

[0182] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of the present application.

[0183] Based on the above as Figure 1 shown in the method, and Figure 5 shown in the virtual device embodiment, for the purpose of achieving the above object, an embodiment of the present application further provides a computer device, which can specifically be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the medical image segmentation model training method as Figure 1 shown above.

[0184] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen, an input unit such as a keyboard, etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0185] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not constitute a limitation on the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0186] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing and storing the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components within the storage medium, as well as communication between other hardware and software in the entity device.

[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or the acquisition of medical image samples can be achieved through hardware. Among them, the medical image samples include a first image sample labeled with organ categories and a second image sample without labeled organ categories; constructing a semi-supervised segmentation framework for an initial segmentation model, where the semi-supervised segmentation framework includes multiple parallel visual encoding branches and a text encoding branch; performing feature interaction and fusion processing on the medical image samples through the initial segmentation model to generate feature maps of the medical image samples; calculating a first loss of the slice matrix of the feature maps of multiple visual encoding branches in the orthogonal direction based on a kernel function; calculating a second loss of the feature maps corresponding to the first image sample under multiple visual encoding branches based on a supervised loss function; calculating a third loss of the feature maps corresponding to the second image sample under multiple visual encoding branches based on a contrastive pseudo-label supervised loss function; training the initial segmentation model with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and outputting a target segmentation model after meeting the iteration termination condition. In the embodiments of this application, on the one hand, after encoding the category text, it is interactively fused with the image visual feature encoding, so that the features recognized by the model can combine the visual-level features and the text description of organ category knowledge. Thus, the text information is fully utilized to assist the model in identifying organs, enhancing the recognition ability of the segmentation model for different organs, improving the accuracy of multi-organ segmentation of medical images, especially improving the segmentation effect of small-sized and complex-shaped organs. On the other hand, the orthogonal geometric characteristics and independent relationships of multiple planes of 3D medical images are fully considered, and the kernel function is used to quantify the dependence relationship between feature sets. Thus, pixel geometry is closely combined with feature learning, solving the problem that the prior art ignores the image geometric structure information, further enhancing the segmentation model's understanding and processing ability of the organ spatial layout and morphological relationship, reducing segmentation errors and boundary blurring problems, and thus improving the overall segmentation accuracy and reliability. At the same time, through the semi-supervised learning method, the limited labeled data and a large amount of unlabeled data are fully utilized to improve the performance of the image segmentation model. On the third hand, training is carried out through the cross-pseudo-supervision method of multiple visual branches, so that the training results of multiple visual branches can learn from each other, improving the generalization ability of the segmentation model. At the same time, multiple loss functions (kernel function, supervised loss function, contrastive pseudo-label supervised loss function) are used to perform joint optimization of two visual branches in parallel. Under the condition of ensuring consistency constraints, the learning process of the model can be constrained from different angles, improving the performance and robustness of the model.

[0188] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the attached drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the devices in the implementation scenario can be distributed in the devices in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0189] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above-disclosed are only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A medical image segmentation model training method, characterized in that: The method comprises: Acquire a medical image sample, wherein the medical image sample includes a first image sample labeled with an organ category and a second image sample not labeled with an organ category; Constructing a semi-supervised segmentation framework of the initial segmentation model, wherein the semi-supervised segmentation framework includes a plurality of parallel visual encoding branches and a text encoding branch; Performing feature interactive fusion processing on the medical image sample through the initial segmentation model to generate a feature map of the medical image sample; Based on the kernel function, calculating the first loss of the slice matrix of the feature graphs of the multiple visual encoding branches in the orthogonal direction, wherein the orthogonal direction includes a first direction and a second direction perpendicular to each other; Calculate the second loss of the feature map corresponding to the first image sample under multiple visual encoding branches based on the supervised loss function; Calculate a third loss of the feature map corresponding to the second image sample under multiple visual encoding branches based on the comparison pseudo-label supervision loss function; The initial segmentation model is trained with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and a target segmentation model is output after an iteration termination condition is met; The step of calculating the first loss of the feature maps of the plurality of visual encoding branches in a slice matrix in an orthogonal direction based on the kernel function comprises: Slicing the feature map along the first direction and the second direction respectively to obtain a first direction slice matrix and a second direction slice matrix; Based on the kernel function, respectively calculating the kernel loss of the first direction slice matrix and the second direction slice matrix under multiple visual encoding branches; The kernel losses corresponding to the multiple visual encoding branches and the total kernel losses corresponding to the multiple visual encoding branches are summed to obtain the first loss.

2. The medical image segmentation model training method according to claim 1, characterized in that: The step of performing feature interactive fusion processing on the medical image sample by using the initial segmentation model to generate a feature map of the medical image sample includes: The text features and the visual features in the medical image sample are respectively encoded by the visual encoding branch and the text encoding branch to obtain text feature encoding and visual feature encoding with the same number of channels, wherein the text features are used to describe the organ category; Performing multimodal semantic fusion processing on the text feature code and the visual feature code to generate a text-visual fusion code; The text-visual fusion code and the visual feature code are connected via a jump connection to form the feature map.

3. The medical image segmentation model training method according to claim 2, characterized in that: The performing multimodal semantic fusion processing on the text feature code and the visual feature code to generate a text-visual fusion code includes: Respectively projecting the text feature code and the visual feature code into a query space to generate a text query code and a visual query code; Respectively projecting the text feature code and the visual feature code into a key-value pair space to generate a text key-value pair code and a visual key-value pair code; Performing an alignment operation on the text query code based on the visual key-value pair code to generate a target text feature code; Performing an alignment operation on the visual query encoding based on the text key-value pair encoding to generate a target visual feature encoding; A fusion operation is performed on the target text feature code and the target visual feature code to determine a text-visual fusion code.

4. The medical image segmentation model training method according to claim 2, characterized in that: The encoding process of the text features and the visual features in the medical image sample by the visual encoding branch and the text encoding branch to obtain the text feature encoding and the visual feature encoding with the same number of channels includes: identifying text features in the medical image sample; Inputting the text feature into the text encoder of the text encoding branch to obtain the text feature encoding; Inputting the medical image sample into the image segmentation network of the visual encoding branch to obtain the image feature encoding of the medical image sample; The image feature code is reshaped based on the channel dimension to which the text feature code belongs to determine the visual feature code.

5. The medical image segmentation model training method according to claim 3, characterized in that: The image segmentation network of the visual coding branch includes a plurality of encoder layers; the multimodal semantic fusion processing of the text feature coding and the visual feature coding to determine the text-visual fusion coding includes: Performing a multimodal semantic fusion operation on the text feature code and the visual feature code corresponding to the last encoder layer to determine a text-visual fusion code of a target encoder layer, wherein the target encoder layer is the penultimate encoder layer of the image segmentation network; The step of connecting the text-visual fusion code and the visual feature code through a skip connection to form the feature map includes: The text-visual fusion encoding and the visual feature encoding belonging to the same encoder layer are connected through a jump connection.

6. The medical image segmentation model training method according to claim 1, characterized in that: The kernel function-based, respectively calculating the kernel loss of the first direction slice matrix and the second direction slice matrix under different visual encoding branches, comprises: Determine a Euclidean distance matrix corresponding to the first direction slice matrix and the second direction slice matrix respectively under a target visual coding branch, wherein the target visual coding branch is any visual coding branch among the multiple visual coding branches; Extracting a triangular matrix from the Euclidean distance matrix; Calculating the gamma value of the first-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the first-direction slice matrix, and calculating the gamma value of the second-direction slice matrix based on the median of the elements in the triangular matrix corresponding to the second-direction slice matrix; Calculate a kernel matrix based on a kernel function, a gamma value of the first-direction slice matrix, and a gamma value of the second-direction slice matrix; A kernel loss under the target visual encoding branch is determined based on the kernel matrix.

7. A medical image segmentation model training device, characterized in that: The device comprises: A data acquisition module, used to acquire medical image samples, wherein the medical image samples include first image samples labeled with organ categories and second image samples not labeled with organ categories; A construction module, used to construct a semi-supervised segmentation framework of an initial segmentation model, wherein the semi-supervised segmentation framework includes a plurality of parallel visual encoding branches and a text encoding branch; A feature extraction module, used for performing feature interactive fusion processing on the medical image sample through the initial segmentation model to generate a feature map of the medical image sample; A training module, for calculating the first loss of the slicing matrix of the feature graphs of the multiple visual coding branches in the orthogonal direction based on the kernel function, wherein the orthogonal directions include a first direction and a second direction perpendicular to each other, and calculating the first loss of the slicing matrix of the feature graphs of the multiple visual coding branches in the orthogonal direction based on the kernel function, comprising: slicing the feature graph along the first direction and the second direction respectively to obtain the first direction slicing matrix and the second direction slicing matrix; calculating the kernel loss of the first direction slicing matrix and the second direction slicing matrix under the multiple visual coding branches respectively based on the kernel function; summing the kernel losses corresponding to the multiple visual coding branches and the total kernel losses corresponding to the multiple visual coding branches to obtain the first loss; and, Calculating a second loss of the feature map corresponding to the first image sample under the plurality of visual encoding branches based on a supervised loss function; and Calculating a third loss of the feature map corresponding to the second image sample under the plurality of visual encoding branches based on the comparison pseudo-label supervision loss function; and The initial segmentation model is trained with the goal of minimizing the weighted sum of the first loss, the second loss, and the third loss, and a target segmentation model is output after an iteration termination condition is met.

8. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the medical image segmentation model training method as described in any one of claims 1 to 6 are implemented.

9. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the program, the medical image segmentation model training method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Creference image segmentation model training method and reference image segmentation method

    CN116993976A

  • Semantic segmentation method oriented to omnibearing supervision based on prompt text

    CN117830638A

  • Medical image segmentation method fusing multi-modal graphic and text information

    CN119205800A