Medical Image Segmentation Method and System Based on Natural Language Processing Technology

By combining natural language processing and deep learning neural network methods, the shortcomings in complexity and diversity of traditional medical image segmentation methods are solved, and more efficient and accurate medical image segmentation is achieved.

CN118429369BActive Publication Date: 2025-07-18ZHEJIANG FEITU IMAGING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410762178.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-07-18
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

Traditional medical image segmentation methods are not good in dealing with images with smooth gray value changes and noise interference or the edges are not obvious, and lack a deep understanding of the image content, so they cannot effectively process the complexity and diversity of medical images.

Method used

Using a method based on natural language processing technology, by acquiring medical images and natural language descriptions, using deep learning neural networks for image and text processing, combining focus area maps and cross-modal learning modules, feature extraction and segmentation of medical images is performed to achieve accurate segmentation of anatomical structures and lesions.

Benefits of technology

It improves the accuracy and efficiency of medical image segmentation, can better understand the anatomical structure and lesions in the image, handles complex textures and fuzzy boundaries, and makes the segmentation results meet the expectations of medical experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118429369B_ABST
    Figure CN118429369B_ABST
Patent Text Reader

Abstract

The present application discloses a medical image segmentation method and system based on natural language processing technology. By obtaining a medical image and a natural language description of the medical image, and using image and text processing and analysis algorithms based on a deep learning neural network to analyze the medical image and process the natural language description of the medical image, the medical image segmentation result can be intelligently obtained according to the image pixel feature information guided by the natural language of the medical image. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex texture and fuzzy boundaries of the medical image can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent segmentation, and more specifically, to a medical image segmentation method and system based on natural language processing technology. Background Art

[0002] Medical image segmentation is a key computer vision technology that plays a crucial role in medical diagnosis and treatment. This technology involves accurately separating specific structures in medical images, such as organs, tumors, or other anatomical features, from their surrounding environment. Through medical image segmentation technology, doctors can better understand and analyze the shape, size, and potential medical problems of these structures.

[0003] However, traditional medical image segmentation methods are based on thresholding and edge detection. Specifically, threshold segmentation methods often rely on the gray-level differences in images. For images with smooth gray-level changes, such as tumor segmentation, it may be difficult to find an appropriate threshold to accurately capture and characterize these complex features, resulting in inaccurate segmentation results. Edge detection methods rely on edge information in images, but in cases with significant noise interference or unclear edges, the segmentation effect is not good. In addition, these traditional methods usually lack a deep understanding of the image content and cannot effectively handle the complexity and diversity in medical images.

[0004] Therefore, a medical image segmentation method based on natural language processing technology is desired. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a medical image segmentation method and system based on natural language processing technology. By obtaining a medical image and a natural language description of the medical image, and using an image and text processing and analysis algorithm based on a deep learning neural network to analyze the medical image and process the natural language description of the medical image, an intelligent medical image segmentation result is obtained based on the image pixel feature information guided by the natural language of the medical image. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex textures and fuzzy boundaries of medical images can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.

[0006] According to one aspect of the present application, a medical image segmentation method based on natural language processing technology is provided, which includes: obtaining a medical image and a natural language description of the medical image; performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image; extracting medical image features from the medical image to obtain a medical image feature map; passing the medical image feature map through a main feature distribution mapping network based on a focus area map to obtain a foreground-salient medical image feature map; performing per-pixel position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; using a cross-modal learning module to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors; aggregating the features of the set of semantic-guided reinforced medical image pixel-level feature vectors to obtain a reinforced medical image feature map as the reinforced medical image feature; and obtaining a medical image segmentation result based on the reinforced medical image feature.

[0007] In the above medical image segmentation method based on natural language processing technology, performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image includes: performing word segmentation on the natural language description of the medical image and then passing it through a semantic encoder including a word embedding layer to obtain the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image.

[0008] In the above medical image segmentation method based on natural language processing technology, extracting medical image features from the medical image to obtain a medical image feature map includes: passing the medical image through a medical image feature extractor based on AlexNet to obtain the medical image feature map.

[0009] In the above-mentioned medical image segmentation method based on natural language processing technology, obtaining a foreground-salient medical image feature map by passing the medical image feature map through a main feature distribution mapping network based on a focus region map includes: calculating spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain a medical image spatial attention feature matrix; performing masking processing on the medical image spatial attention feature matrix based on a predetermined threshold to obtain a medical image masked spatial attention feature matrix; and calculating the element-wise multiplication between each feature matrix along the channel dimension of the medical image feature map and the medical image masked spatial attention feature matrix to obtain the foreground-salient medical image feature map; wherein, calculating spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain a medical image spatial attention feature matrix includes: performing global average pooling processing on the medical image feature map along the channel dimension to obtain a medical image global average pooling feature matrix; inputting the medical image global average pooling feature matrix into a Sigmoid activation unit to obtain the medical image spatial attention feature matrix; wherein, performing masking processing on the medical image spatial attention feature matrix based on a predetermined threshold to obtain a medical image masked spatial attention feature matrix includes: in response to each position feature value in the medical image spatial attention feature matrix being greater than or equal to the predetermined threshold, taking the original value of the position feature value, otherwise setting it to zero.

[0010] In the above-mentioned medical image segmentation method based on natural language processing technology, using a cross-modal learning module to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors includes: calculating the similarity between each foreground-salient medical image pixel-level feature vector in the set of foreground-salient medical image pixel-level feature vectors and all medical image natural semantic descriptor granularity semantic encoding feature vectors in the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors to obtain a set of medical image pixel-semantic similarity row vectors; calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors; and respectively using each medical image pixel-semantic similarity class support weight vector in the set of medical image pixel-semantic similarity class support weight vectors as a weight vector to calculate the weighted sum of the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors to obtain the set of semantic-guided reinforced medical image pixel-level feature vectors.

[0011] In the above medical image segmentation method based on natural language processing technology, calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors includes: taking each eigenvalue in the medical image pixel-semantic similarity row vector as the exponent of the natural constant to calculate the exponential function value with the natural constant as the base according to the position to obtain a medical image pixel-semantic similarity class support row vector; calculating the sum of the respective position eigenvalues in the medical image pixel-semantic similarity class support row vector to obtain a medical image pixel-semantic global similarity; and calculating the position-wise division between each eigenvalue in the medical image pixel-semantic similarity class support row vector and the medical image pixel-semantic global similarity to obtain the semantic-guided enhanced medical image pixel-level feature vector.

[0012] In the above medical image segmentation method based on natural language processing technology, obtaining a medical image segmentation result based on the enhanced medical image features includes: passing the enhanced medical image feature map through a semantic segmenter based on the Softmax classification function to obtain a medical image segmentation result.

[0013] In the above medical image segmentation method based on natural language processing technology, it further includes training the semantic encoder including a word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus area map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function.

[0014] In the above medical image segmentation method based on natural language processing technology, the training step includes: obtaining training data, where the training data includes training medical images and natural language descriptions of the training medical images, as well as the ground truth of the medical image segmentation results; performing word segmentation on the natural language descriptions of the training medical images and then passing them through the semantic encoder including the word embedding layer to obtain a sequence of semantic encoding feature vectors of the training medical images at the word granularity of natural semantic descriptions; passing the training medical images through the medical image feature extractor based on AlexNet to obtain training medical image feature maps; passing the training medical image feature maps through the main feature distribution mapping network based on the focus region map to obtain training foreground saliency medical image feature maps; performing per-pixel position feature dispersion on the training foreground saliency medical image feature maps to obtain a set of training foreground saliency medical image pixel-level feature vectors; using the cross-modal learning module to perform cross-modal learning reinforcement on the set of training foreground saliency medical image pixel-level feature vectors and the sequence of semantic encoding feature vectors of the training medical images at the word granularity of natural semantic descriptions to obtain a set of training semantic-guided reinforcement medical image pixel-level feature vectors; aggregating the features of the set of training semantic-guided reinforcement medical image pixel-level feature vectors to obtain training reinforced medical image feature maps; optimizing the features of the training reinforced medical image feature maps to obtain optimized training reinforced medical image feature maps; passing the optimized training reinforced medical image feature maps through the semantic segmenter based on the Softmax classification function to obtain the value of the classification loss function; and, based on the value of the classification loss function and through backpropagation of gradient descent, optimizing the semantic encoder including the word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus region map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function, where, in each iteration of the model, the training reinforced medical image feature maps are optimized.

[0015] According to another aspect of the present application, there is provided a medical image segmentation system based on natural language processing technology, which includes: a medical image data acquisition module for acquiring a medical image and a natural language description of the medical image; a word-grained semantic encoding module for performing word segmentation on the natural language description of the medical image and then performing word-grained semantic encoding to obtain a sequence of word-grained semantic encoding feature vectors of the natural semantic description of the medical image; a medical image feature extraction module for extracting medical image features from the medical image to obtain a medical image feature map; a feature distribution mapping module for passing the medical image feature map through a main feature distribution mapping network based on a focus area map to obtain a foreground-salient medical image feature map; a feature dispersion module for performing pixel-by-pixel position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; a semantic guidance enhancement module for using a cross-modal learning module to perform cross-modal learning enhancement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of word-grained semantic encoding feature vectors of the natural semantic description of the medical image to obtain a set of semantic guidance enhanced medical image pixel-level feature vectors; a feature aggregation module for aggregating the set of semantic guidance enhanced medical image pixel-level feature vectors to obtain an enhanced medical image feature map as the enhanced medical image feature; and a segmentation result generation module for obtaining a medical image segmentation result based on the enhanced medical image feature.

[0016] Compared with the prior art, a medical image segmentation method and system provided by the present application obtain a medical image and a natural language description of the medical image, and use image and text processing and analysis algorithms based on deep learning neural networks to analyze the medical image and process the natural language description of the medical image, so as to intelligently obtain a medical image segmentation result according to the image pixel feature information guided by the natural language of the medical image. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex texture and fuzzy boundary of the medical image can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation. Brief Description of the Drawings

[0017] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1Flowchart of a medical image segmentation method based on natural language processing technology according to an embodiment of the present application.

[0019] Figure 2 Schematic diagram of the architecture of a medical image segmentation method based on natural language processing technology according to an embodiment of the present application.

[0020] Figure 3 Flowchart of obtaining a foreground-salient medical image feature map by passing the medical image feature map through a main feature distribution mapping network based on a focus region map in a medical image segmentation method based on natural language processing technology according to an embodiment of the present application.

[0021] Figure 4 Flowchart of training the semantic encoder including a word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on a focus region map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function in a medical image segmentation method based on natural language processing technology according to an embodiment of the present application.

[0022] Figure 5 Block diagram of a medical image segmentation system based on natural language processing technology according to an embodiment of the present application. Detailed implementation manners

[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0024] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0025] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0026] It should be noted that the modifiers "a" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0027] Medical image segmentation is a key computer vision technology that plays a crucial role in medical diagnosis and treatment. This technology involves accurately separating specific structures in medical images, such as organs, tumors, or other anatomical features, from their surrounding environment. Through medical image segmentation technology, doctors can better understand and analyze the shape, size, and potential medical problems of these structures.

[0028] However, traditional medical image segmentation methods adopt threshold - based and edge - detection - based approaches. Specifically, the threshold segmentation method often relies on the gray - level differences in the image. For images with smooth gray - level changes, such as tumor segmentation, it may be difficult to find an appropriate threshold to accurately capture and characterize these complex features, resulting in inaccurate segmentation results. The edge - detection method relies on the edge information in the image, but in cases where there is significant noise interference or the edges are not obvious, the segmentation effect is poor. In addition, these traditional methods usually lack a deep understanding of the image content and cannot effectively handle the complexity and diversity in medical images.

[0029] Therefore, to address the above - mentioned technical problems, the technical concept of this application is to obtain a medical image and the natural - language description of the medical image, and adopt image and text processing and analysis algorithms based on deep - learning neural networks to analyze the medical image and process the natural - language description of the medical image, so as to intelligently obtain the medical image segmentation result based on the image pixel feature information guided by the natural language of the medical image. Through this method, the anatomical structure and pathological conditions in the image can be better understood, and at the same time, the complex texture and blurred boundaries of the medical image can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.

[0030] Figure 1 It is a flowchart of a medical image segmentation method based on natural - language processing technology according to an embodiment of this application. Figure 2 It is a schematic diagram of the architecture of a medical image segmentation method based on natural - language processing technology according to an embodiment of this application. As Figure 1 and Figure 2As shown, the medical image segmentation method based on natural language processing technology according to an embodiment of the present application includes: S110, obtaining a medical image and a natural language description of the medical image; S120, performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of medical image natural semantic description word granularity semantic encoding feature vectors; S130, extracting medical image features from the medical image to obtain a medical image feature map; S140, passing the medical image feature map through a main feature distribution mapping network based on a focus area map to obtain a foreground-salient medical image feature map; S150, performing pixel-wise position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; S160, using a cross-modal learning module to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic description word granularity semantic encoding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors; S170, aggregating the features of the set of semantic-guided reinforced medical image pixel-level feature vectors to obtain a reinforced medical image feature map as the reinforced medical image feature; and S180, obtaining a medical image segmentation result based on the reinforced medical image feature.

[0031] In step S110, a medical image and a natural language description of the medical image are obtained. It should be understood that the medical image refers to an image of the internal structure or tissue of the human body obtained by a medical imaging device (such as X-ray, CT scan, MRI, etc.), which provides rich anatomical structure and lesion information. The natural language description of the medical image refers to a textual description of the medical image, usually written by a doctor or a professional, for recording information such as abnormalities, lesions, structures, etc. found in the image. In particular, these descriptions may include the location, size, shape, characteristics, and relationship with surrounding tissues of the lesion, etc., which can further supplement and explain this information and provide a more comprehensive basis for medical diagnosis. Based on this, in order to be able to perform more accurate segmentation of medical images in the technical solution of the present application, a medical image and a natural language description of the medical image are obtained.

[0032] In step S120, after performing word segmentation on the natural language description of the medical image, semantic encoding based on word granularity is performed to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image. Specifically, in the embodiment of the present application, after performing word segmentation on the natural language description of the medical image, semantic encoding based on word granularity is performed to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image, including: after performing word segmentation on the natural language description of the medical image, passing it through a semantic encoder including a word embedding layer to obtain the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image. Correspondingly, considering that the natural language description of the medical image is a complete natural language description composed of multiple words or terms about the medical image, and there is a semantic association in context between the multiple words or terms about the medical image. Therefore, in order to better understand and analyze the text meanings of these words or terms about medical image information, in the technical solution of the present application, word segmentation is performed on the natural language description of the medical image. It should be understood that word segmentation can divide the natural language description of the long-text medical image into smaller semantic units (words or phrases), thereby helping the computer better understand the text meanings and structures of each medical image word, providing support and assistance for subsequent semantic encoding and text understanding. Further, in order to be able to more clearly understand the semantic information between each word about the medical image and thus better represent the natural language of the medical image, in the technical solution of the present application, the natural language description of the medical image after word segmentation is passed through a semantic encoder including a word embedding layer to capture and mine the context information about the medical image between each word, thereby obtaining a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image containing rich semantic information, providing a more accurate and rich semantic feature representation for subsequent medical image analysis and processing.

[0033] In step S130, medical image feature extraction is performed on the medical image to obtain a medical image feature map. Specifically, in the embodiment of the present application, performing medical image feature extraction on the medical image to obtain a medical image feature map includes: passing the medical image through a medical image feature extractor based on AlexNet to obtain the medical image feature map. It should be understood that considering that the medical image contains key feature information, such as the anatomical structure inside the human body, the size and morphological information of organs or lesions, etc., these key feature information play an important role in the medical image segmentation task. And AlexNet performs excellently in the image recognition task and can effectively extract features in the image, including information such as edges, textures, and shapes, which is suitable for the feature extraction of medical images. Based on this, in order to more accurately understand and capture the key feature information in the medical image, in the technical solution of the present application, the medical image is passed through a medical image feature extractor based on AlexNet to obtain a medical image feature map. It is worth mentioning that compared with other CNN models, the AlexNet contains a total of 5 convolutional layers and 3 fully connected layers, and its depth is several times that of the previous neural network, and the ReLU (Rectified Linear Unit) is used as the activation function, which can effectively solve the problem of gradient disappearance. That is to say, by processing the medical image through AlexNet, it is possible to focus on the higher-level and more abstract feature information of the key structures in the medical image, so as to extract and mine the complex structures and patterns in the medical image, which helps to improve the performance of medical image segmentation.

[0034] In step S140, the medical image feature map is passed through a main feature distribution mapping network based on a focus region map to obtain a foreground-salient medical image feature map. Accordingly, considering that the medical image feature map contains key foreground feature information regarding a medical image segmentation task, such as a lesion area, e.g., the location, size, and morphology of a tumor, or functional information of an organ or tissue, etc., but there is also background and redundant information that has little to do with the task. Therefore, in order to focus on the focus feature information of the medical image feature map and achieve foreground saliency of the medical image feature map, that is, to highlight the main object or region of interest in the medical image feature map while reducing the influence and interference of background information, in the technical solution of the present application, the medical image feature map is passed through a main feature distribution mapping network based on a focus region map to obtain a foreground-salient medical image feature map. It is worth mentioning that the main feature distribution mapping network based on the focus region map incorporates a gating mechanism, which can accurately identify and process foreground objects of interest in the image, dynamically adjust the focus on the key region part of the input data by learning the importance of different regions in the image, so as to more accurately identify and analyze the image content. That is to say, through the main feature distribution mapping network based on the focus region map, the important key regions in the medical image feature map can be focused on, the attention to the main foreground object can be improved, so as to distinguish the main foreground object from the background and reduce the interference of redundant background information to obtain a foreground-salient medical image feature map.

[0035] Figure 3 A flowchart of obtaining a foreground-salient medical image feature map by passing the medical image feature map through a main feature distribution mapping network based on a focus region map in a medical image segmentation method based on natural language processing technology according to an embodiment of the present application. Specifically, in the embodiment of the present application, as Figure 3As shown, the medical image feature map is passed through the main feature distribution mapping network based on the focus area map to obtain the foreground salient medical image feature map, including: S210, calculating the spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain the medical image spatial attention feature matrix; S220, performing masking processing on the medical image spatial attention feature matrix based on a predetermined threshold to obtain the medical image masked spatial attention feature matrix; and, S230, calculating the element-wise multiplication between each feature matrix along the channel dimension of the medical image feature map and the medical image masked spatial attention feature matrix to obtain the foreground salient medical image feature map; wherein, calculating the spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain the medical image spatial attention feature matrix includes: performing global average pooling processing on the medical image feature map along the channel dimension to obtain the medical image global average pooling feature matrix; inputting the medical image global average pooling feature matrix into the Sigmoid activation unit to obtain the medical image spatial attention feature matrix; wherein, performing masking processing on the medical image spatial attention feature matrix based on a predetermined threshold to obtain the medical image masked spatial attention feature matrix includes: in response to each position feature value in the medical image spatial attention feature matrix being greater than or equal to the predetermined threshold, taking the original value of the position feature value, otherwise setting it to zero.

[0036] In an embodiment of the present application, preferably, passing the medical image feature map through the main feature distribution mapping network based on the focus area map to obtain the foreground salient medical image feature map includes: using the main feature distribution mapping network based on the focus area map to process and enhance the medical image feature map with the following mapping formula to obtain the foreground salient medical image feature map; wherein, the mapping formula is: ; wherein, represents the medical image feature map, represents the number of channels of the medical image feature map, represents the medical image global average pooling feature matrix, represents the Sigmoid activation unit, represents the medical image spatial attention feature matrix, represents the masking process, is the predetermined threshold, represents the medical image masked spatial attention feature matrix, represents the foreground salient medical image feature map, represents element-wise multiplication.

[0037] In step S150, pixel - by - pixel position feature dispersion is performed on the foreground - salient medical image feature map to obtain a set of foreground - salient medical image pixel - level feature vectors. It should be understood that considering that the foreground - salient medical image feature map usually has a high resolution and each pixel contains rich medical image semantic information. Therefore, in order to be able to perform a deeper analysis and processing on the foreground - salient medical image features of each pixel in the foreground - salient medical image feature map, and to understand and analyze the main feature information of the medical image in a more fine - grained manner, in the technical solution of this application, pixel - by - pixel position feature dispersion is performed on the foreground - salient medical image feature map to obtain a set of foreground - salient medical image pixel - level feature vectors. It should be understood that the pixel - by - pixel position feature dispersion means extracting the features of each pixel in the foreground - salient medical image feature map and transforming these pixel - level features into a set of foreground - salient medical image pixel - level feature vectors. This emphasizes the extraction and integration of each medical image pixel - level feature to more precisely distinguish the tissue types and lesions of different medical images, so as to obtain a richer and more detailed medical image pixel - level feature representation and get a set of foreground - salient medical image pixel - level feature vectors containing more pixel - level fine - grained information.

[0038] In step S160, a cross-modal learning module is used to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors. Accordingly, considering that the set of foreground-salient medical image pixel-level feature vectors expresses the features of each foreground-salient pixel point in the medical image and can describe in detail the feature information of the foreground object (such as the lesion area). The sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors provides semantic information about the content of the medical image, such as lesion type, size, location, growth rate, which helps to understand the medical concepts and clinical features depicted in the image. Therefore, in order to better represent and reinforce the foreground-salient medical image pixel-level features based on the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors, in the technical solution of this application, a cross-modal learning module is used to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors. Specifically, the cross-modal learning module calculates the similarity between the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors, and based on the similarity and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors, performs feature matching update and reinforcement on each foreground-salient medical image pixel-level feature vector in the set of foreground-salient medical image pixel-level feature vectors, so as to use the medical image natural semantic descriptor granularity semantic encoding features to guide the foreground-salient medical image pixel-level feature analysis, thereby improving the interpretability of the foreground-salient medical image pixel-level feature content.

[0039] Specifically, in the embodiments of the present application, a cross-modal learning module is used to perform cross-modal learning reinforcement on the set of foreground saliency medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic coding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors, including: calculating the similarity between each foreground saliency medical image pixel-level feature vector in the set of foreground saliency medical image pixel-level feature vectors and all medical image natural semantic descriptor granularity semantic coding feature vectors in the sequence of medical image natural semantic descriptor granularity semantic coding feature vectors to obtain a set of medical image pixel-semantic similarity row vectors; calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors; and, respectively, using each medical image pixel-semantic similarity class support weight vector in the set of medical image pixel-semantic similarity class support weight vectors as a weight vector, calculating the weighted sum of the sequence of medical image natural semantic descriptor granularity semantic coding feature vectors to obtain the set of semantic-guided reinforced medical image pixel-level feature vectors.

[0040] More specifically, in the embodiments of the present application, calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors, including: taking each eigenvalue in the medical image pixel-semantic similarity row vector as the exponent of the natural constant to calculate the exponential function value with the natural constant as the base according to the position to obtain a medical image pixel-semantic similarity class support row vector; calculating the sum of the respective position eigenvalues in the medical image pixel-semantic similarity class support row vector to obtain a medical image pixel-semantic global similarity; and, calculating the position-wise division between each eigenvalue in the medical image pixel-semantic similarity class support row vector and the medical image pixel-semantic global similarity to obtain the semantic-guided reinforced medical image pixel-level feature vector.

[0041] In the embodiments of the present application, preferably, a cross-modal learning module is used to perform cross-modal learning reinforcement on the set of foreground saliency medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic coding feature vectors to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors, including: using the cross-modal learning module to perform cross-modal learning reinforcement on the set of foreground saliency medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic coding feature vectors with the following cross-modal formula to obtain the set of semantic-guided reinforced medical image pixel-level feature vectors; wherein, the cross-modal formula is: ; wherein, and are the set of foreground saliency medical image pixel-level feature vectors and the th foreground saliency medical image pixel-level feature vector and the th medical image natural semantic descriptor granularity semantic coding feature vector in the sequence of the medical image natural semantic descriptor granularity semantic coding feature vectors respectively, represents calculating the th foreground saliency medical image pixel-level feature vector and the th medical image natural semantic descriptor granularity semantic coding feature vector, and represent a similarity matrix composed of a plurality of the similarity values, represents the number of feature vectors in the sequence of the medical image natural semantic descriptor granularity semantic coding feature vectors, is the th semantic guidance enhanced medical image pixel-level feature vector in the set of the semantic guidance enhanced medical image pixel-level feature vectors, represents the exponential function with the natural constant as the base.

[0042] In step S170, the set of the semantic guidance enhanced medical image pixel-level feature vectors is subjected to feature aggregation to obtain an enhanced medical image feature map as the enhanced medical image feature. It should be understood that in order to comprehensively integrate each of the semantic guidance enhanced medical image pixel-level feature vectors to obtain a comprehensive medical image feature display, in the technical solution of the present application, the set of the semantic guidance enhanced medical image pixel-level feature vectors is subjected to feature aggregation to obtain an enhanced medical image feature map. That is, aggregating the set of the semantic guidance enhanced medical image pixel-level feature vectors can provide a more comprehensive understanding of the medical image content, which helps to improve the accuracy of subsequent segmentation.

[0043] In step S180, based on the enhanced medical image features, a medical image segmentation result is obtained. Specifically, in the embodiment of the present application, obtaining a medical image segmentation result based on the enhanced medical image features includes: passing the enhanced medical image feature map through a semantic segmenter based on the Softmax classification function to obtain a medical image segmentation result. That is, classification processing is performed on the enhanced medical image features obtained by aggregating the set of semantic-guided enhanced medical image pixel-level feature vectors, so as to intelligently obtain a medical image segmentation result according to the image pixel feature information guided by natural language of the medical image. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex textures and blurred boundaries of medical images can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.

[0044] In the technical solution of the present application, in a preferred embodiment, passing the enhanced medical image feature map through a semantic segmenter based on the Softmax classification function to obtain a medical image segmentation result includes: calculating the mean matrix and variance matrix of the enhanced medical image feature vectors after the enhanced medical image feature map is unfolded, where the value at each position of the mean matrix is the mean of a pair of eigenvalue at two positions corresponding to the position coordinates of the enhanced medical image feature vector respectively, and the value at each position of the variance matrix is the variance of a pair of eigenvalue at two positions corresponding to the position coordinates of the enhanced medical image feature vector respectively; multiplying the transposed vector of the enhanced medical image feature vector by the mean matrix to obtain a first intermediate vector and multiplying the variance matrix by the enhanced medical image feature vector to obtain a second intermediate vector, where the enhanced medical image feature vector is a column vector; calculating the sum of the dot product of the first intermediate vector and the transposed vector of the second intermediate vector to obtain a third intermediate vector; calculating the matrix product of the mean matrix and the variance matrix, and multiplying the transposed vector of the enhanced medical image feature vector by the matrix product to obtain a fourth intermediate vector; calculating the sum of the dot product of the third intermediate vector and the transposed vector of the fourth intermediate vector to obtain an optimized enhanced medical image feature vector; restoring the optimized enhanced medical image feature vector to an optimized enhanced medical image feature map; and passing the optimized enhanced medical image feature map through a semantic segmenter based on the Softmax classification function to obtain a medical image segmentation result.

[0045] That is, in order to improve the feature expression effect of the enhanced medical image feature map in the complex feature representation dimension of the cross-modal semantic-guided association representation based on the image semantic features and text semantic features of the medical image, in the above steps, the mean matrix and variance matrix for the group aggregation statistical evaluation of the feature value granularity local distribution of the enhanced medical image feature map are used as the retrieval and response distribution enhancement of the enhanced medical image feature map, and a group aggregation-based feature distribution open-domain reference-free distribution retrieval response framework for the enhanced medical image feature map is constructed to avoid the distribution response redundancy caused by the local overflow features of the enhanced medical image feature map through response superposition, so as to realize the fidelity constraint related to the response self-aggregation statistics of the enhanced medical image feature map from the cross-modal semantic-guided association representation of the complex feature representation dimension to the target image semantic segmentation domain during the iterative process, thereby improving the feature expression effect of the complex feature representation dimension of the enhanced medical image feature map during the model iteration process, and improving the accuracy of the medical image segmentation result obtained by the enhanced medical image feature map through the semantic segmenter based on the Softmax classification function.

[0046] It is worth mentioning that those of ordinary skill in the art should be aware that before applying the deep neural network model for inference, the deep neural network model needs to be trained first so that the deep neural network can implement specific functional capabilities.

[0047] Specifically, in the embodiment of the present application, it also includes training the semantic encoder including the word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus area map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function.

[0048] Figure 4 It is a flowchart for training the semantic encoder including the word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus area map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function in the medical image segmentation method based on natural language processing technology according to the embodiment of the present application. As Figure 4As shown, the training steps include: S310, obtaining training data, where the training data includes training medical images, natural language descriptions of the training medical images, and the ground truth of the medical image segmentation results; S320, performing word segmentation on the natural language description of the training medical image and then passing it through the semantic encoder including a word embedding layer to obtain a sequence of semantic encoding feature vectors at the word granularity of the natural semantic description of the training medical image; S330, passing the training medical image through the medical image feature extractor based on AlexNet to obtain a training medical image feature map; S340, passing the training medical image feature map through the main feature distribution mapping network based on the focus area map to obtain a training foreground salient medical image feature map; S350, performing per-pixel position feature dispersion on the training foreground salient medical image feature map to obtain a set of training foreground salient medical image pixel-level feature vectors; S360, using the cross-modal learning module to perform cross-modal learning reinforcement on the set of training foreground salient medical image pixel-level feature vectors and the sequence of semantic encoding feature vectors at the word granularity of the natural semantic description of the training medical image to obtain a set of training semantic-guided reinforced medical image pixel-level feature vectors; S370, aggregating the features of the set of training semantic-guided reinforced medical image pixel-level feature vectors to obtain a training reinforced medical image feature map; S380, optimizing the features of the training reinforced medical image feature map to obtain an optimized training reinforced medical image feature map; S390, passing the optimized training reinforced medical image feature map through the semantic segmenter based on the Softmax classification function to obtain a classification loss function value; and S400, based on the classification loss function value and through backpropagation of gradient descent to the semantic encoder including a word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus area map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function, where in each iteration of the model, the training reinforced medical image feature map is optimized.

[0049] Specifically, in the technical solution of the present application, during each iteration of the model, optimizing the training enhanced medical image feature map to obtain an optimized training enhanced medical image feature map includes: dividing each feature value of the training enhanced medical image feature map by the maximum feature value of the training enhanced medical image feature map to obtain a training enhanced medical image semantic interaction representation map; dividing the mean value of the feature values of the training enhanced medical image feature map by the standard deviation of the feature values of the training enhanced medical image feature map to obtain a statistical interaction value corresponding to the training enhanced medical image feature map; subtracting the training statistical interaction value from each feature value of the training enhanced medical image semantic interaction representation map, taking the absolute value, and calculating the logarithm value at each position to obtain a training enhanced medical image semantic interaction information representation map; adding the training statistical interaction value to each feature value of the training enhanced medical image semantic interaction representation map and then multiplying by a predetermined weight hyperparameter to obtain a training enhanced medical image semantic interaction pattern representation map; and adding the training enhanced medical image semantic interaction information representation map and the training enhanced medical image semantic interaction pattern representation map point by point to obtain the optimized training enhanced medical image feature map.

[0050] That is, in order to improve the feature expression effect of the training enhanced medical image feature map in the complex feature representation dimension of the cross-modal semantic guidance correlation representation based on the image semantic features and text semantic features of medical images, in the above steps, a short sequence including an interaction representation containing probability statistical features represented by the mean value and standard value and a distribution interaction representation of the feature value and the maximum feature value as latent variable features is used as a sub-manifold latent motif under the complex manifold network of the training enhanced medical image feature map, and the potential motif feature information pattern and feature distribution pattern of the training enhanced medical image feature map based on the training enhanced medical image semantic interaction information representation map and the training enhanced medical image semantic interaction pattern representation map are used as its global structure inference unit, so as to reconstruct the complex manifold structure of the training enhanced medical image feature map in the form of a global structure latent motif dictionary based on connection, so as to improve the model's ability to generate and evolve the understanding of the manifold structure corresponding to the features in the complex feature representation dimension during the iteration process, improve the feature expression effect of the model during the iteration process for the complex feature representation dimension of the training enhanced medical image feature map, and thus improve the accuracy of the medical image segmentation result obtained by the training enhanced medical image feature map through the semantic segmenter based on the Softmax classification function. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex texture and fuzzy boundaries of medical images can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.

[0051] In summary, the medical image segmentation method based on natural language processing technology according to the embodiments of the present application is elucidated. It obtains a medical image and a natural language description of the medical image, and uses image and text processing and analysis algorithms based on deep learning neural networks to analyze the medical image and process the natural language description of the medical image, so as to intelligently obtain a medical image segmentation result based on the image pixel feature information guided by the natural language of the medical image. Through this method, the anatomical structure and lesion conditions in the image can be better understood, and at the same time, the complex texture and fuzzy boundaries of the medical image can be processed, making the segmentation result more in line with the expectations of medical experts, thereby improving the accuracy and efficiency of segmentation.

[0052] Figure 5 FIG. is a block diagram of a medical image segmentation system based on natural language processing technology according to an embodiment of the present application. As Figure 5 shown, the medical image segmentation system 100 based on natural language processing technology according to an embodiment of the present application includes: a medical image data acquisition module 110 for obtaining a medical image and a natural language description of the medical image; a word-level semantic encoding module 120 for performing word segmentation on the natural language description of the medical image and then performing word-level semantic encoding to obtain a sequence of word-level semantic encoding feature vectors of the natural semantic description of the medical image; a medical image feature extraction module 130 for extracting medical image features from the medical image to obtain a medical image feature map; a feature distribution mapping module 140 for mapping the medical image feature map through a main feature distribution mapping network based on a focus area map to obtain a foreground-salient medical image feature map; a feature dispersion module 150 for performing per-pixel position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; a semantic guidance enhancement module 160 for using a cross-modal learning module to perform cross-modal learning enhancement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of word-level semantic encoding feature vectors of the natural semantic description of the medical image to obtain a set of semantic guidance enhanced medical image pixel-level feature vectors; a feature aggregation module 170 for aggregating the set of semantic guidance enhanced medical image pixel-level feature vectors to obtain an enhanced medical image feature map as an enhanced medical image feature; and a segmentation result generation module 180 for obtaining a medical image segmentation result based on the enhanced medical image feature.

[0053] Here, those skilled in the art can understand that the specific operations of each step in the above-mentioned medical image segmentation system based on natural language processing technology have been introduced in detail in the description of the medical image segmentation method based on natural language processing technology above, and therefore, the repeated description thereof will be omitted. Figures 1 to 4 of the medical image segmentation method based on natural language processing technology, and thus, the repeated description thereof will be omitted.

[0054] As described above, the medical image segmentation system 100 based on natural language processing technology according to an embodiment of the present disclosure can be implemented in various wireless terminals, such as a server having a medical image segmentation algorithm based on natural language processing technology. In a possible implementation, the medical image segmentation system 100 based on natural language processing technology according to an embodiment of the present disclosure can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the medical image segmentation system 100 based on natural language processing technology can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the medical image segmentation system 100 based on natural language processing technology can also be one of the many hardware modules of the wireless terminal.

[0055] Alternatively, in another example, the medical image segmentation system 100 based on natural language processing technology and the wireless terminal can also be separate devices, and the medical image segmentation system 100 based on natural language processing technology can be connected to the wireless terminal through a wired and / or wireless network, and transmit interaction information in accordance with a predefined data format.

[0056] The above are only examples of the principles of the present disclosure, and those skilled in the art can make various modifications without departing from the scope of the present disclosure. The above embodiments are presented for illustrative purposes rather than for limitation. The present disclosure can also take many forms other than those explicitly described herein. Therefore, it should be emphasized that the present disclosure is not limited to the explicitly disclosed methods, systems, and devices, but is intended to include variations and modifications within the spirit of the appended claims.

Claims

1. A medical image segmentation method based on natural language processing technology, characterized in that Including: Obtaining a medical image and a natural language description of the medical image; Performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image; performing medical image feature extraction on the medical image to obtain a medical image feature map; Passing the medical image feature map through a main feature distribution mapping network based on a focus area map to obtain a foreground-salient medical image feature map; Performing per-pixel position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; Using a cross-modal learning module to perform cross-modal learning reinforcement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image to obtain a set of semantic-guided reinforced medical image pixel-level feature vectors, including: calculating the similarity between each foreground-salient medical image pixel-level feature vector in the set of foreground-salient medical image pixel-level feature vectors and all word granularity semantic encoding feature vectors in the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image to obtain a set of medical image pixel-semantic similarity row vectors; calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors; respectively using each medical image pixel-semantic similarity class support weight vector in the set of medical image pixel-semantic similarity class support weight vectors as a weight vector, calculating the weighted sum of the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image to obtain the set of semantic-guided reinforced medical image pixel-level feature vectors; Aggregating the set of semantic-guided reinforced medical image pixel-level feature vectors to obtain a reinforced medical image feature map as the reinforced medical image feature; Based on the reinforced medical image feature, obtaining a medical image segmentation result.

2. The medical image segmentation method based on natural language processing technology according to claim 1, wherein Performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image, including: performing word segmentation on the natural language description of the medical image and then passing it through a semantic encoder including a word embedding layer to obtain the sequence of word granularity semantic encoding feature vectors of the natural semantic description of the medical image.

3. The medical image segmentation method based on natural language processing technology according to claim 2, wherein Performing medical image feature extraction on the medical image to obtain a medical image feature map, including: passing the medical image through a medical image feature extractor based on AlexNet to obtain the medical image feature map.

4. The medical image segmentation method based on natural language processing technology according to claim 3, characterized in that The medical image feature map is passed through a main feature distribution mapping network based on the focus region map to obtain a foreground salient medical image feature map, including: calculating the spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain a medical image spatial attention feature matrix; performing a masking process on the medical image spatial attention feature matrix based on a predetermined threshold to obtain a medical image masked spatial attention feature matrix; calculating the element-wise multiplication between each feature matrix along the channel dimension in the medical image feature map and the medical image masked spatial attention feature matrix to obtain the foreground salient medical image feature map; wherein, calculating the spatial attention features of each feature matrix along the channel dimension in the medical image feature map to obtain a medical image spatial attention feature matrix includes: performing global average pooling processing on the medical image feature map along the channel dimension to obtain a medical image global average pooling feature matrix; inputting the medical image global average pooling feature matrix into a Sigmoid activation unit to obtain the medical image spatial attention feature matrix; wherein, performing a masking process on the medical image spatial attention feature matrix based on a predetermined threshold to obtain a medical image masked spatial attention feature matrix includes: in response to each position feature value in the medical image spatial attention feature matrix being greater than or equal to the predetermined threshold, taking the original value of the position feature value, otherwise setting it to zero.

5. The medical image segmentation method based on natural language processing technology according to claim 4, wherein Calculating the class support weight vectors of each medical image pixel-semantic similarity row vector in the set of medical image pixel-semantic similarity row vectors to obtain a set of medical image pixel-semantic similarity class support weight vectors, including: taking the exponential of each eigenvalue in the medical image pixel-semantic similarity row vector as the base of the natural constant to calculate the element-wise exponential function value with the base of the natural constant to obtain a medical image pixel-semantic similarity class support row vector; calculating the sum of each position feature value in the medical image pixel-semantic similarity class support row vector to obtain a medical image pixel-semantic global similarity; calculating the element-wise division between each eigenvalue in the medical image pixel-semantic similarity class support row vector and the medical image pixel-semantic global similarity to obtain the semantic-guided enhanced medical image pixel-level feature vector.

6. The medical image segmentation method based on natural language processing technology according to claim 5, wherein, Based on the enhanced medical image features, obtaining a medical image segmentation result, including: passing the enhanced medical image feature map through a semantic segmenter based on the Softmax classification function to obtain a medical image segmentation result.

7. The medical image segmentation method based on natural language processing technology according to claim 6, wherein It also includes training the semantic encoder including a word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus region map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function.

8. The medical image segmentation method based on natural language processing technology according to claim 7, characterized in that, The training steps include: obtaining training data, where the training data includes training medical images and natural language descriptions of the training medical images, as well as the ground truth of the medical image segmentation results; performing word segmentation on the natural language description of the training medical image and then passing it through the semantic encoder including a word embedding layer to obtain a sequence of semantic encoding feature vectors at the word granularity of the natural semantic description of the training medical image; passing the training medical image through the medical image feature extractor based on AlexNet to obtain a training medical image feature map; passing the training medical image feature map through the main feature distribution mapping network based on the focus area map to obtain a training foreground-salient medical image feature map; performing per-pixel position feature dispersion on the training foreground-salient medical image feature map to obtain a set of training foreground-salient medical image pixel-level feature vectors; using the cross-modal learning module to perform cross-modal learning enhancement on the set of training foreground-salient medical image pixel-level feature vectors and the sequence of semantic encoding feature vectors at the word granularity of the natural semantic description of the training medical image to obtain a set of training semantic-guided enhanced medical image pixel-level feature vectors; aggregating the features of the set of training semantic-guided enhanced medical image pixel-level feature vectors to obtain a training enhanced medical image feature map; optimizing the features of the training enhanced medical image feature map to obtain an optimized training enhanced medical image feature map; passing the optimized training enhanced medical image feature map through the semantic segmenter based on the Softmax classification function to obtain a classification loss function value; and performing backpropagation based on the classification loss function value through gradient descent on the semantic encoder including a word embedding layer, the medical image feature extractor based on AlexNet, the main feature distribution mapping network based on the focus area map, the cross-modal learning module, and the semantic segmenter based on the Softmax classification function. Among them, in each iteration of the model, the training enhanced medical image feature map is optimized.

9. A medical image segmentation system based on natural language processing technology for performing the medical image segmentation method based on natural language processing technology according to any one of claims 1-8, characterized in that, including: A medical image data acquisition module for obtaining medical images and natural language descriptions of medical images; a word granularity semantic encoding module for performing word segmentation on the natural language description of the medical image and then performing semantic encoding based on word granularity to obtain a sequence of semantic encoding feature vectors at the word granularity of the natural semantic description of the medical image; a medical image feature extraction module for extracting medical image features from the medical image to obtain a medical image feature map; a feature distribution mapping module for passing the medical image feature map through the main feature distribution mapping network based on the focus area map to obtain a foreground-salient medical image feature map; a feature dispersion module for performing per-pixel position feature dispersion on the foreground-salient medical image feature map to obtain a set of foreground-salient medical image pixel-level feature vectors; A semantic guidance enhancement module, which is used to perform cross-modal learning enhancement on the set of foreground-salient medical image pixel-level feature vectors and the sequence of medical image natural semantic descriptor granularity semantic encoding feature vectors by using a cross-modal learning module to obtain a set of semantic guidance enhanced medical image pixel-level feature vectors; A feature aggregation module, which is used to aggregate the features of the set of semantic guidance enhanced medical image pixel-level feature vectors to obtain an enhanced medical image feature map as the enhanced medical image feature; A segmentation result generation module, which is used to obtain a medical image segmentation result based on the enhanced medical image feature.

Citation Information

Patent Citations

  • Semantic segmentation method with second-order pooling

    US20150104102A1

  • Harmonizing composite images utilizing a semantic-guided transformer neural network

    US20240161240A1