Ultrasonic image processing method and device, medium and equipment
Through ultrasonic endoscopic image processing method, segmented branches and classification branches are used to generate cross-region feature maps, which solves the problems of operator dependence and insufficient information utilization of ultrasonic image processing in the prior art, achieves more accurate tumor staging and model generalization capabilities, and enhances the support for clinical diagnosis.
Patent Information
- Application Number
- CN202510498651.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, in ultrasonic image processing, there are problems such as strong operator dependence, insufficient information utilization, poor model interpretability and insufficient prior knowledge application, resulting in the inability to accurately obtain effective medical information.
Ultrasonic endoscopic image processing method is used to generate cross-region feature maps through segmentation branches and classification branches of the ultrasonic endoscopic staging diagnostic model of esophageal cancer by mathematical operations, and dynamically adjust the segmentation weight of the loss function with the importance heat map to achieve multi-task classification.
It improves the accuracy of ultrasound image segmentation and the accuracy of tumor staging, enhances the generalization ability of the model, and provides richer medical information to support clinical diagnosis.
Smart Images

Figure CN120451062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ultrasonic image processing, and in particular to an ultrasonic image processing method, device, medium and equipment. Background Art
[0002] Esophageal cancer is a common malignant tumor of the digestive system, with high morbidity and mortality, posing a serious threat to human health. Early diagnosis and accurate staging are crucial for developing treatment plans and predicting patient prognosis. Endoscopic ultrasound (EUS), as an important imaging modality, can visually demonstrate the multilayered structure of the esophageal wall and the depth of tumor invasion, and therefore plays a crucial role in the preoperative staging of esophageal cancer.
[0003] Traditional EUS examination relies on the experience and skills of the endoscopist to assess the tumor stage by observing and interpreting EUS images. However, this method has the following shortcomings:
[0004] High operator dependence: The interpretation of EUS images is highly dependent on the experience of the endoscopist. The observation results between different physicians may vary greatly, affecting the accuracy and consistency of staging determination.
[0005] Task separation and insufficient information utilization: Traditional methods usually perform image segmentation and T stage classification tasks separately, failing to fully utilize the shared feature information between the two, resulting in insufficient feature extraction in key areas (such as the junction of the tumor and the esophageal wall).
[0006] Insufficient application of prior knowledge: Clinically, the structure of the various layers of the esophageal wall and their interrelationships provide an important basis for staging. However, existing algorithms rarely use medical prior knowledge to characterize the intersection area and are unable to capture the details of the tumor boundary well.
[0007] Insufficient model interpretability: Although deep learning technology has been widely used in the field of medical image processing, most model outputs are only black box results, which makes it difficult to provide clinicians with intuitive and reliable auxiliary information.
[0008] These deficiencies result in the inability of existing technologies to accurately obtain effective medical information based on ultrasound images. Summary of the Invention
[0009] The present invention provides an ultrasonic image processing method, device, medium and equipment to solve the problem in the prior art that effective medical information cannot be accurately obtained from ultrasonic images.
[0010] In a first aspect, the present application provides an ultrasound image processing method, comprising:
[0011] Obtain endoscopic ultrasound images;
[0012] Inputting the ultrasound endoscopic image into a segmentation branch of a preset ultrasound endoscopic staging diagnostic model for esophageal cancer, so that the segmentation branch outputs each basic segmentation map of the ultrasound endoscopic image and then outputs each cross-region feature map of the ultrasound endoscopic image according to a preset mathematical operation method and each basic segmentation map;
[0013] The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process;
[0014] Inputting each of the cross-region feature maps into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network;
[0015] The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
[0016] This application uses the segmentation branch of the preset esophageal cancer endoscopic ultrasound staging diagnosis model to process endoscopic ultrasound images, which can output high-quality basic segmentation maps. These segmentation maps provide an accurate basis for subsequent cross-region feature extraction. Then, through a preset mathematical operation method, cross-region feature maps are obtained from the basic segmentation map. These feature maps can highlight the relationship between the tumor and the various layers of the esophageal wall structure, providing richer information for tumor staging. In addition, the segmentation branch dynamically adjusts the segmentation weight of the loss function according to the importance heat map during the training process. This innovative mechanism enables the model to pay more attention to key areas, thereby improving the accuracy of segmentation. Finally, the cross-region feature map is input into the classification branch to complete multi-task classification, including key staging tasks such as T1a / T1b, T2 / T3-4 and T1a-T4. This multi-task learning method not only improves the accuracy of classification, but also enhances the generalization ability of the model. This application effectively solves the problem that the existing technology cannot accurately obtain effective medical information based on ultrasound images.
[0017] As a preferred embodiment of the first aspect, the outputting of the characteristic maps of each intersection area of the ultrasound endoscopy image according to the preset mathematical operation method and the each basic segmentation map is specifically:
[0018] The segmentation branch performs an element-by-element product operation on the multiple segmentation results in each basic segmentation map, and outputs each intersection region feature map;
[0019] The multiple segmentation results include a tumor region segmentation map, a submucosal layer segmentation map, an adventitia segmentation map, and various cropping frames.
[0020] In this preferred embodiment, the present application performs an element-by-element multiplication operation on multiple segmentation results in the basic segmentation map, which include segmentation maps of key anatomical structures such as the tumor area, submucosal layer, outer membrane, and cropping frame. This operation can highlight the relationships and interaction areas between different anatomical structures, so that the generated cross-region feature map can provide richer image feature information than a single segmentation result. These feature maps can not only reveal the complex interactions between the tumor and its surrounding tissues, but also emphasize those subtle structural differences that are crucial for tumor staging. Therefore, the cross-region feature map generated by this method can provide more accurate and discriminative features for subsequent tumor staging classification tasks, thereby improving the diagnostic performance of the entire system and enhancing the accuracy of clinicians' judgment of tumor staging.
[0021] As a preferred embodiment of the first aspect, the segmentation branch is trained on the initial segmentation branch based on historical ultrasound endoscopic images, and the segmentation weights of the loss function are dynamically adjusted according to preset importance heat maps during the training process, specifically:
[0022] Obtain historical endoscopic ultrasound images;
[0023] Inputting the historical ultrasound endoscopy image into the initial segmentation branch for training, so that the initial segmentation branch outputs each initial basic segmentation image;
[0024] Calculating the center point of the esophageal structure according to the probability distribution of the submucosal segmentation map and the adventitia segmentation map of each initial basic segmentation map;
[0025] Dividing each of the initial basic segmentation images into N sector-shaped regions according to the center point of the esophageal structure; wherein N is a positive integer greater than 1;
[0026] Calculating an importance heat map of each initial basic segmentation map according to the fan-shaped region and the tumor region segmentation map of each initial basic segmentation map;
[0027] Calculate each weight based on the importance heat map and a preset threshold;
[0028] According to the weights, the loss function of the initial segmentation branch is dynamically adjusted. If the value of the loss function no longer decreases within a preset time period, the training is stopped to obtain the segmentation branch.
[0029] The loss function is specifically:
[0030] The loss function is a combination of binary cross entropy loss and Dice loss;
[0031] The formula of the loss function is:
[0032]
[0033] Where, ω c represents the weight of the channel, represents the binary cross entropy loss, Denotes Dice loss.
[0034] In this preferred embodiment, the present application acquires historical ultrasound endoscopic images and inputs them into the initial segmentation branch for training, generating an initial basic segmentation map. The probability distribution of the submucosal and adventitial layers in these segmentation maps is then used to calculate the center point of the esophageal structure. Based on this, the image is divided into N sectors. Combining the sector and tumor region segmentation maps, an importance heatmap is calculated for each region, which helps identify areas of the image that are critical for tumor staging. Next, the segmentation weights in the loss function are dynamically adjusted based on these importance heatmaps and weights calculated using a preset threshold. This dynamic adjustment mechanism ensures that the model focuses on optimizing feature recognition in key areas during training. Ultimately, when the loss function value no longer decreases significantly within a preset timeframe, training ceases, resulting in the optimized segmentation branch. The loss function combines a binary cross-entropy loss and a Dice loss. This combination of two losses not only considers pixel-level classification accuracy but also the degree of overlap between the predicted and true segmented regions, thereby enhancing the model's ability to identify key regions. Therefore, this approach not only improves segmentation accuracy but also enhances the model's ability to predict tumor staging by emphasizing the features of key regions, providing a more reliable basis for clinical diagnosis.
[0035] As a preferred embodiment of the first aspect, the present invention further includes:
[0036] The preset esophageal cancer endoscopic ultrasound staging diagnosis model fuses the loss of the segmentation branch with the loss of the classification branch through a joint loss function and performs collaborative optimization;
[0037] The joint loss function is:
[0038]
[0039] Where, is the loss function of the segmentation branch, which consists of binary cross entropy loss and Dice loss; and are the weighted cross entropy losses in the classification tasks respectively; seg ,λ binary and λ multi is the weight coefficient of each loss item.
[0040] In this preferred embodiment, the present application organically combines the loss of the segmentation task (such as Dice loss and binary cross entropy loss) with the loss of the classification task (such as weighted cross entropy loss) by designing a joint loss function, so that the two branches can promote each other and optimize together during the training process. This collaborative optimization strategy not only improves the model's ability to extract key regional features, but also enhances the model's ability to discriminate tumor stages, ultimately achieving more precise image segmentation and more accurate tumor staging, providing strong support for clinical diagnosis and treatment.
[0041] In a second aspect, the present application provides an ultrasonic image processing device. The ultrasonic image processing device includes an acquisition module, a segmentation module, and a classification module;
[0042] The acquisition module is used to obtain an ultrasound endoscopic image;
[0043] The segmentation module is used to input the ultrasonic endoscopic image into the segmentation branch of the preset ultrasonic endoscopic staging diagnosis model for esophageal cancer, so that after the segmentation branch outputs each basic segmentation map of the ultrasonic endoscopic image, it outputs each cross-region feature map of the ultrasonic endoscopic image according to a preset mathematical operation method and each basic segmentation map;
[0044] The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process;
[0045] The classification module is used to input the respective cross-region feature maps into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network;
[0046] The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
[0047] This device uses three modules to divide the work and coordinate work to more accurately process ultrasound images. This application uses the segmentation branch of the preset esophageal cancer ultrasound endoscopic staging diagnosis model to process ultrasound endoscopic images, which can output high-quality basic segmentation maps. These segmentation maps provide an accurate basis for subsequent cross-region feature extraction. Then, through a preset mathematical operation method, cross-region feature maps are obtained from the basic segmentation map. These feature maps can highlight the relationship between the tumor and the various layers of the esophageal wall structure, providing richer information for tumor staging. In addition, the segmentation branch dynamically adjusts the segmentation weight of the loss function according to the importance heat map during training. This innovative mechanism enables the model to pay more attention to key areas, thereby improving the accuracy of segmentation. Finally, the cross-region feature map is input into the classification branch to complete multi-task classification, including key staging tasks such as T1a / T1b, T2 / T3-4 and T1a-T4. This multi-task learning method not only improves the accuracy of classification, but also enhances the generalization ability of the model. This application effectively solves the problem that the existing technology cannot accurately obtain effective medical information based on ultrasound images.
[0048] As a preferred embodiment of the second aspect, the outputting of the characteristic maps of each intersection area of the ultrasound endoscopy image according to the preset mathematical operation method and the each basic segmentation map is specifically:
[0049] The segmentation branch performs an element-by-element product operation on the multiple segmentation results in each basic segmentation map, and outputs each intersection region feature map;
[0050] The multiple segmentation results include a tumor region segmentation map, a submucosal layer segmentation map, an adventitia segmentation map, and various cropping frames.
[0051] In this preferred embodiment, the present application performs an element-by-element multiplication operation on multiple segmentation results in the basic segmentation map, which include segmentation maps of key anatomical structures such as the tumor area, submucosal layer, outer membrane, and cropping frame. This operation can highlight the relationships and interaction areas between different anatomical structures, so that the generated cross-region feature map can provide richer image feature information than a single segmentation result. These feature maps can not only reveal the complex interactions between the tumor and its surrounding tissues, but also emphasize those subtle structural differences that are crucial for tumor staging. Therefore, the cross-region feature map generated by this method can provide more accurate and discriminative features for subsequent tumor staging classification tasks, thereby improving the diagnostic performance of the entire system and enhancing the accuracy of clinicians' judgment of tumor staging.
[0052] As a preferred embodiment of the second aspect, the segmentation branch is trained on the initial segmentation branch based on historical ultrasound endoscopic images, and the segmentation weights of the loss function are dynamically adjusted according to preset importance heat maps during the training process, specifically:
[0053] Obtain historical endoscopic ultrasound images;
[0054] Inputting the historical ultrasound endoscopy image into the initial segmentation branch for training, so that the initial segmentation branch outputs each initial basic segmentation image;
[0055] Calculating the center point of the esophageal structure according to the probability distribution of the submucosal segmentation map and the adventitia segmentation map of each initial basic segmentation map;
[0056] Dividing each of the initial basic segmentation images into N sector-shaped regions according to the center point of the esophageal structure; wherein N is a positive integer greater than 1;
[0057] Calculating an importance heat map of each initial basic segmentation map according to the fan-shaped region and the tumor region segmentation map of each initial basic segmentation map;
[0058] Calculate each weight based on the importance heat map and a preset threshold;
[0059] According to the weights, the loss function of the initial segmentation branch is dynamically adjusted. If the value of the loss function no longer decreases within a preset time period, the training is stopped to obtain the segmentation branch.
[0060] In this preferred embodiment, the present application trains the initial segmentation branch by acquiring and utilizing historical ultrasound endoscopic images to generate an initial basic segmentation map. Next, based on the probability distribution of the submucosal layer and adventitia in these segmentation maps, the center points of the esophageal structures are calculated. These center points reflect the locations of key structures in the image. Based on these center points, the image is then divided into multiple sector-shaped regions. Each region is combined with the segmentation map of the tumor region to calculate an importance heat map, thereby identifying regions critical for tumor staging. Based on these heat maps and preset thresholds, weights are calculated for each region, which are used to dynamically adjust the segmentation weights in the loss function. This dynamic adjustment ensures that the model focuses more on the regions most critical for improving segmentation accuracy and tumor staging accuracy during training. Ultimately, when the loss function value stops decreasing significantly within a certain period of time, training is terminated, resulting in the optimized segmentation branch. This approach not only improves segmentation accuracy but also enhances the model's ability to predict tumor staging by emphasizing the characteristics of key regions.
[0061] In a third aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium including a stored computer program. When the computer program is executed, the computer-readable storage medium controls a device containing the computer-readable storage medium to execute the ultrasound image processing method described above. The beneficial effects thereof are the same as those of the ultrasound image processing method provided in the first aspect of the present application.
[0062] In a fourth aspect, the present application provides a terminal device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the ultrasound image processing methods described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 : A schematic flow chart of an embodiment of the ultrasound image processing method provided in this application;
[0064] Figure 2 : A flow chart of an embodiment of the ultrasound image data processing process provided by this application;
[0065] Figure 3 : A structural diagram of an embodiment of the overall structure of the model provided in this application;
[0066] Figure 4 : A schematic structural diagram of an embodiment of the tumor heat map enhancement mechanism provided in this application;
[0067] Figure 5 : A structural schematic diagram of an embodiment of the ultrasonic image processing device provided in this application. DETAILED DESCRIPTION
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0069] Example 1
[0070] Please refer to Figure 1 , which is an ultrasonic image processing method provided by an embodiment of the present invention.
[0071] In this embodiment, the process of the ultrasound image processing method in this application is described in detail through steps S01-S03.
[0072] Esophageal cancer is a common malignant tumor with a poor prognosis in the digestive system. Early diagnosis and staging evaluation are crucial for the selection of treatment options. Traditional endoscopic ultrasound (EUS) examination can visually display the multi-layer structure of the esophageal wall and the depth of tumor infiltration (including the mucosal layer, muscularis mucosa, submucosal layer, muscularis propria and adventitial layer), and is therefore regarded as the gold standard for preoperative staging of esophageal cancer. According to data from the World Health Organization, the early detection rate of digestive tract cancer in my country is less than 20%, and accurate identification of T1-T4 stages, especially distinguishing T1a / T1b substages, is of great value for treatment decision-making. T classification (Tumor Classification) is an indicator for evaluating the depth of local invasion in the tumor TNM staging system, which is divided according to the degree of tumor invasion of the hierarchical structure of the esophageal wall:
[0073] -T1: The tumor is confined to the mucosa (T1a) or invades the submucosa (T1b);
[0074] -T2: The tumor invades the muscularis propria;
[0075] -T3: The tumor penetrates the muscularis propria and reaches the adventitia;
[0076] -T4: The tumor invades adjacent organs;
[0077] Accurately distinguishing T stages is crucial for developing treatment plans. For example, endoscopic resection is an option for stage T1a, while surgery is considered for stage T1b. Endoscopic ultrasound, which demonstrates the layered structure of the esophageal wall, has become the preferred imaging method for determining T stage.
[0078] This application addresses the problems existing in the field of EUS image processing for esophageal cancer, proposes a deep learning model for joint segmentation and classification, and introduces a tumor heat map enhancement mechanism.
[0079] The overall structure of this application is as follows Figure 2 As shown, this application adopts a "dual-branch" model design, using a homogeneous backbone network to extract global features, serving the segmentation branch and the classification branch respectively.
[0080] More specifically, Figure 3 As shown, the backbone feature network of this application can use the pre-trained ConvNeXt-Base model as the backbone network of both the segmentation branch and the classification branch, or can use the backbone of other classification models. After a series of convolutions and nonlinear activations, a feature map with multi-layer semantic information is generated.
[0081] As an example, the segmentation branch can be designed based on the UNet structure, using skip connections and upsampling mechanisms to achieve multi-scale feature fusion. The model inputs an ultrasound endoscopic image and outputs segmentation results in multiple channels, including a basic segmentation map (4 channels) and a cross-region segmentation map (4 channels).
[0082] The classification branch can use the cross-region feature map (4 channels) output by the segmentation branch as input, extract low-dimensional feature vectors through ConvNeXt-Base and MLP modules, and design a multi-task classification head to complete the T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks respectively;
[0083] This application also uses the probability map of the submucosal layer and the outer membrane in the segmentation results to calculate the center position of the image, and divides the image into sector regions. The tumor area is calculated in each region, and weights are assigned according to a preset threshold. Finally, an importance heat map is generated to guide the weighting of the segmentation loss;
[0084] The model usage process for this application is as follows:
[0085] Input image: Obtain esophageal endoscopic ultrasound image and perform preprocessing.
[0086] Generating lesion and esophageal structure segmentation: The image is fed into the segmentation branch network, which outputs eight channels of segmentation results, including four basic segmentation maps and four cross-region segmentation maps. The four basic segmentation map channels are used for lesion and esophageal structure visualization, while the four cross-region segmentation maps are used for downstream staging classification prediction tasks.
[0087] Stage classification prediction: The four cross-region channels are input into the classification branch network to obtain prediction results for three tasks: T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification task.
[0088] S01: Obtain an endoscopic ultrasound image.
[0089] S02: inputting the ultrasound endoscopic image into a segmentation branch of a preset ultrasound endoscopic staging diagnostic model for esophageal cancer, so that the segmentation branch outputs each basic segmentation map of the ultrasound endoscopic image and then outputs each intersection region feature map of the ultrasound endoscopic image according to a preset mathematical operation method and each basic segmentation map;
[0090] The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process.
[0091] As a preferred embodiment of the first embodiment, the endoscopic ultrasound image is input into a segmentation branch of a preset endoscopic ultrasound staging diagnosis model for esophageal cancer, so that the segmentation branch outputs each basic segmentation map of the endoscopic ultrasound image, specifically:
[0092] The final output of the segmentation branch network includes four basic segmentation maps: tumor area L, submucosal layer SM, adventitia AD and cropping box BBOX.
[0093] Based on the above output, four intersection region segmentation maps are generated through mathematical operations:
[0094] L \SM =L⊙(1-SM),
[0095] SM \L =SM⊙(1-L),
[0096] L \AD =L⊙(1-AD),
[0097] AD \L =AD⊙(1-L),
[0098] where ⊙ represents the element-wise product.
[0099] Finally, these four cross channels are concatenated with the basic segmentation map to form an enhanced feature map for subsequent classification tasks.
[0100] In this preferred embodiment, the four basic segmentation maps output by the segmentation branch network of the present application provide detailed anatomical structure information for subsequent image analysis, which is a prerequisite for accurate tumor staging. Then, based on these basic segmentation maps, four cross-region segmentation maps are generated through mathematical operations (such as element-by-element multiplication). These cross-region segmentation maps can highlight the relationship between different anatomical structures, especially the boundary area between the tumor and the surrounding normal tissue, which is extremely important for accurate tumor staging. Ultimately, the enhanced feature map formed by splicing these four cross channels with the basic segmentation map not only contains rich anatomical structure information, but also emphasizes the cross-regions that are crucial for tumor staging, thereby providing more abundant and discriminative features for subsequent classification tasks.
[0101] As a preferred embodiment of the first embodiment, the segmentation branch is trained on the initial segmentation branch based on historical ultrasound endoscopic images, and the segmentation weight of the loss function is dynamically adjusted according to preset importance heat maps during the training process, specifically:
[0102] Using ConvNeXt-Base as the backbone network, multi-layer features are extracted. After multiple convolution modules and pooling operations, the spatial resolution is gradually reduced and the abstraction ability of semantic features is improved.
[0103] Each stage outputs intermediate features and uses UNet as a decoder for feature fusion;
[0104] The segmentation task is optimized using a combination of binary cross entropy (BCE) and Dice loss, where Dice loss can effectively handle the problem of class imbalance.
[0105] At the same time, tumor heatmap weighting is introduced in some areas, and a larger loss weight is applied to key areas to encourage the model to focus on the boundary details between the tumor and surrounding tissues. Let the predicted probability be p and the true value be t (the ignored label is masked), then for each channel c, define:
[0106]
[0107] Where Ω represents the effective pixel set (for some channels, Ω is multiplied by the importance heat map H), ω c Indicates the weight of the channel. In this application, the weight of the BBox channel is set to 0.5, and the rest are 1. The final combination is:
[0108]
[0109] The respective importance heat maps, i.e., tumor heat maps, are specifically as follows Figure 4 As shown:
[0110] 1. Center positioning:
[0111] The probability map M of the submucosa SM and adventitia AD output by the segmentation branch SM and M AD ,
[0112] Calculate the weighted center:
[0113]
[0114] This center position serves as the reference point for subsequent sector-shaped area division.
[0115] 2. Sector area division:
[0116] Taking the center point as the reference, the image is evenly divided into N sector-shaped areas, and the coverage angle of each area is
[0117] Number each sector area S0, S1, ..., S N-1 .
[0118] 3. Regional tumor area calculation and weight allocation:
[0119] For each sector area S k , calculate tumor area Among them Lij is the probability value of the corresponding pixel in the tumor segmentation map;
[0120] The preset threshold T is used to determine whether the tumor area in the region reaches a significant level, and the weight is defined:
[0121]
[0122] The weights of each region are used to construct a heatmap H of the overall image importance, and the key regions are weighted in the segmentation loss calculation, so that the model pays more attention to these regions during training.
[0123] In this preferred embodiment, ConvNeXt-Base is used as the backbone network. Multiple convolutional modules and pooling operations are used to extract multi-layer features, and the spatial resolution is gradually reduced to improve the abstraction of semantic features. The intermediate features output by each stage are then fused through the UNet decoder to generate a high-quality basic segmentation map. Next, a combination of binary cross entropy (BCE) and Dice loss is used for optimization. The Dice loss can effectively address class imbalance and improve segmentation accuracy. In addition, a tumor heatmap weighting mechanism is introduced to assign greater loss weights to key regions, encouraging the model to pay more attention to the boundary details between the tumor and surrounding tissue. Specifically, the image is divided into multiple sector-shaped regions by calculating the weighted center of the probability map of the submucosal layer and the adventitia. The tumor area within each region is calculated. A preset threshold is used to determine whether the tumor area within the region reaches a significant level, and weights are defined accordingly. These weights are used to construct a global importance heatmap and weight key regions in the segmentation loss calculation, thereby increasing the model's attention to these regions during training. Through this method, the model can more accurately identify and segment key structures in ultrasound endoscopic images, providing a more reliable basis for clinical diagnosis. It has significant clinical application value and broad application prospects.
[0124] S03: inputting each of the cross-region feature maps into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network;
[0125] The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
[0126] As a preferred embodiment of the first embodiment, the inputting of each cross-region feature map into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network, is specifically as follows:
[0127] The classification branch is mainly responsible for determining the T stage of the image after cross-region feature extraction. Its main design is as follows:
[0128] Multi-channel segmentation feature extraction and two-stage model design:
[0129] The segmentation branch outputs an 8-channel segmentation map to fully extract the intrinsic structural information of the ultrasound endoscopic image.
[0130] The classification branch makes staging judgments based on the structural information mined by the segmentation branch, which is more accurate and more interpretable than the single classification model that directly performs staging on ultrasound endoscopic images.
[0131] For different T-stage tasks, three classification heads are designed:
[0132] T1a / T1b binary classification head: used for early staging determination.
[0133] T2 / T3-4 binary classification head: used to distinguish between middle and late stages.
[0134] T1a-T4 five-category head: enables more fine-grained staging determination.
[0135] Each classification head adopts the MLP structure, and the specific formula is as follows:
[0136] MLP(x)=W2Dropout(ReLU(W1x+b1))+b2,
[0137] Among them, W1, W2, b1 and b2 are corresponding parameters.
[0138] For different tasks, cross entropy losses with different category weights are used to effectively balance the problem of uneven sample size.
[0139] As a preferred embodiment of the first embodiment, it also includes:
[0140] The classification loss used in this application is weighted cross-entropy loss. Since the dataset uses a balanced sampling strategy to ensure a balanced number of samples across categories in each minibatch, we set weights such as [4.0, 1.0] for the T1a / T1b tasks and [1.0, 1.5] for the T2 / T3-4 tasks. Standard cross-entropy loss is used for the T1a-T4 tasks.
[0141] The classification loss and segmentation loss are fused in the joint loss function through preset weight parameters, and jointly participate in back-propagation optimization, thereby realizing multi-task collaborative training.
[0142] To achieve collaborative optimization of segmentation and classification tasks, this application designs a joint loss function:
[0143]
[0144] in: It is the segmentation loss, which is composed of BCE and Dice loss, and is combined with the heat map H to perform key area weighting; and are the weighted cross entropy losses for each classification task; seg ,λ binary and λ multi is the weight coefficient of each loss item, and the optimal value is determined through experimental tuning.
[0145] During training, the AdamW optimizer was used, combined with a cosine annealing schedule to dynamically adjust the learning rate. The first five epochs served as a warm-up phase, using a linear ramp-up strategy. Subsequently, the learning rate was gradually reduced using a cosine function to ensure smooth convergence throughout the training process.
[0146] In this preferred embodiment, the present application organically combines the loss of the segmentation task (such as Dice loss and binary cross entropy loss) with the loss of the classification task (such as weighted cross entropy loss) by designing a joint loss function, so that the two branches can promote each other and optimize together during the training process. This collaborative optimization strategy not only improves the model's ability to extract key regional features, but also enhances the model's ability to discriminate tumor stages, ultimately achieving more precise image segmentation and more accurate tumor staging, providing strong support for clinical diagnosis and treatment.
[0147] As a preferred embodiment of Example 1, this application conducted experiments on an EUS image dataset of 271 esophageal cancer patients using three-fold cross validation, and obtained the following main experimental results:
[0148] 1. Segmentation performance:
[0149] -Lesion segmentation: The Lesion Dice coefficient based on the ConvNeXt+UNet model was 87.01%, which was increased to 89.20% after the introduction of the joint segmentation and classification model, and then reached 90.32% through the tumor heat map mechanism.
[0150] -SM segmentation: The Dice coefficient of the base model is 92.72%, and the joint model is improved to 94.52%. Although there is a slight fluctuation after heat map enhancement, the overall performance remains above 93%.
[0151] -AD segmentation: increased from 92.22% to 94.56%, indicating that the cross-region features and heat map mechanism have a significant promoting effect on boundary region recognition.
[0152] 2. Classification performance:
[0153] -T1a / T1b binary classification: The best classification baseline had an accuracy of 73.12% and an AUC of 69.26%. The combined segmentation and classification model improved these to 75.95% and 79.79%, respectively. After introducing the heatmap mechanism, the accuracy reached 81.01% and the AUC was 82.45%.
[0154] -T2 / T3-4 binary classification: The joint model achieved an accuracy of 88.54% and an AUC exceeding 94%. The heatmap mechanism was slightly improved to 89.06% and 94.44%.
[0155] -T1a-T4 five-category classification: The overall accuracy increased from 69.45% of the baseline to 73.06% of the joint model. The heatmap mechanism increased the accuracy to 76.01%, and the AUC also improved accordingly.
[0156] 3. Evaluation indicators:
[0157] - Segmentation indicators: Dice coefficient, IoU, precision, recall and other indicators are used for evaluation;
[0158] -Classification indicators: using multi-dimensional indicators such as accuracy, sensitivity, specificity, AUC and confusion matrix;
[0159] - Cross-validation: The average performance, standard deviation, and 95% confidence interval of each fold are recorded in detail in the experimental report, which fully verifies the robustness and stability of the model.
[0160] The above experimental results show that the joint optimization strategy of fine segmentation and multi-task classification in this application is significantly better than the traditional single-task processing method, especially in the judgment of early stages (T1a / T1b), the model's capture effect of subtle boundary features is significantly improved.
[0161] This application uses the segmentation branch of the preset esophageal cancer endoscopic ultrasound staging diagnosis model to process endoscopic ultrasound images, which can output high-quality basic segmentation maps. These segmentation maps provide an accurate basis for subsequent cross-region feature extraction. Then, through a preset mathematical operation method, cross-region feature maps are obtained from the basic segmentation map. These feature maps can highlight the relationship between the tumor and the various layers of the esophageal wall structure, providing richer information for tumor staging. In addition, the segmentation branch dynamically adjusts the segmentation weight of the loss function according to the importance heat map during the training process. This innovative mechanism enables the model to pay more attention to key areas, thereby improving the accuracy of segmentation. Finally, the cross-region feature map is input into the classification branch to complete multi-task classification, including key staging tasks such as T1a / T1b, T2 / T3-4 and T1a-T4. This multi-task learning method not only improves the accuracy of classification, but also enhances the generalization ability of the model. This application effectively solves the problem that the existing technology cannot accurately obtain effective medical information based on ultrasound images.
[0162] Example 2
[0163] Please refer to Figure 5 , is an ultrasonic image processing device provided in an embodiment of the present application.
[0164] In this embodiment, the ultrasound image processing apparatus includes an acquisition module 10 , a segmentation module 20 and a classification module 30 .
[0165] Esophageal cancer is a common malignant tumor with a poor prognosis in the digestive system. Early diagnosis and staging evaluation are crucial for the selection of treatment options. Traditional endoscopic ultrasound (EUS) examination can visually display the multi-layer structure of the esophageal wall and the depth of tumor infiltration (including the mucosal layer, muscularis mucosa, submucosal layer, muscularis propria and adventitial layer), and is therefore regarded as the gold standard for preoperative staging of esophageal cancer. According to data from the World Health Organization, the early detection rate of digestive tract cancer in my country is less than 20%, and accurate identification of T1-T4 stages, especially distinguishing T1a / T1b substages, is of great value for treatment decision-making. T classification (Tumor Classification) is an indicator for evaluating the depth of local invasion in the tumor TNM staging system, which is divided according to the degree of tumor invasion of the hierarchical structure of the esophageal wall:
[0166] -T1: The tumor is confined to the mucosa (T1a) or invades the submucosa (T1b);
[0167] -T2: The tumor invades the muscularis propria;
[0168] -T3: The tumor penetrates the muscularis propria and reaches the adventitia;
[0169] -T4: The tumor invades adjacent organs;
[0170] Accurately distinguishing T stages is crucial for developing treatment plans. For example, endoscopic resection is an option for stage T1a, while surgery is considered for stage T1b. Endoscopic ultrasound, which demonstrates the layered structure of the esophageal wall, has become the preferred imaging method for determining T stage.
[0171] This application addresses the problems existing in the field of EUS image processing for esophageal cancer, proposes a deep learning model for joint segmentation and classification, and introduces a tumor heat map enhancement mechanism.
[0172] The overall structure of this application is as follows Figure 2 As shown, this application adopts a "dual-branch" model design, using a homogeneous backbone network to extract global features, serving the segmentation branch and the classification branch respectively.
[0173] More specifically, Figure 3 As shown, the backbone feature network of this application can use the pre-trained ConvNeXt-Base model as the backbone network of both the segmentation branch and the classification branch, or can use the backbone of other classification models. After a series of convolutions and nonlinear activations, a feature map with multi-layer semantic information is generated.
[0174] As an example, the segmentation branch can be designed based on the UNet structure, using skip connections and upsampling mechanisms to achieve multi-scale feature fusion. The model inputs an ultrasound endoscopic image and outputs segmentation results in multiple channels, including a basic segmentation map (4 channels) and a cross-region segmentation map (4 channels).
[0175] The classification branch can use the cross-region feature map (4 channels) output by the segmentation branch as input, extract low-dimensional feature vectors through ConvNeXt-Base and MLP modules, and design a multi-task classification head to complete the T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks respectively;
[0176] This application also uses the probability map of the submucosal layer and the outer membrane in the segmentation results to calculate the center position of the image, and divides the image into sector regions. The tumor area is calculated in each region, and weights are assigned according to a preset threshold. Finally, an importance heat map is generated to guide the weighting of the segmentation loss;
[0177] The model usage process for this application is as follows:
[0178] Input image: Obtain esophageal endoscopic ultrasound image and perform preprocessing.
[0179] Generating lesion and esophageal structure segmentation: The image is fed into the segmentation branch network, which outputs eight channels of segmentation results, including four basic segmentation maps and four cross-region segmentation maps. The four basic segmentation map channels are used for lesion and esophageal structure visualization, while the four cross-region segmentation maps are used for downstream staging classification prediction tasks.
[0180] Stage classification prediction: The four cross-region channels are input into the classification branch network to obtain prediction results for three tasks: T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification task.
[0181] The acquisition module 10 is used to acquire an ultrasonic endoscopic image.
[0182] The segmentation module 20 is configured to input the endoscopic ultrasound image into a segmentation branch of a preset endoscopic ultrasound staging diagnostic model for esophageal cancer, so that after the segmentation branch outputs each basic segmentation map of the endoscopic ultrasound image, it outputs each cross-region feature map of the endoscopic ultrasound image based on a preset mathematical operation method and each basic segmentation map;
[0183] The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process.
[0184] As a preferred embodiment of the second embodiment, the endoscopic ultrasound image is input into a segmentation branch of a preset endoscopic ultrasound staging diagnosis model for esophageal cancer, so that the segmentation branch outputs each basic segmentation map of the endoscopic ultrasound image, specifically:
[0185] The final output of the segmentation branch network includes four basic segmentation maps: tumor area L, submucosal layer SM, adventitia AD and cropping box BBOX.
[0186] Based on the above output, four intersection region segmentation maps are generated through mathematical operations:
[0187] L \SM =L⊙(1-SM),
[0188] SM \L =SM⊙(1-L),
[0189] L \AD =L⊙(1-AD),
[0190] AD \L =AD⊙(1-L),
[0191] where ⊙ represents the element-wise product.
[0192] Finally, these four cross channels are concatenated with the basic segmentation map to form an enhanced feature map for subsequent classification tasks.
[0193] In this preferred embodiment, the four basic segmentation maps output by the segmentation branch network of the present application provide detailed anatomical structure information for subsequent image analysis, which is a prerequisite for accurate tumor staging. Then, based on these basic segmentation maps, four cross-region segmentation maps are generated through mathematical operations (such as element-by-element multiplication). These cross-region segmentation maps can highlight the relationship between different anatomical structures, especially the boundary area between the tumor and the surrounding normal tissue, which is extremely important for accurate tumor staging. Ultimately, the enhanced feature map formed by splicing these four cross channels with the basic segmentation map not only contains rich anatomical structure information, but also emphasizes the cross-regions that are crucial for tumor staging, thereby providing more abundant and discriminative features for subsequent classification tasks.
[0194] As a preferred embodiment of the second embodiment, the segmentation branch is trained on the initial segmentation branch based on the historical ultrasound endoscopy image, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process, specifically:
[0195] Using ConvNeXt-Base as the backbone network, multi-layer features are extracted. After multiple convolution modules and pooling operations, the spatial resolution is gradually reduced and the abstraction ability of semantic features is improved.
[0196] Each stage outputs intermediate features and uses UNet as a decoder for feature fusion;
[0197] The segmentation task is optimized using a combination of binary cross entropy (BCE) and Dice loss, where Dice loss can effectively handle the problem of class imbalance.
[0198] At the same time, tumor heatmap weighting is introduced in some areas, and a larger loss weight is applied to key areas to encourage the model to focus on the boundary details between the tumor and surrounding tissues. Let the predicted probability be p and the true value be t (the ignored label is masked), then for each channel c, define:
[0199]
[0200] Where Ω represents the effective pixel set (for some channels, Ω is multiplied by the importance heat map H), ω c Indicates the weight of the channel. In this application, the weight of the BBox channel is set to 0.5, and the rest are 1. The final combination is:
[0201]
[0202] The respective importance heat maps, i.e., tumor heat maps, are specifically as follows Figure 4 As shown:
[0203] 1. Center positioning:
[0204] The probability map M of the submucosa SM and adventitia AD output by the segmentation branch SM and M AD ,
[0205] Calculate the weighted center:
[0206]
[0207] This center position serves as the reference point for subsequent sector-shaped area division.
[0208] 2. Sector area division:
[0209] Taking the center point as the reference, the image is evenly divided into N sector-shaped areas, and the coverage angle of each area is
[0210] Number each sector area S0, S1, ..., S N-1 .
[0211] 3. Regional tumor area calculation and weight allocation:
[0212] For each sector area S k , calculate tumor area Among them L ij is the probability value of the corresponding pixel in the tumor segmentation map;
[0213] The preset threshold T is used to determine whether the tumor area in the region reaches a significant level, and the weight is defined:
[0214]
[0215] The weights of each region are used to construct a heatmap H of the overall image importance, and the key regions are weighted in the segmentation loss calculation, so that the model pays more attention to these regions during training.
[0216] In this preferred embodiment, ConvNeXt-Base is used as the backbone network. Multiple convolutional modules and pooling operations are used to extract multi-layer features, and the spatial resolution is gradually reduced to improve the abstraction of semantic features. The intermediate features output by each stage are then fused through the UNet decoder to generate a high-quality basic segmentation map. Next, a combination of binary cross entropy (BCE) and Dice loss is used for optimization. The Dice loss can effectively address class imbalance and improve segmentation accuracy. In addition, a tumor heatmap weighting mechanism is introduced to assign greater loss weights to key regions, encouraging the model to pay more attention to the boundary details between the tumor and surrounding tissue. Specifically, the image is divided into multiple sector-shaped regions by calculating the weighted center of the probability map of the submucosal layer and the adventitia. The tumor area within each region is calculated. A preset threshold is used to determine whether the tumor area within the region reaches a significant level, and weights are defined accordingly. These weights are used to construct a global importance heatmap and weight key regions in the segmentation loss calculation, thereby increasing the model's attention to these regions during training. Through this method, the model can more accurately identify and segment key structures in ultrasound endoscopic images, providing a more reliable basis for clinical diagnosis. It has significant clinical application value and broad application prospects.
[0217] The classification module 30 is used to input the cross-region feature maps into the classification branch of the preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes the multi-task classification of the endoscopic ultrasound image according to the preset lightweight classification network;
[0218] The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
[0219] As a preferred embodiment of the second embodiment, the inputting of each cross-region feature map into the classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network is specifically as follows:
[0220] The classification branch is mainly responsible for determining the T stage of the image after cross-region feature extraction. Its main design is as follows:
[0221] Multi-channel segmentation feature extraction and two-stage model design:
[0222] The segmentation branch outputs an 8-channel segmentation map to fully extract the intrinsic structural information of the ultrasound endoscopic image.
[0223] The classification branch makes staging judgments based on the structural information mined by the segmentation branch, which is more accurate and more interpretable than the single classification model that directly performs staging on ultrasound endoscopic images.
[0224] For different T-stage tasks, three classification heads are designed:
[0225] T1a / T1b binary classification head: used for early staging determination.
[0226] T2 / T3-4 binary classification head: used to distinguish between middle and late stages.
[0227] T1a-T4 five-category head: enables more fine-grained staging determination.
[0228] Each classification head adopts the MLP structure, and the specific formula is as follows:
[0229] MLP(x)=W2Dropout(ReLU(W1x+b1))+b2,
[0230] Among them, W1, W2, b1 and b2 are corresponding parameters.
[0231] For different tasks, cross entropy losses with different category weights are used to effectively balance the problem of uneven sample size.
[0232] As a preferred embodiment of the second embodiment, it also includes:
[0233] The classification loss used in this application is weighted cross-entropy loss. Since the dataset uses a balanced sampling strategy to ensure a balanced number of samples across categories in each minibatch, we set weights such as [4.0, 1.0] for the T1a / T1b tasks and [1.0, 1.5] for the T2 / T3-4 tasks. Standard cross-entropy loss is used for the T1a-T4 tasks.
[0234] The classification loss and segmentation loss are fused in the joint loss function through preset weight parameters, and jointly participate in back-propagation optimization, thereby realizing multi-task collaborative training.
[0235] To achieve collaborative optimization of segmentation and classification tasks, this application designs a joint loss function:
[0236]
[0237] in: It is the segmentation loss, which is composed of BCE and Dice loss, and is combined with the heat map H to perform key area weighting; and are the weighted cross entropy losses for each classification task; seg ,λ binary and λ multi is the weight coefficient of each loss item, and the optimal value is determined through experimental tuning.
[0238] During training, the AdamW optimizer was used, combined with a cosine annealing schedule to dynamically adjust the learning rate. The first five epochs served as a warm-up phase, using a linear ramp-up strategy. Subsequently, the learning rate was gradually reduced using a cosine function to ensure smooth convergence throughout the training process.
[0239] In this preferred embodiment, the present application organically combines the loss of the segmentation task (such as Dice loss and binary cross entropy loss) with the loss of the classification task (such as weighted cross entropy loss) by designing a joint loss function, so that the two branches can promote each other and optimize together during the training process. This collaborative optimization strategy not only improves the model's ability to extract key regional features, but also enhances the model's ability to discriminate tumor stages, ultimately achieving more precise image segmentation and more accurate tumor staging, providing strong support for clinical diagnosis and treatment.
[0240] As a preferred embodiment of the second embodiment, the present application conducted experiments on an EUS image dataset of 271 esophageal cancer patients using three-fold cross validation, and obtained the following main experimental results:
[0241] 1. Segmentation performance:
[0242] -Lesion segmentation: The Lesion Dice coefficient based on the ConvNeXt+UNet model was 87.01%, which was increased to 89.20% after the introduction of the joint segmentation and classification model, and then reached 90.32% through the tumor heat map mechanism.
[0243] -SM segmentation: The Dice coefficient of the base model is 92.72%, and the joint model is improved to 94.52%. Although there is a slight fluctuation after heat map enhancement, the overall performance remains above 93%.
[0244] -AD segmentation: increased from 92.22% to 94.56%, indicating that the cross-region features and heat map mechanism have a significant promoting effect on boundary region recognition.
[0245] 2. Classification performance:
[0246] -T1a / T1b binary classification: The best classification baseline had an accuracy of 73.12% and an AUC of 69.26%. The combined segmentation and classification model improved these to 75.95% and 79.79%, respectively. After introducing the heatmap mechanism, the accuracy reached 81.01% and the AUC was 82.45%.
[0247] -T2 / T3-4 binary classification: The joint model achieved an accuracy of 88.54% and an AUC exceeding 94%. The heatmap mechanism was slightly improved to 89.06% and 94.44%.
[0248] -T1a-T4 five-category classification: The overall accuracy increased from 69.45% of the baseline to 73.06% of the joint model. The heatmap mechanism increased the accuracy to 76.01%, and the AUC also improved accordingly.
[0249] 3. Evaluation indicators:
[0250] - Segmentation indicators: Dice coefficient, IoU, precision, recall and other indicators are used for evaluation;
[0251] -Classification indicators: using multi-dimensional indicators such as accuracy, sensitivity, specificity, AUC and confusion matrix;
[0252] - Cross-validation: The average performance, standard deviation, and 95% confidence interval of each fold are recorded in detail in the experimental report, which fully verifies the robustness and stability of the model.
[0253] The above experimental results show that the joint optimization strategy of fine segmentation and multi-task classification in this application is significantly better than the traditional single-task processing method, especially in the judgment of early stages (T1a / T1b), the model's capture effect of subtle boundary features is significantly improved.
[0254] This device uses three modules to divide the work and coordinate work to more accurately process ultrasound images. This application uses the segmentation branch of the preset esophageal cancer ultrasound endoscopic staging diagnosis model to process ultrasound endoscopic images, which can output high-quality basic segmentation maps. These segmentation maps provide an accurate basis for subsequent cross-region feature extraction. Then, through a preset mathematical operation method, cross-region feature maps are obtained from the basic segmentation map. These feature maps can highlight the relationship between the tumor and the various layers of the esophageal wall structure, providing richer information for tumor staging. In addition, the segmentation branch dynamically adjusts the segmentation weight of the loss function according to the importance heat map during training. This innovative mechanism enables the model to pay more attention to key areas, thereby improving the accuracy of segmentation. Finally, the cross-region feature map is input into the classification branch to complete multi-task classification, including key staging tasks such as T1a / T1b, T2 / T3-4 and T1a-T4. This multi-task learning method not only improves the accuracy of classification, but also enhances the generalization ability of the model. This application effectively solves the problem that the existing technology cannot accurately obtain effective medical information based on ultrasound images.
[0255] Example 3:
[0256] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the ultrasonic image processing method;
[0257] Wherein, the ultrasonic image processing method, if implemented in the form of a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0258] Example 4
[0259] The present application provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, any one of the ultrasound image processing methods described in Example 1 is implemented.
[0260] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for processing an ultrasonic image, characterized in that: include: Obtain endoscopic ultrasound images; Inputting the ultrasound endoscopic image into a segmentation branch of a preset ultrasound endoscopic staging diagnostic model for esophageal cancer, so that the segmentation branch outputs each basic segmentation map of the ultrasound endoscopic image and then outputs each cross-region feature map of the ultrasound endoscopic image according to a preset mathematical operation method and each basic segmentation map; The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process; Inputting each of the cross-region feature maps into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network; The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
2. The ultrasonic image processing method according to claim 1, wherein: The method of outputting characteristic maps of each cross region of the ultrasound endoscopy image according to the preset mathematical operation method and each basic segmentation map is specifically as follows: The segmentation branch performs an element-by-element product operation on the multiple segmentation results in each basic segmentation map, and outputs each intersection region feature map; The multiple segmentation results include a tumor region segmentation map, a submucosal layer segmentation map, an adventitia segmentation map, and various cropping frames.
3. The ultrasonic image processing method according to claim 1, wherein: The segmentation branch is trained on the initial segmentation branch based on historical ultrasound endoscopic images, and the segmentation weights of the loss function are dynamically adjusted according to the preset importance heat maps during the training process. Specifically, Obtain historical endoscopic ultrasound images; Inputting the historical ultrasound endoscopy image into the initial segmentation branch for training, so that the initial segmentation branch outputs each initial basic segmentation image; Calculating the center point of the esophageal structure according to the probability distribution of the submucosal segmentation map and the adventitia segmentation map of each initial basic segmentation map; Dividing each of the initial basic segmentation images into N sector-shaped regions according to the center point of the esophageal structure; wherein N is a positive integer greater than 1; Calculating an importance heat map of each initial basic segmentation map according to the fan-shaped region and the tumor region segmentation map of each initial basic segmentation map; Calculate each weight based on the importance heat map and a preset threshold; According to the weights, the loss function of the initial segmentation branch is dynamically adjusted. If the value of the loss function no longer decreases within a preset time period, the training is stopped to obtain the segmentation branch.
4. The ultrasonic image processing method according to claim 3, characterized in that: The loss function is specifically: The loss function is a combination of binary cross entropy loss and Dice loss; The formula of the loss function is: Where, ω c represents the weight of the channel, represents the binary cross entropy loss, Denotes Dice loss.
5. The ultrasonic image processing method according to any one of claims 1 to 4, characterized in that: Also includes: The preset esophageal cancer endoscopic ultrasound staging diagnosis model fuses the loss of the segmentation branch with the loss of the classification branch through a joint loss function and performs collaborative optimization; The joint loss function is: Where, is the loss function of the segmentation branch, which consists of binary cross entropy loss and Dice loss; and are the weighted cross entropy losses in the classification tasks respectively; seg ,λ binary and λ multi is the weight coefficient of each loss item.
6. An ultrasonic image processing device, characterized in that: Includes acquisition module, segmentation module and classification module; The acquisition module is used to obtain an ultrasound endoscopic image; The segmentation module is used to input the ultrasonic endoscopic image into the segmentation branch of the preset ultrasonic endoscopic staging diagnosis model for esophageal cancer, so that after the segmentation branch outputs each basic segmentation map of the ultrasonic endoscopic image, it outputs each cross-region feature map of the ultrasonic endoscopic image according to a preset mathematical operation method and each basic segmentation map; The segmentation branch is trained on the initial segmentation branch according to the historical ultrasound endoscopy images, and the segmentation weight of the loss function is dynamically adjusted according to the preset importance heat maps during the training process; The classification module is used to input the respective cross-region feature maps into a classification branch of a preset esophageal cancer endoscopic ultrasound staging diagnosis model, so that the classification branch completes multi-task classification of the endoscopic ultrasound image according to a preset lightweight classification network; The multi-task classification includes T1a / T1b binary classification, T2 / T3-4 binary classification, and T1a-T4 five-classification tasks.
7. The ultrasonic image processing device according to claim 6, characterized in that: The method of outputting characteristic maps of each cross region of the ultrasound endoscopy image according to the preset mathematical operation method and each basic segmentation map is specifically as follows: The segmentation branch performs an element-by-element product operation on the multiple segmentation results in each basic segmentation map, and outputs each intersection region feature map; The multiple segmentation results include a tumor region segmentation map, a submucosal layer segmentation map, an adventitia segmentation map, and various cropping frames.
8. The ultrasonic image processing device according to claim 6, characterized in that: The segmentation branch is trained on the initial segmentation branch based on historical ultrasound endoscopic images, and the segmentation weights of the loss function are dynamically adjusted according to the preset importance heat maps during the training process. Specifically, Obtain historical endoscopic ultrasound images; Inputting the historical ultrasound endoscopy image into the initial segmentation branch for training, so that the initial segmentation branch outputs each initial basic segmentation image; Calculating the center point of the esophageal structure according to the probability distribution of the submucosal segmentation map and the adventitia segmentation map of each initial basic segmentation map; Dividing each of the initial basic segmentation images into N sector-shaped regions according to the center point of the esophageal structure; wherein N is a positive integer greater than 1; Calculating an importance heat map of each initial basic segmentation map according to the fan-shaped region and the tumor region segmentation map of each initial basic segmentation map; Calculate each weight based on the importance heat map and a preset threshold; According to the weights, the loss function of the initial segmentation branch is dynamically adjusted. If the value of the loss function no longer decreases within a preset time period, the training is stopped to obtain the segmentation branch.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the ultrasound image processing method according to any one of claims 1 to 5.
10. A terminal device, characterized in that: The ultrasonic image processing method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the ultrasonic image processing method according to any one of claims 1 to 5 is implemented.