Class activation graph generation model training method, generation method and related equipment

Through multi-prototype parametric modeling and medical attention guidance technology, the problem of overlapping class activation maps in multi-lesion skin ultrasound images is solved, the interpretability and accuracy of class activation maps are improved, and the interpretation effect of medical images is enhanced.

CN120672750AActive Publication Date: 2025-09-19脉得智能科技(无锡)有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511164391.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-09-19
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

When there are multiple lesions in skin ultrasound images, existing class activation map technologies often overlap between different class activation maps, resulting in distorted interpretation.

Method used

Multi-prototype parametric modeling is adopted. By assigning multiple learnable prototype vectors to each category, combining the medical attention guidance module and the channel attention unit, the skin structure mask is used to guide the spatial attention of the feature map, feature fusion is performed, and temporal consistency loss is introduced to optimize the class activation map generation model.

Benefits of technology

It improves the category specificity and spatial precision of class activation maps, enhances the accuracy and clinical credibility of medical image interpretation, reduces the overlap of class activation maps, and improves the clarity of image interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672750A_ABST
    Figure CN120672750A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a class activation graph generation model training method, a class activation graph generation method and related equipment, and relates to the technical field of image processing. A plurality of groups of sample data are acquired, a plurality of prototype vectors corresponding to each labeling category are initialized, a skin feature map and a corresponding skin structure mask are input into a medical attention guiding module, attention mechanism processing is performed to obtain a fusion feature map after attention fusion, and the fusion feature map corresponding to a key frame is input into a full-connection classification layer to obtain a full-connection classification layer; the method comprises the steps of obtaining a prediction category of a key frame, inputting a fusion feature graph and each prototype vector into a class activation graph generation module, carrying out weighted fusion to obtain a class activation graph of each annotation category, and calculating joint training loss of each prediction category, each annotation category, each focus mask and each class activation graph so as to update a class activation graph generation model. Through a multi-prototype activation mechanism, overlapping among different types of activation graphs can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a class activation map generation model training method, a generation method and related equipment. Background Art

[0002] The incidence of skin diseases continues to rise, becoming a major global health issue. Traditional pathological examination, the gold standard for diagnosis, faces challenges such as high invasiveness, high costs, and high physician subjectivity.

[0003] In recent years, the combination of non-invasive skin ultrasound imaging and artificial intelligence-assisted diagnosis has become a research hotspot. In medical image interpretability research, class activation map technology has attracted widespread attention due to its ability to intuitively express the model's focus areas.

[0004] However, in current class activation map technology, when there are multiple lesions in ultrasound images, different class activation maps often overlap, resulting in distorted interpretation. Summary of the Invention

[0005] In view of this, an object of an embodiment of the present invention is to provide a class activation map generation model training method, generation method and related devices to at least partially improve the above-mentioned problems.

[0006] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, an embodiment of the present invention provides a method for training a class activation map generation model, comprising: Acquire multiple sets of sample data; wherein each set of the sample data includes a skin ultrasound video and a labeled category of the skin ultrasound video, the skin ultrasound video includes multiple consecutive skin ultrasound images, the skin ultrasound images include key frames and non-key frames, and each set of the sample data also includes a skin structure mask of each skin ultrasound image and a lesion mask of the key frame; Initializing a plurality of prototype vectors corresponding to each of the labeled categories according to the labeled categories of all the sample data and preset prototype hyperparameters; wherein the prototype hyperparameters are used to specify the number of prototype vectors corresponding to each of the labeled categories, and the prototype vectors represent different morphological subtype features under the same labeled category; For each skin ultrasound image in each set of sample data, the skin feature map and the corresponding skin structure mask are input into a medical attention guidance module, and processed by an attention mechanism to obtain a fused feature map after fusion of attention; wherein the skin feature map is extracted from the skin ultrasound image; Inputting the fused feature map corresponding to the key frame into a fully connected classification layer to obtain a predicted category of the key frame; Inputting the fused feature map and each of the prototype vectors into a class activation map generation module for weighted fusion to obtain a class activation map of each of the labeled categories; A joint training loss is calculated for each of the predicted categories, each of the labeled categories, each of the lesion masks, and each of the class activation maps, and the class activation map generation model is updated according to the joint training loss.

[0007] Optionally, the step of inputting the skin feature map and the corresponding skin structure mask into a medical attention guidance module, and processing them through an attention mechanism to obtain a fused feature map after fusion of attention, includes: Inputting the skin structure mask and the skin feature map into a spatial attention unit, and performing spatial attention processing to obtain a spatial attention map; Inputting the skin feature map into a channel attention unit and performing channel attention processing to obtain a channel attention map; Perform dot multiplication on the skin feature map, the spatial attention map, and the channel attention map to obtain a fused feature map after fusion of attention.

[0008] Optionally, the spatial attention map is calculated by the following formula:

[0009] in, for the skin structure mask, is the skin feature map, is a tunable hyperparameter, for convolution, is the activation function; The channel attention map is calculated by the following formula:

[0010] in, is the skin feature map, is the global average pooling, is the maximum pooling, and for convolution, is the ReLU activation.

[0011] Optionally, the step of inputting the fused feature map and each prototype vector into a class activation map generation module for weighted fusion to obtain a class activation map of each labeled category includes: Calculate the inner product matching between each feature point of the fused feature map and each prototype vector to obtain a prototype activation map of each prototype vector; Inputting the fused feature map into a weight prediction unit to obtain a weight map of each prototype vector; For each of the labeled categories, the prototype activation maps and the weight maps are weighted and summed to obtain a class activation map of the labeled category.

[0012] Optionally, the prototype activation map is calculated by the following formula:

[0013] in, is the j-th prototype activation map of the labeled category k, is the fusion feature map, is the j-th prototype vector of the labeled category k, express and The inner product of The weight map is calculated by the following formula:

[0014] in, is the j-th prototype activation map of the labeled category k, is the fusion feature map, is the activation function, for convolution; The class activation map is calculated by the following formula:

[0015] in, is the class activation map of the labeled category k, and K is the number of prototype vectors of the labeled category k.

[0016] Optionally, the calculation formula of the joint training loss is:

[0017]

[0018]

[0019]

[0020]

[0021] in, is the cross entropy loss, is the probability of predicting the labeled category i, is the reciprocal of the proportion of samples labeled category i, is the auxiliary adjustment factor; is the IoU similarity loss, is the class activation graph, mask for the lesion; is the temporal consistency loss, are the class activation maps of two consecutive skin ultrasound images in the skin ultrasound video, is the L2 norm; is the multi-lesion decoupling regularization term, express and For different annotation types, are class activation maps of different annotation types for the same skin ultrasound image, is the Hadamard product, is the L1 norm; , , , is the weight of each loss.

[0022] In a second aspect, an embodiment of the present invention provides a method for generating a class activation map, comprising: Obtain key frames in the skin ultrasound video to be classified; Inputting the key frame into a pre-trained feature extraction model to obtain key frame features; Inputting the key frame features into a pre-trained hierarchical model to obtain a skin structure mask; The key frame features and the skin structure mask are input into a class activation map generation model to obtain class activation maps of each category of the key frame; the class activation map generation model is trained by the class activation map generation model training method described in the first aspect.

[0023] In a third aspect, an embodiment of the present invention provides a class activation map generation model training device, comprising: A sample data acquisition unit, configured to acquire multiple sets of sample data; wherein each set of sample data includes a skin ultrasound video and a labeled category of the skin ultrasound video, the skin ultrasound video includes multiple consecutive skin ultrasound images, the skin ultrasound images include key frames and non-key frames, and each set of sample data also includes a skin structure mask of each skin ultrasound image and a lesion mask of the key frame; A prototype vector initialization unit is used to initialize multiple prototype vectors corresponding to each of the labeled categories according to the labeled categories of all the sample data and preset prototype hyperparameters; wherein the prototype hyperparameters are used to specify the number of prototype vectors corresponding to each of the labeled categories, and the prototype vectors represent different morphological subtype features under the same labeled category; a feature fusion unit, configured to input, for each skin ultrasound image in each set of sample data, a skin feature map and a corresponding skin structure mask into a medical attention guidance module, and process the skin feature map and the corresponding skin structure mask through an attention mechanism to obtain a fused feature map after fusion of attention; wherein the skin feature map is extracted from the skin ultrasound image; A category prediction unit, configured to input the fused feature map corresponding to the key frame into a fully connected classification layer to obtain a predicted category of the key frame; A class activation map prediction unit is used to input the fused feature map and each prototype vector into a class activation map generation module for weighted fusion to obtain a class activation map of each labeled category; The loss calculation unit is used to calculate the joint training loss of each of the predicted categories, each of the labeled categories, each of the lesion masks, and each of the class activation maps, and update the class activation map generation model according to the joint training loss.

[0024] In a fourth aspect, an embodiment of the present invention provides a class activation map generating device, comprising: A key frame acquisition unit, used to acquire key frames in the skin ultrasound video to be classified; A key frame feature extraction unit, configured to input the key frame into a pre-trained feature extraction model to obtain key frame features; a skin structure extraction unit, configured to input the key frame features into a pre-trained hierarchical model to obtain a skin structure mask; A class activation map generation unit is used to input the key frame features and the skin structure mask into a class activation map generation model to obtain class activation maps of each category of the key frame; the class activation map generation model is trained by the class activation map generation model training method described in the first aspect.

[0025] In a fifth aspect, an embodiment of the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the class activation map generation model training method described in any one of the first aspects above is implemented.

[0026] The embodiments of the present invention provide a class activation map generation model training method, generation method, and related equipment. Through a multi-prototype parameterized activation mechanism, multiple learnable prototype vectors are assigned to each category, thereby capturing the diversity within the category and more finely matching the local features related to the category in the image. The generated class activation map can reduce the overlapping parts and improve the interpretability of the class activation.

[0027] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0029] Figure 1 A schematic structural block diagram of an electronic device provided by an embodiment of the present invention; Figure 2 One of the flow diagrams of a class activation map generation model training method provided by an embodiment of the present invention; Figure 3 One of the schematic structural block diagrams of a class activation map generation model provided by an embodiment of the present invention; Figure 4 A second schematic structural diagram of a class activation map generation model provided by an embodiment of the present invention; Figure 5 A second flow chart of a method for training a class activation map generation model provided by an embodiment of the present invention; Figure 6 A third flow chart of a method for training a class activation map generation model provided by an embodiment of the present invention; Figure 7 A third schematic structural diagram of a class activation map generation model provided by an embodiment of the present invention; Figure 8 A training flow chart of a class activation map generation model provided by an embodiment of the present invention; Figure 9 A schematic diagram of a process for generating a class activation map provided by an embodiment of the present invention; Figure 10 A schematic structural block diagram of a class activation map generation model training device provided by an embodiment of the present invention; Figure 11 A schematic structural block diagram of a class activation map generation device provided by an embodiment of the present invention.

[0030] Icons: 100-electronic device; 101-memory; 102-communication interface; 103-processor; 104-communication bus; 300-class activation map generation model; 310-medical attention guidance module; 311-spatial attention unit; 312-channel attention processing; 320-class activation map generation module; 321-weight prediction unit; 330-fully connected classification layer; 500-class activation map generation model training device; 510-sample data acquisition unit; 520-prototype vector initialization unit; 530-feature fusion unit; 540-category prediction unit; 550-class activation map prediction unit; 560-loss calculation unit; 600-class activation map generation device; 610-key frame acquisition unit; 620-key frame feature extraction unit; 630-skin structure extraction unit; 640-class activation map generation unit. DETAILED DESCRIPTION

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0032] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0033] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0034] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0035] As described in the background technology, in traditional class activation map technology, class activation map generation relies on the weighted sum of classifier weights and feature maps (usually based on GAP). This method ignores morphological differences within the category (such as different subtypes of melanoma), resulting in poor performance of the generated class activation map when there are multiple lesions or large intra-category variations.

[0036] Based on the above situation, the embodiments of the present invention provide a class activation map generation model training method, generation method and related equipment. Through multi-prototype parametric modeling, the category specificity, spatial accuracy and multi-lesion decoupling capability of the class activation map are improved, and the accuracy and clinical credibility of medical image interpretation are enhanced.

[0037] To implement the process steps and functions of each example of the present invention, please refer to Figure 1 , Figure 1 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103. The memory 101 and processor 103 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101, thereby performing various functional applications and data processing.

[0038] The electronic device 100 may be, but is not limited to, a personal computer (PC), a server, a distributed computer, or the like. It is understood that the electronic device 100 is not limited to a physical server and may also be a virtual machine on a physical server, a virtual machine built on a cloud platform, or other computer that provides the same functionality as the server or virtual machine. The operating system of the electronic device 100 may be, but is not limited to, Windows, Linux, or the like.

[0039] The memory 101 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0040] The communication connection between the electronic device 100 and an external device is achieved through at least one communication interface 102 (which can be wired or wireless).

[0041] Processor 103 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the embodiments of the present invention may be completed by hardware integrated logic circuits in processor 103 or by software instructions. Processor 103 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0042] I understand. Figure 1 The structure shown is for illustration only. The electronic device 100 may further include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0043] The following is an exemplary description of the class activation map generation model training method provided by the present invention. Figure 2 , the execution subject of this method can be the above Figure 1 The electronic device 100 shown in FIG. 1 includes the following steps: Figure 2 The following steps are described: S210: Acquire multiple groups of sample data.

[0044] Among them, each set of sample data includes skin ultrasound video and the labeled category of skin ultrasound video. Skin ultrasound video includes multiple continuous skin ultrasound images. Skin ultrasound images include key frames and non-key frames. Each set of sample data also includes the skin structure mask of each skin ultrasound image and the lesion mask of the key frame.

[0045] S220: Initialize multiple prototype vectors corresponding to each labeled category based on the labeled categories of all sample data and preset prototype hyperparameters.

[0046] Among them, the prototype hyperparameter is used to specify the number of prototype vectors corresponding to each annotation category, and the prototype vector represents different morphological subtype characteristics under the same annotation category.

[0047] S230: For each skin ultrasound image in each set of sample data, the skin feature map and the corresponding skin structure mask are input into the medical attention guidance module, and processed by the attention mechanism to obtain a fused feature map after fusion of attention.

[0048] The skin feature map is extracted from the skin ultrasound image.

[0049] S240: Input the fused feature map corresponding to the key frame into the fully connected classification layer to obtain the predicted category of the key frame.

[0050] S250: Input the fused feature map and each prototype vector into the class activation map generation module for weighted fusion to obtain the class activation map of each labeled category.

[0051] S260: Calculate the joint training loss for each predicted category, each labeled category, each lesion mask, and each type of activation map, and update the class activation map generation model according to the joint training loss.

[0052] To train the class activation map generation model, first prepare multiple sets of sample data and initialize multiple prototype vectors.

[0053] Each set of sample data is a skin ultrasound video including multiple consecutive skin ultrasound images, wherein one skin ultrasound image is a key frame, for example, a frame with a blood flow peak is used as a key frame.

[0054] In addition, each skin ultrasound image has a corresponding skin structure mask, representing the anatomical regions of the epidermis, dermis, and subcutaneous tissue. This skin structure mask provides prior guidance for subsequent feature extraction from the skin ultrasound image. The skin structure mask in the sample data can be obtained by a doctor's annotation of the skin ultrasound image or by segmenting the skin ultrasound image using a pre-trained hierarchical model.

[0055] In addition, each skin ultrasound video is also annotated with a label category, for example, melanoma, hemangioma, basal cell carcinoma, and a lesion mask of the key frame, which is obtained by manual labeling by a doctor.

[0056] At this point, the sample data for training is ready. Next, initialize multiple prototype vectors, which are used for intra-class variation modeling, for example, different morphological subtypes of melanoma. The initialized prototype vectors are obtained based on the sample data and the prototype hyperparameters. In the sample data, how many labeled categories are marked, that is, how many categories are there in the classification of the skin ultrasound video. For example, there are three types of skin ultrasound classification: melanoma, hemangioma, and basal cell carcinoma (of course, there are more than three in reality, this is just an example). The prototype hyperparameter is how many prototypes there are in each category. This is predefined. The number of prototypes for each category can be different or the same. For the sake of simplicity, the embodiment of the present invention is explained as if the number of prototypes for each category is the same. For example, each category has 3 prototype vectors. Then 9 prototype vectors will be initialized. . Later in the training process, Update and optimize until the training is completed.

[0057] After preparing the sample data and initializing the prototype vector, the training process of the class activation map generation model begins. Figure 3 The class activation map generation model 300 includes a medical attention guidance module 310, a class activation map generation module 320, and a fully connected classification layer 330. After forward propagation through the medical attention guidance module 310 and the class activation map generation module 320 to calculate the joint training loss, backpropagation is then used to update the parameters of the medical attention guidance module 310 and the class activation map generation module 320.

[0058] Each skin ultrasound image in each set of sample data is processed to generate a class activation map for each labeled category. Loss calculation and backpropagation update are then performed.

[0059] A skin feature map is extracted from the ultrasound image using a well-trained feature extraction model, such as a MixConv neural network or a ResNet neural network. The skin feature map and the corresponding skin structure mask are then input into the medical attention guidance module. After processing with the attention mechanism, the different structural layers of the epidermis, dermis, and subcutaneous tissue are highlighted, resulting in a fused feature map.

[0060] Since the sample data only classifies the skin ultrasound video as a whole, we only need to classify the key frames in the skin ultrasound video and perform subsequent loss calculations based on the overall classification. It is understandable that each ultrasound image in the skin ultrasound video can also be classified and labeled, and all frames in the skin ultrasound video can be classified during processing. In this embodiment of the present invention, the fused feature map corresponding to the key frame is input into the fully connected classification layer to obtain the predicted category of the key frame.

[0061] The class activation map generation module processes the fused feature map and each prototype vector to generate the class activation map of each labeled category.

[0062] Finally, the joint training loss is calculated for each predicted category, each labeled category, each lesion mask, and each type of activation map, and the class activation map generation model is updated based on the joint training loss.

[0063] This method assigns multiple learnable prototype vectors to each category through a multi-prototype parameterized activation mechanism, thereby capturing the diversity within the category, more finely matching the category-related local features in the image, reducing the overlap in the generated class activation map, and improving the interpretability of the class activation.

[0064] In order to capture richer information in the skin feature map, spatial attention processing and channel attention processing can be performed on the skin feature map. Figure 4 , the medical attention guidance module 310 may include a spatial attention unit 311 and a channel attention processing 312. Figure 5 , the above step S230 may include the following steps: S231: Input the skin structure mask and the skin feature map into the spatial attention unit, and obtain a spatial attention map after spatial attention processing.

[0065] S232: Input the skin feature map into the channel attention unit, and obtain a channel attention map after channel attention processing.

[0066] S233: Perform dot multiplication on the skin feature map, the spatial attention map, and the channel attention map to obtain a fused feature map after fusion of attention.

[0067] There are many ways to perform spatial attention processing on the skin feature map using the skin structure mask as prior information. In one optional implementation, the processing process can be expressed by the following formula:

[0068] in, For skin structure mask, is the skin feature map, is a tunable hyperparameter, for convolution, is the activation function.

[0069] By substituting the skin structure mask and the skin feature map into the formula, we can obtain the spatial attention map.

[0070] The channel attention processing of the skin feature map adopts the CBAM channel attention mechanism. For example, the channel attention map is calculated by the following formula:

[0071] in, is the skin feature map, is the global average pooling, is the maximum pooling, and for convolution, is the ReLU activation.

[0072] After the spatial attention map and the channel attention map, the skin feature map can be fused. The fusion process is as follows:

[0073] in, is the fusion feature map, is the skin feature map, is the spatial attention map, is the channel attention map.

[0074] For the fusion of prototype vector and fusion feature map, please refer to Figure 6 , the above step S250 may include the following steps: S251: Calculate the inner product matching between each feature point of the fusion feature map and each prototype vector to obtain the prototype activation map of each prototype vector.

[0075] S252: Input the fused feature map into the weight prediction unit to obtain the weight map of each prototype vector.

[0076] S253: For each labeled category, perform weighted summation of the prototype activation maps and the weight maps to obtain a class activation map of the labeled category.

[0077] For example, the embodiment of the present invention includes three categories, namely melanoma, hemangioma, and basal cell carcinoma. Each category has three prototype vectors. We illustrate the generation of a class activation map for a category. For example, melanoma has three prototype vectors. 、 、 .

[0078] First calculate the fusion feature map Each feature point and 、 、 The inner product matching of 、 、 The prototype activation map is calculated by the following formula:

[0079] in, is the activation map of the j-th prototype of the labeled category k, is the fusion feature map, is the j-th prototype vector of the labeled category k, express and The inner product of .

[0080] For example, the three prototype activation maps of melanoma are calculated as 、 、 .

[0081] At the same time, the weight prediction unit predicts the weight map of the prototype vector, see Figure 7 , the class activation map generation module 320 includes a weight prediction unit 321. The fusion feature map The weight map of each prototype vector is predicted by the input weight prediction unit 321. The weight prediction unit 321 can use a light convolution. For example, the weight map can be calculated by the following formula:

[0082] in, is the activation map of the j-th prototype of the labeled category k, is the fusion feature map, is the activation function, for convolution.

[0083] The parameters of the weight prediction unit 321 will also be updated during the training process.

[0084] For example, the weight graph of the three prototype vectors predicting melanoma is 、 、 .

[0085] Finally, the dot product of each prototype activation map and each weight map of each category is summed to obtain the class activation map of each category. The class activation map can be calculated by the following formula:

[0086] in, is the class activation map of the labeled category k, and K is the number of prototype vectors of the labeled category k.

[0087] For example, the class activation map for melanoma is .

[0088] The calculation formula of the joint training loss in step S260 can be:

[0089]

[0090]

[0091]

[0092]

[0093] in, is the cross entropy loss, is the probability of predicting the labeled category i, is the reciprocal of the proportion of samples labeled category i, is the auxiliary adjustment factor; is the IoU similarity loss, is the class activation graph, Mask for lesions; is the temporal consistency loss, are the class activation maps of two consecutive skin ultrasound images in the skin ultrasound video, is the L2 norm; is the multi-lesion decoupling regularization term, express and For different annotation types, They are the class activation maps of different annotation types for the same skin ultrasound image, is the Hadamard product, is the L1 norm; , , , is the weight of each loss.

[0094] Continuing with the above example to illustrate the joint training loss, if the skin ultrasound videos labeled as melanoma, hemangioma, and basal cell carcinoma account for 20%, 20%, and 40% of the sample data respectively, then They are 、 、 、 The percentage of keyframes in this batch of sample data that were classified as melanoma, hemangioma, or basal cell carcinoma. This data is used to calculate the cross-entropy loss. The IoU similarity loss is calculated by comparing the activation maps of the keyframes with the annotated lesion masks. Therefore, the cross-entropy loss and IoU similarity loss use the keyframe results.

[0095] The temporal consistency loss is designed to prevent the class activation map from jumping over time by introducing a temporal smoothing loss. The temporal consistency loss is calculated using two consecutive frames of the same skin ultrasound video.

[0096] The multi-lesion decoupling regularization term is calculated using the activation maps corresponding to the same skin ultrasound image. For example, the activation maps of a skin ultrasound image include the activation maps of melanoma, hemangioma, and basal cell carcinoma. The sum of the L2 norms between each of these activation maps is calculated.

[0097] See also Figure 8 , and further explain the training method of the class activation map generation model from the perspective of forward propagation and back propagation. First, initialize the prototype vector according to the labeled category and prototype hyperparameters, input the skin feature map and skin structure mask extracted from the skin ultrasound image in the sample into the medical attention guidance module to obtain the feature fusion map, input the prototype vector and the feature fusion map into the class activation map generation module to generate the class activation map, and input the feature fusion map into the fully connected classification layer to obtain the key frame prediction category. Calculate the joint training loss based on the key frame prediction category, class activation map, labeled category, and lesion mask. Finally, perform directional propagation based on the joint training loss to update the parameters of the fully connected classification layer, class activation map generation module, medical attention guidance module, and prototype vector.

[0098] Furthermore, the embodiment of the present invention also provides a method for generating a class activation map, see Figure 9 , the activation map generation method includes the following steps: S410: Acquire key frames in the skin ultrasound video to be classified.

[0099] S420: Input the key frame into a pre-trained feature extraction model to obtain key frame features.

[0100] S430: Input the key frame features into the pre-trained hierarchical model to obtain a skin structure mask.

[0101] S440: Inputting the key frame features and the skin structure mask into the class activation map generation model to obtain the class activation map of each category of the key frame; the class activation map generation model is trained by the above-mentioned class activation map generation model training method.

[0102] The skin is scanned by ultrasound equipment to obtain an ultrasound video of the skin to be classified, and the key frames are extracted. The features of the key frames are extracted using a pre-trained feature extraction model to obtain key frame features. The key frame features are input into a pre-trained hierarchical model to obtain a skin structure mask. Finally, the key frame features and the skin structure mask are input into a class activation map generation model to obtain class activation maps of each category of the key frames.

[0103] Furthermore, the embodiment of the present invention also provides a class activation map generation model training device, see Figure 10 , the activation map generation model training device 500 includes: The sample data acquisition unit 510 is used to acquire multiple groups of sample data; wherein each group of sample data includes a skin ultrasound video and a labeled category of the skin ultrasound video, the skin ultrasound video includes multiple continuous skin ultrasound images, the skin ultrasound images include key frames and non-key frames, and each group of sample data also includes a skin structure mask of each skin ultrasound image and a lesion mask of the key frame.

[0104] The prototype vector initialization unit 520 is used to initialize multiple prototype vectors corresponding to each annotation category based on the annotation categories of all sample data and preset prototype hyperparameters; wherein the prototype hyperparameters are used to specify the number of prototype vectors corresponding to each annotation category, and the prototype vectors represent different morphological subtype features under the same annotation category.

[0105] The feature fusion unit 530 is used to input the skin feature map and the corresponding skin structure mask into the medical attention guidance module for each skin ultrasound image in each set of sample data, and obtain a fused feature map after fusion of attention through the attention mechanism; wherein the skin feature map is extracted from the skin ultrasound image.

[0106] The category prediction unit 540 is used to input the fused feature map corresponding to the key frame into the fully connected classification layer to obtain the predicted category of the key frame.

[0107] The class activation map prediction unit 550 is used to input the fused feature map and each prototype vector into the class activation map generation module for weighted fusion to obtain the class activation map of each labeled category.

[0108] The loss calculation unit 560 is used to calculate the joint training loss for each predicted category, each labeled category, each lesion mask, and each activation map, and update the class activation map generation model according to the joint training loss.

[0109] Furthermore, the embodiment of the present invention also provides a class activation map generation device, see Figure 11 , the activation map generating device 600 includes: The key frame acquisition unit 610 is configured to acquire key frames in the skin ultrasound video to be classified.

[0110] The key frame feature extraction unit 620 is used to input the key frame into the pre-trained feature extraction model to obtain the key frame features.

[0111] The skin structure extraction unit 630 is used to input the key frame features into the pre-trained hierarchical model to obtain a skin structure mask.

[0112] The class activation map generation unit 640 is used to input the key frame features and the skin structure mask into the class activation map generation model to obtain the class activation map of each category of the key frame; the class activation map generation model is trained by the above-mentioned class activation map generation model training method.

[0113] In summary, the embodiments of the present invention provide a class activation map generation model training method, generation method and related equipment. By introducing a multi-prototype parameterized modeling mechanism, multiple learnable prototype vectors are assigned to each category, which can better capture the diversity within the same category, improve the matching accuracy between the class activation map and the local area of ​​the image, and effectively reduce the overlapping area between multiple class activation maps; introduce a medical attention guidance module, including a spatial attention unit and a channel attention unit, and use the skin structure mask to guide the spatial attention of the feature map to improve the focus on the skin structure area. Combined with medical prior knowledge (skin structure mask) for feature fusion, the class activation map can be more consistent with the anatomical structure of the medical image; use temporal consistency loss to constrain the time dimension of the class activation map in the video sequence to prevent the class activation map from jumping drastically over time.

[0114] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.

[0115] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0116] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0117] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0118] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A class activation map generation model training method, characterized in that: include: Acquire multiple sets of sample data; wherein each set of the sample data includes a skin ultrasound video and a labeled category of the skin ultrasound video, the skin ultrasound video includes multiple consecutive skin ultrasound images, the skin ultrasound images include key frames and non-key frames, and each set of the sample data also includes a skin structure mask of each skin ultrasound image and a lesion mask of the key frame; Initializing a plurality of prototype vectors corresponding to each of the labeled categories according to the labeled categories of all the sample data and preset prototype hyperparameters; wherein the prototype hyperparameters are used to specify the number of prototype vectors corresponding to each of the labeled categories, and the prototype vectors represent different morphological subtype features under the same labeled category; For each skin ultrasound image in each set of sample data, the skin feature map and the corresponding skin structure mask are input into a medical attention guidance module, and processed by an attention mechanism to obtain a fused feature map after fusion of attention; wherein the skin feature map is extracted from the skin ultrasound image; Inputting the fused feature map corresponding to the key frame into a fully connected classification layer to obtain a predicted category of the key frame; Inputting the fused feature map and each of the prototype vectors into a class activation map generation module for weighted fusion to obtain a class activation map of each of the labeled categories; A joint training loss is calculated for each of the predicted categories, each of the labeled categories, each of the lesion masks, and each of the class activation maps, and the class activation map generation model is updated according to the joint training loss.

2. The class activation map generation model training method according to claim 1, characterized in that The skin feature map and the corresponding skin structure mask are input into the medical attention guidance module, and processed by the attention mechanism to obtain a fused feature map after fusion of attention, including: Inputting the skin structure mask and the skin feature map into a spatial attention unit, and performing spatial attention processing to obtain a spatial attention map; Inputting the skin feature map into a channel attention unit and performing channel attention processing to obtain a channel attention map; Perform dot multiplication on the skin feature map, the spatial attention map, and the channel attention map to obtain a fused feature map after fusion of attention.

3. The class activation map generation model training method according to claim 2, characterized in that The spatial attention map is calculated by the following formula: in, for the skin structure mask, is the skin feature map, is a tunable hyperparameter, for convolution, is the activation function; The channel attention map is calculated by the following formula: in, is the skin feature map, is the global average pooling, is the maximum pooling, and for convolution, is the ReLU activation.

4. The class activation map generation model training method according to claim 1, wherein The step of inputting the fused feature map and each prototype vector into a class activation map generation module for weighted fusion to obtain a class activation map of each labeled category includes: Calculate the inner product matching between each feature point of the fused feature map and each prototype vector to obtain a prototype activation map of each prototype vector; Inputting the fused feature map into a weight prediction unit to obtain a weight map of each prototype vector; For each of the labeled categories, the prototype activation maps and the weight maps are weighted and summed to obtain a class activation map of the labeled category.

5. The class activation map generation model training method according to claim 4, characterized in that The prototype activation map is calculated by the following formula: in, is the j-th prototype activation map of the labeled category k, is the fusion feature map, is the j-th prototype vector of the labeled category k, express and The inner product of The weight map is calculated by the following formula: in, is the j-th prototype activation map of the labeled category k, is the fusion feature map, is the activation function, for convolution; The class activation map is calculated by the following formula: in, is the class activation map of the labeled category k, and K is the number of prototype vectors of the labeled category k.

6. The class activation map generation model training method according to claim 1, characterized in that The calculation formula of the joint training loss is: in, is the cross entropy loss, is the probability of predicting the labeled category i, is the reciprocal of the proportion of samples labeled category i, is the auxiliary adjustment factor; is the IoU similarity loss, is the class activation graph, mask for the lesion; is the temporal consistency loss, are the class activation maps of two consecutive skin ultrasound images in the skin ultrasound video, is the L2 norm; is the multi-lesion decoupling regularization term, express and For different annotation types, are class activation maps of different annotation types for the same skin ultrasound image, is the Hadamard product, is the L1 norm; , , , is the weight of each loss.

7. A method for generating a class activation map, characterized in that: include: Obtain key frames in the skin ultrasound video to be classified; Inputting the key frame into a pre-trained feature extraction model to obtain key frame features; Inputting the key frame features into a pre-trained hierarchical model to obtain a skin structure mask; The key frame features and the skin structure mask are input into a class activation map generation model to obtain class activation maps of each category of the key frame; the class activation map generation model is trained by the class activation map generation model training method according to any one of claims 1 to 6.

8. A class activation map generation model training device, characterized in that: include: A sample data acquisition unit, configured to acquire multiple sets of sample data; wherein each set of sample data includes a skin ultrasound video and a labeled category of the skin ultrasound video, the skin ultrasound video includes multiple consecutive skin ultrasound images, the skin ultrasound images include key frames and non-key frames, and each set of sample data also includes a skin structure mask of each skin ultrasound image and a lesion mask of the key frame; A prototype vector initialization unit is used to initialize multiple prototype vectors corresponding to each of the labeled categories according to the labeled categories of all the sample data and preset prototype hyperparameters; wherein the prototype hyperparameters are used to specify the number of prototype vectors corresponding to each of the labeled categories, and the prototype vectors represent different morphological subtype features under the same labeled category; a feature fusion unit, configured to input, for each skin ultrasound image in each set of sample data, a skin feature map and a corresponding skin structure mask into a medical attention guidance module, and process the skin feature map and the corresponding skin structure mask through an attention mechanism to obtain a fused feature map after fusion of attention; wherein the skin feature map is extracted from the skin ultrasound image; A category prediction unit, configured to input the fused feature map corresponding to the key frame into a fully connected classification layer to obtain a predicted category of the key frame; A class activation map prediction unit is used to input the fused feature map and each prototype vector into a class activation map generation module for weighted fusion to obtain a class activation map of each labeled category; The loss calculation unit is used to calculate the joint training loss of each of the predicted categories, each of the labeled categories, each of the lesion masks, and each of the class activation maps, and update the class activation map generation model according to the joint training loss.

9. A class activation map generating device, characterized in that: include: A key frame acquisition unit, used to acquire key frames in the skin ultrasound video to be classified; A key frame feature extraction unit, configured to input the key frame into a pre-trained feature extraction model to obtain key frame features; a skin structure extraction unit, configured to input the key frame features into a pre-trained hierarchical model to obtain a skin structure mask; A class activation map generation unit is used to input the key frame features and the skin structure mask into a class activation map generation model to obtain class activation maps of each category of the key frame; the class activation map generation model is trained by the class activation map generation model training method according to any one of claims 1 to 6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the class activation map generation model training method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Ophthalmology ultrasound image classification method and device based on neural network model

    CN115631367A

  • Weak supervision semantic segmentation method and device based on attention mask

    CN116935055A

  • Weakly supervised pathological image tissue segmentation method and device based on all-digital pathological section inaccurate point labeling

    CN117152068A

  • Pathology classification model training method, pathology classification method and electronic equipment

    CN119027761A

  • Image segmentation model training method, image segmentation method, electronic equipment and storage medium

    CN120259341A