A training method, application method and electronic device for a medical image retrieval network

By introducing multi-scale modules and convolutional self-attention modules into medical image retrieval networks, the problem of insufficient consideration of the interaction of medical image multi-scale information and channel domain information in the prior art is solved, and more efficient and accurate medical image retrieval is achieved.

CN115757844BActive Publication Date: 2025-06-27WUHAN UNIV OF TECH CHONGQING RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211512870.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-06-27
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing deep hash medical image retrieval method does not fully consider the multi-scale information of medical images, resulting in information loss and ignore information interaction in the channel domain, affecting the quality of the hash code, thereby reducing the retrieval efficiency.

Method used

A medical image retrieval network training method is proposed. By constructing a network including multi-scale module, convolutional self-attention module and hash tag module, multi-scale feature extraction and attention supervision and constraint processing are performed to generate high-quality hash codes.

Benefits of technology

By fully extracting multi-scale information of medical images and enhancing information interaction in the channel domain, network performance is improved and the accuracy and efficiency of medical image retrieval is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115757844B_ABST
    Figure CN115757844B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, an application method and an electronic device for a medical image retrieval network, including: obtaining medical triple instances and medical triple instance labels; constructing a medical image retrieval network including a multi-scale module, a convolutional self-attention module and a hash tag module, using the medical triple instances and labels as the input of the medical image retrieval network, obtaining multi-scale feature maps based on the multi-scale module, obtaining an attention supervision constraint map based on the convolutional self-attention module, and obtaining corresponding hash codes and predicted labels based on the hash tag module; setting the parameters of the overall loss function, training the medical image retrieval network until the loss no longer decreases, and obtaining a trained and complete medical image retrieval network. The present invention captures multi-scale information of medical images and enhances the information interaction in the channel domain based on the multi-scale module and the convolutional self-attention module, optimizes and improves the medical image retrieval network, and improves the accuracy of medical image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image retrieval, and in particular to a method for training a medical image retrieval network, an application method, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of ray imaging technology, medical data has gradually become electronic, and the number of medical images has increased sharply. In order to better assist medical diagnosis and evaluation, it is crucial to mine useful information from large-scale medical images. Therefore, medical image retrieval has attracted wide attention.

[0003] The goal of medical image retrieval is to retrieve similar medical images in a large medical image database, which can present the context of the queried medical image and thus help with diagnosis. Since it is for a large-scale medical image dataset, medical image retrieval algorithms need to have good scalability and accuracy. The deep hashing technology projects high-dimensional features into low-dimensional binary codes by using a deep neural network, accelerating the retrieval process and improving the retrieval efficiency. Therefore, deep hash code learning has been widely applied in medical image retrieval.

[0004] Although the existing deep hashing medical image retrieval methods have achieved good performance, there are still the following problems: 1. They do not fully consider the multi-scale information of medical images, resulting in a large amount of information loss; 2. They ignore the information interaction in the channel domain to capture discriminable regions, affecting the quality of the hash code. Summary of the Invention

[0005] In view of this, it is necessary to propose a method for training a medical image retrieval network, an application method, an electronic device, and a storage medium, which are used to solve the technical problems in the prior art that the multi-scale information of medical images is not considered, resulting in a large amount of information loss, and the information interaction in the channel domain is ignored, affecting the quality of the hash code and thus the low efficiency of medical image retrieval.

[0006] To solve the above problems, the present invention provides a method for training a medical image retrieval network, including:

[0007] Obtaining an anchor medical instance, a positive medical instance, and a negative medical instance that constitute a medical triple instance, as well as an anchor medical instance label, a positive medical instance label, and a negative medical instance label that constitute a medical triple instance label;

[0008] Construct a medical image retrieval network including a multi-scale module, a convolutional self-attention module, and a hash tag module. Use medical triples instances and medical triples labels as the input of the medical image retrieval network. Based on the multi-scale module, perform multi-scale feature extraction on the medical triples instances and medical triples labels to obtain a multi-scale feature map. Based on the convolutional self-attention module, perform attention supervision and constraint processing on the multi-scale feature map to obtain an attention supervision and constraint map. Based on the hash tag module, generate corresponding hash codes and prediction labels from the attention supervision and constraint map;

[0009] Set the parameters of the overall loss function of the medical image retrieval network. Train the medical image retrieval network based on the hash codes and prediction labels until the loss no longer decreases, and obtain a trained and complete medical image retrieval network.

[0010] Further, the performing multi-scale feature extraction on the medical triples instances and medical triples labels based on the multi-scale module to obtain a multi-scale feature image includes:

[0011] Perform feature extraction on the input medical triples instances and medical triples labels to obtain a triple feature map, and perform convolution operations on the triple feature map through three parallel convolutional layers, three parallel first dense blocks, and a max pooling layer to obtain three output feature maps with the same height and width. Then, splice the three feature maps along the channel direction to obtain the multi-scale feature map.

[0012] Further, the performing attention supervision and constraint processing on the multi-scale feature image based on the convolutional self-attention module to obtain an attention supervision and constraint map includes:

[0013] Input the multi-scale feature map, learn attention weights in multiple parallel ways through three parallel convolutional layers based on the attention mechanism to obtain an attention map, add and splice the input multi-scale feature map and the attention map to obtain the attention supervision and constraint map, which is used as the output of the convolutional self-attention module.

[0014] Further, the generating corresponding hash codes and prediction labels from the attention supervision and constraint map based on the hash tag module includes:

[0015] Use the attention supervision and constraint map as the input, input it to the fully connected layer after passing through the second dense block. The fully connected layer is respectively connected to the hash layer and the classification layer;

[0016] Convert the attention supervision and constraint map into hash codes based on the hash layer;

[0017] Perform prediction on the attention supervision and constraint map based on the classification layer to obtain corresponding prediction labels.

[0018] Furthermore, the overall loss function consists of a hierarchical similarity function, a semantic learning function, and a class-level preservation function.

[0019] Furthermore, set the parameters of the overall loss function of the medical image retrieval network, including:

[0020] Convert the binary code into a class hash code based on the hierarchical similarity function and use the Euclidean norm to replace the Hamming distance, and determine the similarity of the hash code and the similarity of the deep features according to the class hash code and the Hamming distance respectively, as the similarity parameter;

[0021] Obtain semantic information that enhances the potential correlation of the class hash code based on the semantic learning function, as the semantic information parameter.

[0022] Obtain the learned features based on the class-level preservation function to project and distinguish the class-level information of the hash code, as the class-level information parameter.

[0023] Furthermore, train the medical image retrieval network based on the hash code and the predicted label until the loss no longer decreases, to obtain a trained complete medical image retrieval network, including:

[0024] Use the similarity parameter, the semantic information parameter, and the class-level information parameter as the loss parameters of the loss function. When training based on the hash code and the predicted label, adjust the loss parameters of the overall loss function and the margin threshold of the triplet loss, optimize the loss using the Adam function, perform training for a set number of rounds or until the loss no longer decreases, to obtain a trained complete medical image retrieval network.

[0025] The present invention also provides a method for applying a medical image retrieval network, including:

[0026] Obtain the medical image to be retrieved;

[0027] Input the medical image to be retrieved into the trained complete medical image retrieval network to retrieve similar medical images, where the trained complete medical image retrieval network is determined according to the medical image retrieval network training method described in any one of the above.

[0028] The medical image retrieval network outputs to obtain similar medical images.

[0029] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the medical image retrieval network training method described in any one of the above, and / or the medical image retrieval network application method described above.

[0030] The present invention also provides a computer - storable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the medical image retrieval network training method described in any one of the above, and / or the medical image retrieval network application method described above.

[0031] Compared with the prior art, the beneficial effects of adopting the above - mentioned embodiments are as follows: In the medical image retrieval network training provided by the present invention, first, an anchor medical instance, a positive medical instance, and a negative medical instance are obtained from a training sample set to form a medical triplet label, and an anchor medical instance label, a positive medical instance label, and a negative medical instance label are obtained to form a medical triplet instance label; then, a medical image retrieval network composed of a multi - scale module, a convolutional self - attention module, and a hash tag module is constructed, and the triplet instance and the triplet label are used as the input of the medical image retrieval network; among them, multi - scale feature extraction is performed based on the multi - scale module to obtain a multi - scale feature map, and attention supervision and constraint are performed based on the convolutional self - attention module to obtain an attention supervision and constraint map, which can capture the multi - scale information of medical images and enhance the information interaction in the channel domain; finally, the parameters of the overall loss function are set, and the medical image retrieval network is trained based on the hash code and the predicted label until the loss no longer decreases, and a trained - complete medical image retrieval network is obtained. In the application method provided by the present invention, first, a medical image to be measured is obtained, and then the above - trained - complete medical image retrieval network is used to retrieve similar medical images, and then the similar medical images can be output. In summary, by introducing a multi - scale module and a convolutional self - attention module, the present invention optimizes and improves the medical image retrieval network, fully extracts the multi - scale information of medical images and enhances the information interaction in the channel domain, improves the network performance, and enhances the accuracy of medical image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following - described drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of an embodiment of the medical image retrieval network training method provided by the present invention;

[0034] Figure 2 It is a flowchart of an embodiment of step S102 of the present invention;

[0035] Figure 3 It is an overall network architecture diagram of an embodiment of the medical image retrieval network provided by the present invention;

[0036] Figure 4Schematic structural diagram of an embodiment of the medical image retrieval network training device provided by the present invention;

[0037] Figure 5 Schematic structural diagram of an embodiment of the medical image retrieval network application device provided by the present invention;

[0038] Figure 6 Schematic structural diagram of an embodiment of the electronic device provided by the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0040] It should be understood that the accompanying drawings of the schematic diagrams are not drawn to the actual scale. The flowcharts used in the present invention illustrate the operations implemented according to some embodiments of the present invention. It should be understood that the operations of the flowcharts may not be implemented in sequence, and steps without logical context relationships may be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present invention.

[0041] Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0042] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0043] The embodiment of the present invention adopts a medical image retrieval method based on multi-scale triplet hashing, which will be described below.

[0044] Figure 1 Schematic flowchart of an embodiment of the medical image retrieval network training method provided by the present invention. As Figure 1 shown, the medical image retrieval network training method includes:

[0045] S101. Obtain the anchor medical instance, positive medical instance, and negative medical instance that constitute the medical triple instance, as well as the anchor medical instance label, positive medical instance label, and negative medical instance label that constitute the medical triple instance label;

[0046] S102. Construct a medical image retrieval network including a multi-scale module, a convolutional self-attention module, and a hash tag module. Use the medical triple instance and the medical triple label as the input of the medical image retrieval network. Perform multi-scale feature extraction on the medical triple instance and the medical triple label based on the multi-scale model to obtain a multi-scale feature map. Perform attention supervision and constraint processing on the multi-scale feature map based on the convolutional self-attention module to obtain an attention supervision and constraint map. Generate a corresponding hash code from the attention supervision and constraint map based on the hash tag module;

[0047] S103. Set the parameters of the overall loss function of the medical image retrieval network, and train the medical image retrieval network until the loss no longer decreases to obtain a trained and complete medical image retrieval network.

[0048] In the training of the medical image retrieval network provided by the present invention, multi-scale feature extraction is performed based on the multi-scale module to obtain a multi-scale feature map, and attention supervision and constraint are performed based on the convolutional self-attention module to obtain an attention supervision and constraint map, which can capture the multi-scale information of medical images and enhance the information interaction in the channel domain. Compared with the prior art, by introducing the multi-scale module and the convolutional self-attention module, the present invention optimizes and improves the medical image retrieval network, can fully extract the multi-scale information of medical images and enhance the information interaction in the channel domain, and obtains higher medical image retrieval accuracy.

[0049] In a specific embodiment of the present invention, step S101 of obtaining the anchor medical instance, positive medical instance, and negative medical instance that constitute the medical triple instance, as well as the anchor medical instance label, positive medical instance label, and negative medical instance label that constitute the medical triple instance label includes:

[0050] Use 3 datasets, namely the Curated X-Ray of the combined selected dataset of COVID-19 chest X-ray images, the Skin Cancer MNISTS Dataset dataset, and the COVID-19 RadiographyDataset dataset. For each dataset, 70% of the data is selected as the training set, and the remaining 30% is used as the test and retrieval set. The medical images in the same dataset are of the same type of medical images, and the medical images in different datasets are of different types of medical images.

[0051] For the given I medical triple units and the corresponding triple tags , where represents the i-th anchor medical instance, represents the tag of this anchor medical instance; represents the i-th positive medical instance, represents the tag of this positive medical instance; represents the i-th negative medical instance, represents the tag of this negative medical instance. Compared with the negative medical instance, the anchor medical instance is more similar to the positive medical instance. The purpose is to learn the mapping relationship that maps medical instances to binary codes and at the same time maintain the similarity of similar medical instances in the Hamming space. Further, that is is smaller than , where represents the Hamming distance, , and respectively represent and 's k-bit binary codes.

[0052] In a specific embodiment of the present invention, step S102 includes a medical image retrieval network with a multi-scale module, a convolutional self-attention module, and a hash tag module. The medical triple instances and medical triple tags are used as the input of the medical image retrieval network. As Figure 2 shown, step S102 includes:

[0053] S201. Perform multi-scale feature extraction on the medical triple instances and medical triple tags based on the multi-scale module to obtain a multi-scale feature map;

[0054] S202. Perform attention supervision and constraint processing on the multi-scale feature map based on the convolutional self-attention module to obtain an attention supervision and constraint map;

[0055] S203. Generate corresponding hash codes from the attention supervision and constraint map based on the hash tag module.

[0056] In a specific embodiment of the present invention, performing multi-scale feature extraction on the medical triple instances and medical triple tags based on the multi-scale module to obtain a multi-scale feature image includes:

[0057] Perform feature extraction on the input medical triple instances and medical triple tags to obtain a triple feature map, and perform convolution operations on the triple feature map through three parallel convolutional layers, three parallel first dense blocks, and a max pooling layer to obtain three output feature maps with the same height and width. Then, splice the three feature maps along the channel direction to obtain a multi-scale feature map.

[0058] Specifically, the three parallel convolutional layers respectively contain 16 1×1 convolutions with a stride of 1, 16 3×3 convolutions with a stride of 1, and a 5×5 convolution with a stride of 1. The padding of the above three parallel convolutional layers is 0, 1, and 2 respectively; each first dense block contains four bottleneck layers, and the consecutive operations of each bottleneck layer include: batch normalization, ReLU (Linear rectification function, a commonly used activation function in artificial neural networks, also known as the rectified linear unit), and a 1×1 convolution, batch normalization, ReLU, and a 3×3 convolution; the max pooling layer compresses the input feature map to remove redundant information. For the feature map input to the multi-scale module, its number of channels can be , then the number of channels of the output feature map is , where c is the growth rate and d is the number of layers. The height and width of the feature map passing through the first dense block remain unchanged, but the number of channels changes. That is, the input image is processed through 3 parallel convolutional operations to obtain 3 output feature maps with the same height and width, and then the 3 output feature maps are concatenated along the channel direction to obtain the output multi-scale feature map.

[0059] In a specific embodiment of the present invention, an attention supervision and constraint map is obtained by performing attention supervision and constraint processing on the multi-scale feature image based on a convolutional self-attention module, including:

[0060] Input the multi-scale feature map, pass through three parallel convolutional layers and learn attention weights in multiple parallel ways through an attention mechanism to obtain an attention map, and add and concatenate the input multi-scale feature map with the obtained attention map to obtain the attention supervision and constraint map, which is used as the output of the convolutional self-attention module.

[0061] Specifically, the multi-scale feature map is used as the input of the convolutional self-attention module. The convolutional self-attention module contains three parallel convolutions and uses Q, K, and V (Query, Key, and Value, lookup vector, retrieved vector, content vector) to learn attention weights in multiple parallel ways. For each head , the input is transformed into through a learnable parameter matrix , where and are -dimensional vectors; the convolutional self-attention module can be expressed as:

[0062]

[0063] Among them, , represents a learnable parameter matrix, represents the softmax (logistic regression model) function, where the dot product in the formula is scaled by for scaling.

[0064] Specifically, the multi-scale feature map is passed through three parallel 1×1 convolutions to obtain three vectors Q, K, and V; the attention map is calculated through the attention module equation, and then the original feature map and the obtained attention map are added and concatenated to obtain the attention supervision constraint map, which is used as the final output of the convolutional self-attention module.

[0065] In a specific embodiment of the present invention, based on the hash tag module, the attention supervision constraint map is used to generate corresponding hash codes and prediction labels, including:

[0066] Taking the attention supervision constraint map as the input of the hash tag module, after passing through the second dense block, it is input to the fully connected layer, and the fully connected layer is respectively connected to the hash layer and the classification layer; among them, the hash layer converts the attention supervision constraint map into a hash code, and the classification layer predicts the attention supervision constraint map to obtain the corresponding prediction label.

[0067] Specifically, the activation function of the second dense block is the ReLU function, the fully connected layer includes 1024 nodes, the classification layer contains s nodes, where s represents the number of sample categories, the activation function is the softmax function, and the hash layer uses a function similar to the sign function (such as the tanh function) as the activation function, and k-bit class hash codes are generated during the training process.

[0068] Correspondingly, for the i-th medical triple unit , the deep hash function can be written as:

[0069]

[0070] where

[0071] where represents the k-bit hash code of the medical instance , , represents the deep hash function, represents the deep feature of the medical instance , represents the weight, represents the tanh function, represents the sign function.

[0072] It should be further noted that although the present invention fully extracts the multi-scale information of medical images and enhances the information interaction in the channel domain through the multi-scale module and the convolutional self-attention module, improving the accuracy of the medical image retrieval network, in the loss function, there is still a problem that the loss function only considers the similarity of hash codes, but the similarity of deep features is not compatible with the similarity of hash codes. Therefore, the present invention also uses a total objective loss function composed of a hierarchical similarity function, a semantic learning function, and a class-level preservation function, which can not only preserve the class-level information of deep features and the semantic information of hash codes, but also improve the similarity compatibility between hierarchical similarity deep features and hash codes, further improving the accuracy of the medical image retrieval network. The following is a detailed description.

[0073] In a specific embodiment of the present invention, the overall loss function is composed of a hierarchical similarity function, a semantic learning function, and a class-level preservation function; among them, the hierarchical similarity function converts the binary code into a class hash code, uses the second norm instead of the Hamming distance, and determines the hash code similarity and the deep feature similarity respectively according to the class hash code and the Hamming distance as the similarity parameters; the semantic learning function obtains the semantic information that enhances the potential correlation of the class hash code as the semantic information parameter; the class-level preservation function obtains the learned features to project and distinguish the class-level information of the hash code as the class-level information parameter.

[0074] Specifically, the multi-scale triplet hashing method learns the mapping relationship, maps medical instances to hash codes, and at the same time needs to maintain the similarity between similar medical images. To achieve this goal, compared with the hash code of the negative medical instance the hash code of the anchor medical instance should be more similar to the hash code of the positive medical instance In addition, compared with the deep feature of the negative medical instance the deep feature of the anchor medical instance should be more similar to the deep feature of the positive medical instance To achieve this goal, it is necessary to capture the similarity of hash codes at the same time. The hierarchical similarity function can be written as: the deep feature of the negative medical instance the deep feature of the anchor medical instance the deep feature of the positive medical instance should be more similar. To achieve this goal, it is necessary to capture the similarity of hash codes at the same time. The hierarchical similarity function can be written as: the deep feature of the positive medical instance To achieve this goal, it is necessary to simultaneously capture the similarity of hash codes. The hierarchical similarity function can be written as:

[0075]

[0076] where represents the Hamming distance, represents the margin threshold for hash code similarity learning, represents the margin threshold for deep feature similarity learning, represents the maximum function, Indicates the second normal form.

[0077] However, the above hierarchical similarity function is difficult to optimize during training. Therefore, the binary code , and are transformed into class hash codes , and the second normal form is used to replace the Hamming distance. Then the hierarchical similarity function can be written as:

[0078]

[0079] Among them,

[0080] Among them, represents the hierarchical similarity function, which can capture the similarity of hash codes and the similarity of deep features simultaneously. represents the medical instance 's class hash code, and , represents the weight.

[0081] Semantic information is crucial for enhancing the potential correlation of class hash codes. Therefore, in the specific embodiments of the present invention, label information is used to provide semantic information for learning the hash function. Then the semantic learning function can be expressed as:

[0082]

[0083] Among them, represents the semantic learning function, which maintains the semantic information of the hash code during the hash code learning process. represents the cross-entropy function. represents the projection function. represents 's label. represents 's label. represents 's label.

[0084] Category-level information helps to learn good features to project and distinguish hash codes. The more effective the deep features are, the stronger the hash code discrimination ability is. In the specific embodiments of the present invention, the category-level retention function can be expressed as:

[0085]

[0086] Among them, represents the category-level retention function, which can maintain the category-level information of the deep features. represents the softmax function. represents 's category layer weight. Represents the category layer weight of Represents the category layer weight of

[0087] Taking into account the hierarchical similarity function, semantic learning function and category-level preservation function, the overall loss function can be expressed as:

[0088]

[0089] Among them, and represent hyperparameters that control the loss term weights, represents the objective function.

[0090] In a specific embodiment of the present invention, the medical retrieval network model is trained until the loss no longer decreases, and a trained complete medical image retrieval network is obtained, including:

[0091] Taking the similarity parameter, semantic information parameter and category-level information parameter as the loss parameters of the loss function, when the sample is input for training, the loss parameters of the overall loss function and the margin threshold of the triplet loss are adjusted, and the Adam function is used to optimize the loss, and training is carried out for a set number of rounds or until the loss no longer decreases, and a trained complete network model is obtained.

[0092] Specifically, when training the overall network model, the size of the medical triplet image is adjusted to 256×256, and random sampling is performed in each round of training as the input of the network. The margin threshold and of the triplet loss are both set to 0.5, and the parameters and of the overall loss function are set to 0.8 and 1 respectively. The network uses the Adam function to optimize the loss, the learning rate is 0.001, the performance of the evaluation hash code bits from 8, 16, 32, 48 to 64 and the performance of the most similar images from 5, 10, 15, 20, 25 to 30 are evaluated, and training is carried out for 100 rounds or until the loss no longer decreases, and a trained model is obtained.

[0093] Next, a specific application scenario will be combined to more clearly illustrate the technical solution of the present invention and evaluate the effectiveness of the present invention. The specific process is as follows:

[0094] I. Preparation of the dataset:

[0095] Three datasets are used, namely the Curated X-Ray, a combined and selected dataset of COVID-19 chest X-ray images, the Skin Cancer MNISTS Dataset, and the COVID-19 Radiography Dataset. For each dataset, 70% of the data is selected as the training set to train a complete medical image retrieval network, and the remaining 30% is used as the test and retrieval set to test and evaluate the effectiveness of the medical image retrieval network.

[0096] II. Application process:

[0097] According to the specific embodiments of the above-mentioned medical image retrieval network training method, the medical image retrieval network is trained using the medical triplet images and medical triplet labels composed of the training set data until the loss function of the medical image retrieval network no longer decreases, and a trained medical image retrieval network model is obtained.

[0098] Figure 3 This is the overall network architecture diagram of a specific embodiment of the medical image retrieval network of the present invention. The structure of the medical image retrieval network used in the application scenario is as Figure 3 shown. In the medical image retrieval network, it includes:

[0099] Medical triplet instance images (triplet images) composed of positive medical instances (positive), anchor medical instances (anchor), and negative medical instances (negative) are connected to a convolutional self-attention module after passing through a multi-scale module (Multi-Scale DenseBlock). The convolutional self-attention module adds and splices the multi-scale feature map (feature map) and the attention map (attention map) to obtain an attention supervision constraint map, which is then passed into the hash label module. The hash label module includes: a second dense block (DenseBlock), a fully connected layer (fully connected layer), a hash layer (hash layer), and a classification layer (predict layer). Both the hash layer and the classification layer are connected to the fully connected layer to obtain the output hash code and prediction label respectively, and the medical image retrieval network is optimized through the overall loss function.

[0100] The average hit rate (mHR), average average precision (mAP), and average reciprocal rank (mRR) of the sample images in the test dataset are calculated using the trained medical image retrieval neural network, and the retrieval performance is evaluated based on these three metrics. Among them, the hit rate (HR) is used to measure how many images in the returned list are similar to the query image; in the returned list, the average precision (AP) averages the ranking positions of the images similar to the query image to measure the ranking quality; the reciprocal rank (RR) refers to the reciprocal position of the sorting of the first similar image in the returned list.

[0101] III. Result Analysis:

[0102] To evaluate the effectiveness of the method of the present invention, the present invention is compared with advanced methods such as ASH, ATH, DHN, DPSH, DSH, DTSH, and IDHN in terms of retrieval performance.

[0103] Table 1

[0104]

[0105] Table 1 shows the comparison test results of the present invention and other methods on the Curated X-Ray dataset, Skin Cancer MNIST dataset, and COVID-19 Radiography dataset through the test metrics mHR@10 (average hit rate with 10 similar images), mAP@10 (average average precision with 10 similar images), and mRR@10 (average reciprocal rank with 10 similar images). It can be seen from the comparison results that the average precision index of the top 10 retrieval results of the specific embodiment of the present invention is the highest.

[0106] The embodiment of the present invention also provides a medical image retrieval network training device. Figure 4 FIG. is a schematic structural diagram of an embodiment of the medical image retrieval network training device adopted by the present invention. The medical image retrieval network training device 400 includes:

[0107] A first acquisition unit 401, configured to acquire an anchor medical instance, a positive medical instance, and a negative medical instance that make up a medical triple instance, as well as an anchor medical instance label, a positive medical instance label, and a negative medical instance label that make up a medical triple instance label;

[0108] The first processing unit 402 is configured to construct a medical image retrieval network including a multi-scale module, a convolutional self-attention module, and a hash tag module, take a medical triple instance and a medical triple label as inputs of the medical image retrieval network, perform multi-scale feature extraction on the medical triple instance and the medical triple label based on the multi-scale module to obtain a multi-scale feature map, perform attention supervision and constraint processing on the multi-scale feature map based on the convolutional self-attention module to obtain an attention supervision and constraint map, and generate a corresponding hash code based on the attention supervision and constraint map by the hash tag module;

[0109] The training unit 403 is configured to set parameters of the overall loss function of the medical image retrieval network, and train the medical image retrieval network until the loss no longer decreases, so as to obtain a trained complete medical image retrieval network.

[0110] An embodiment of the present invention further provides a medical image retrieval network application device, Figure 5 which is a schematic structural diagram of an embodiment of the medical image retrieval network application device provided by the present invention. The medical image retrieval network application device 500 includes:

[0111] The second acquisition unit 501 is configured to acquire a medical image to be retrieved;

[0112] The second processing unit 502 inputs the medical image to be retrieved into the trained complete medical image retrieval network to retrieve similar medical images, where the trained complete medical image retrieval network is determined according to the medical image retrieval network training method described above;

[0113] The image retrieval unit 503 outputs the similar medical images retrieved by the medical image retrieval network.

[0114] The present invention further provides an electronic device, as Figure 6 shown, Figure 6 which is a schematic structural diagram of an embodiment of the electronic device provided by the present invention. The electronic device 600 includes a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the processor 601 executes the program, the medical image retrieval network training method described above and / or the medical image retrieval network application method described above are implemented.

[0115] As a preferred embodiment, the above-mentioned electronic device 600 further includes a display 603 for displaying the medical image retrieval network training method described above and / or the medical image retrieval network application method described above executed by the processor 601.

[0116] Exemplarily, a computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory 602 and executed by the processor 601 to implement the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 600. For example, the computer program can be divided into the first acquisition unit 401, the first processing unit 402, the training unit 403, the second acquisition unit 501, the second processing unit 502, and the image retrieval unit 503 in the above embodiments. The specific functions of each unit are as described above and will not be elaborated here one by one.

[0117] The electronic device 600 can be a desktop computer with a camera module, a notebook, a palm computer, a smart phone, or other devices.

[0118] Among them, the processor 601 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC). It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can also be a microprocessor, or the processor can also be any conventional processor, etc.

[0119] Among them, the memory 602 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a secure digital (SD) card, a flash card, etc. The memory 602 is used to store programs. After receiving the execution instruction, the processor 601 executes the program. The method defined by the process disclosed in any of the embodiments of the present invention can be applied to the processor 601 or implemented by the processor 601.

[0120] Among them, the display 603 can be an LED display screen, a liquid crystal display, or a touch display, etc. The display 603 is used to display various information in the electronic device 600.

[0121] It can be understood that Figure 6 The structure shown is only a schematic diagram of the structure of the electronic device 600. The electronic device 600 may also include moreFigure 6 More or fewer components as shown. Figure 6 Each component shown in can be implemented by hardware, software, or a combination thereof.

[0122] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned medical image retrieval network training method and / or the above-mentioned medical image retrieval network application method are implemented.

[0123] Generally speaking, computer instructions for implementing the method of the present invention can be carried by any combination of one or more computer-readable storage media. A non-transitory computer-readable storage medium can include any computer-readable medium except for the signal itself in transient propagation.

[0124] The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the context of the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0125] The computer program code for performing the operations of the present invention can be written in one or more programming languages or a combination thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. In particular, the Python language suitable for neural network computing and platform frameworks based on TensorFlow, PyTorch, etc. can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0126] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for training a medical image retrieval network, characterized in that Including: Obtaining the anchor medical instance, positive medical instance, and negative medical instance that constitute the medical triple instance, as well as the anchor medical instance label, positive medical instance label, and negative medical instance label that constitute the medical triple instance label; Constructing a medical image retrieval network including a multi-scale module, a convolutional self-attention module, and a hash tag module, taking the medical triple instance and the medical triple label as the input of the medical image retrieval network, performing multi-scale feature extraction on the medical triple instance and the medical triple label based on the multi-scale module to obtain a multi-scale feature map, performing attention supervision and constraint processing on the multi-scale feature map based on the convolutional self-attention module to obtain an attention supervision and constraint map, and generating corresponding hash codes and prediction labels based on the attention supervision and constraint map by the hash tag module; Among them, the performing attention supervision and constraint processing on the multi-scale feature map based on the convolutional self-attention module to obtain an attention supervision and constraint map includes: inputting the multi-scale feature map, learning attention weights in multiple parallel ways through three parallel convolutional layers and based on the attention mechanism to obtain an attention map, adding and splicing the input multi-scale feature map with the attention map to obtain an attention supervision and constraint map, which is used as the output of the convolutional self-attention module; The generating corresponding hash codes and prediction labels based on the attention supervision and constraint map by the hash tag module includes: taking the attention supervision and constraint map as the input, inputting it into a fully connected layer after passing through a second dense block, and the fully connected layer is respectively connected to a hash layer and a classification layer; converting the attention supervision and constraint map into hash codes based on the hash layer; predicting the attention supervision and constraint map based on the classification layer to obtain corresponding prediction labels; Setting the parameters of the overall loss function of the medical image retrieval network, and training the medical image retrieval network based on the hash codes and prediction labels until the loss no longer decreases to obtain a trained medical image retrieval network; Among them, the overall loss function is composed of a hierarchical similarity function, a semantic learning function, and a class-level preservation function; The setting the parameters of the overall loss function of the medical image retrieval network includes: converting the binary code into a class hash code based on the hierarchical similarity function and using the Euclidean norm to replace the Hamming distance, and respectively determining the similarity of the hash codes and the similarity of the deep features based on the class hash code and the Hamming distance as the similarity parameter; obtaining semantic information that enhances the potential correlation of the class hash codes based on the semantic learning function as the semantic information parameter; obtaining the learned features to project and distinguish the class-level information of the hash codes based on the class-level preservation function as the class-level information parameter.

2. The medical image retrieval network training method according to claim 1, wherein The performing multi-scale feature extraction on the medical triple instance and the medical triple label based on the multi-scale module to obtain a multi-scale feature map includes: Feature extraction is performed on the input medical triple instances and medical triple labels to obtain triple feature maps, and convolution operations on the triple feature maps are performed through three parallel convolutional layers, three parallel first dense blocks, and max pooling layers to obtain three output feature maps with the same height and width. Then, the three feature maps are concatenated along the channel direction to obtain the multi-scale feature map.

3. The medical image retrieval network training method according to claim 1, wherein, Based on the hash code and predicted label, the medical image retrieval network is trained until the loss no longer decreases, and a trained complete medical image retrieval network is obtained, including: Using the similarity parameter, semantic information parameter, and class-level information parameter as the loss parameters of the loss function, when training based on the hash code and predicted label, adjust the loss parameters of the overall loss function and the margin threshold of the triple loss, optimize the loss using the Adam function, perform training for a set number of rounds or until the loss no longer decreases, and obtain a trained complete medical image retrieval network.

4. A method for applying a medical image retrieval network, characterized in that, Including: Obtain the medical image to be retrieved; Input the medical image to be retrieved into the trained complete medical image retrieval network to retrieve similar medical images, where the trained complete medical image retrieval network is determined according to the medical image retrieval network training method described in any one of claims 1 to 3; The medical image retrieval network outputs to obtain similar medical images.

5. An electronic device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the medical image retrieval network training method described in any one of claims 1 to 3, and / or the medical image retrieval network application method described in claim 4.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the medical image retrieval network training method described in any one of claims 1 to 3, and / or the medical image retrieval network application method described in claim 4.

Citation Information

Patent Citations

  • Semantic enhanced hash medical image retrieval method based on mixed attention

    CN113889228A

  • Systems and methods for identifying a target object in an image

    US20180330198A1