An image retrieval method, device, equipment and storage medium

CN116010643BActive Publication Date: 2026-08-11GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

针对单张图像中感兴趣区域的特征比较容易识别,但两张图像的感兴趣区域一起比较相似性却很有挑战性,从而导致图像检索效率和准确度不高

Benefits of technology

[0024]The image retrieval method provided in this invention extracts image features from the image to be retrieved, extracts cross-channel dependencies based on a channel attention mechanism to obtain channel attention features, extracts cross-spatial dependencies based on a spatial attention mechanism to obtain spatial attention features, fuses the channel attention features and spatial attention features to obtain fused features, and retrieves historical images matching the image to be retrieved from a pre-established historical image fused feature library based on the fused features. By enhancing the attention scores of channels with important information through the channel attention mechanism and enhancing the attention scores of regions with important information through the spatial attention mechanism, the method can effectively extract discriminative visual information from images, accurately distinguish between two highly sensitive images, and improve the efficiency and accuracy of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116010643B_ABST
    Figure CN116010643B_ABST
Patent Text Reader

Abstract

This invention discloses an image retrieval method, apparatus, device, and storage medium. Image features are extracted from the image to be retrieved. Cross-channel dependencies are extracted from these features using a channel attention mechanism to obtain channel attention features. Cross-spatial dependencies are extracted from these features using a spatial attention mechanism to obtain spatial attention features. The channel attention features and spatial attention features are fused to obtain fused features. Based on these fused features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fused feature database. By enhancing the attention scores of channels with important information using the channel attention mechanism and enhancing the attention scores of regions with important information using the spatial attention mechanism, this method effectively extracts discriminative visual information from images, accurately distinguishes two highly sensitive images, and improves image retrieval efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to computer-aided image processing technology, and more particularly to an image retrieval method, apparatus, device and storage medium. Background Technology

[0002] With the rapid development of information technology, the medium for carrying information has gradually shifted from text to images. Images can typically carry more information and are more vivid and intuitive. This trend is driving the development of image retrieval.

[0003] Traditional image retrieval employs text-based methods, which first describe the image in text (establishing a correspondence between text and image), then input keywords to search and return ranked results. However, text often fails to accurately and completely describe the content of an image. Therefore, this "text-based image search" method suffers from semantic discrepancies between the text description and the image content itself, affecting search effectiveness.

[0004] With the development of computer vision technology, content-based image retrieval (CBIR) methods have begun to emerge. This "image-to-image" method retrieves images based on their own features such as color, shape, and texture, avoiding semantic differences between text descriptions and image content.

[0005] However, in some fields (e.g., medical imaging), the visual features of regions of interest (ROIs) in images are highly sensitive, meaning that most of the information in the ROIs is consistent, with only subtle distinctions. While identifying ROIs in a single image is relatively easy, comparing the similarity of ROIs between two images is very challenging, resulting in low efficiency and accuracy in image retrieval. Summary of the Invention

[0006] This invention provides an image retrieval method, apparatus, device, and storage medium to improve image retrieval efficiency and accuracy.

[0007] In a first aspect, embodiments of the present invention provide an image retrieval method, including:

[0008] Extract image features from the image to be retrieved;

[0009] Based on the channel attention mechanism, cross-channel dependencies in the image features are extracted to obtain channel attention features;

[0010] Based on the spatial attention mechanism, cross-space dependencies in the image features are extracted to obtain spatial attention features;

[0011] The channel attention features and the spatial attention features are fused to obtain the fused features;

[0012] Based on the fusion features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fusion feature library.

[0013] Secondly, embodiments of the present invention also provide an image retrieval device, comprising:

[0014] The feature extraction module is used to extract image features from the image to be retrieved;

[0015] The channel attention module is used to extract cross-channel dependencies in the image features based on the channel attention mechanism to obtain channel attention features;

[0016] The spatial attention module is used to extract cross-space dependencies in the image features based on the spatial attention mechanism to obtain spatial attention features;

[0017] The feature fusion module is used to fuse the channel attention features and the spatial attention features to obtain fused features;

[0018] The matching module is used to retrieve historical images that match the image to be retrieved from a pre-established historical image historical fusion feature library based on the fusion features.

[0019] Thirdly, embodiments of the present invention also provide a computer device, comprising:

[0020] One or more processors;

[0021] Storage device for storing one or more programs;

[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the image retrieval method as provided in the first aspect of the present invention.

[0023] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image retrieval method as provided in the first aspect of the present invention.

[0024] The image retrieval method provided in this invention extracts image features from the image to be retrieved, extracts cross-channel dependencies based on a channel attention mechanism to obtain channel attention features, extracts cross-spatial dependencies based on a spatial attention mechanism to obtain spatial attention features, fuses the channel attention features and spatial attention features to obtain fused features, and retrieves historical images matching the image to be retrieved from a pre-established historical image fused feature library based on the fused features. By enhancing the attention scores of channels with important information through the channel attention mechanism and enhancing the attention scores of regions with important information through the spatial attention mechanism, the method can effectively extract discriminative visual information from images, accurately distinguish between two highly sensitive images, and improve the efficiency and accuracy of image retrieval. Attached Figure Description

[0025] Figure 1 This is a flowchart of an image retrieval method provided in Embodiment 1 of the present invention;

[0026] Figure 2A This is a flowchart of an image retrieval method provided in Embodiment 2 of the present invention;

[0027] Figure 2B This is a schematic diagram of an image retrieval model structure provided in an embodiment of the present invention;

[0028] Figure 2C This is a schematic diagram of the structure of a feature extraction network provided in an embodiment of the present invention;

[0029] Figure 2D A data processing flowchart of a channel attention mechanism provided in an embodiment of the present invention;

[0030] Figure 2E A data processing flowchart of a spatial attention mechanism provided in an embodiment of the present invention;

[0031] Figure 3A This is a flowchart of a channel attention mechanism processing method provided in an embodiment of the present invention;

[0032] Figure 3B A data processing flowchart for another channel attention mechanism provided in an embodiment of the present invention;

[0033] Figure 4 This is a schematic diagram of the structure of an image retrieval device provided in Embodiment 4 of the present invention;

[0034] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 5 of the present invention. Detailed Implementation

[0035] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0036] Example 1

[0037] Figure 1 This is a flowchart of an image retrieval method provided in Embodiment 1 of the present invention. This embodiment is applicable to image retrieval with highly sensitive visual features. The method can be executed by the image retrieval device provided in this embodiment of the present invention. This device can be implemented in software and / or hardware, and is typically configured in a computer device, such as... Figure 1 As shown, the method specifically includes the following steps:

[0038] S101. Extract image features from the image to be retrieved.

[0039] In this embodiment of the invention, an image to be retrieved is acquired, which has highly sensitive visual features. For example, in this embodiment, the image to be retrieved is a medical image. The visual features of the lesion area in a medical image are highly sensitive; most of the information about the lesion area in different images is consistent, with only subtle differences.

[0040] After obtaining the image to be retrieved, a pre-trained feature extraction network can be used to process the image and extract image features that characterize the image. These image features may include color features, texture features, shape features, and spatial relationship features, etc., which are not limited to these specific features in this embodiment of the invention.

[0041] For example, the feature extraction network can be a deep learning-based neural network, such as VGG Net, ResNet, ResNeXt, SE-Net, etc., and the embodiments of the present invention are not limited thereto.

[0042] S102. Extract cross-channel dependencies in image features based on the channel attention mechanism to obtain channel attention features.

[0043] In this embodiment of the invention, cross-channel dependencies in image features are extracted based on a channel attention mechanism. Based on these cross-channel dependencies, each channel of the image features is weighted to enhance the attention score of channels with important information, thereby obtaining channel attention features.

[0044] S103. Extract cross-space dependencies in image features based on spatial attention mechanism to obtain spatial attention features.

[0045] In this embodiment of the invention, spatial attention mechanisms are used to extract cross-spatial dependencies in image features. Based on these cross-spatial dependencies, each element in the image features is weighted to enhance the attention score of regions with important information, thereby obtaining spatial attention features.

[0046] S104. Merge channel attention features and spatial attention features to obtain fused features.

[0047] In this embodiment of the invention, after obtaining the channel attention features and spatial attention features, the channel attention features and spatial attention features are fused to obtain fused features. Fusion can refer to the spatial concatenation of two features to increase the dimensionality of the fused features, thereby improving the accuracy of retrieval and matching.

[0048] S105. Based on the fusion features, retrieve historical images that match the image to be retrieved from the pre-established historical image fusion feature library.

[0049] In this embodiment of the invention, the historical images (i.e., images that have already existed) are processed in advance using the same steps as described above, such as feature extraction, channel attention, spatial attention, and feature fusion, to obtain the fusion features of the historical images (referred to as historical fusion features). A large number of historical images' historical fusion features form a historical fusion feature library.

[0050] Based on fusion features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fusion feature library. For example, the similarity between the fusion feature and each historical fusion feature in the historical fusion feature library is calculated, and historical images matching the image to be retrieved are determined based on the similarity.

[0051] The image retrieval method provided in this invention extracts image features from the image to be retrieved, extracts cross-channel dependencies based on a channel attention mechanism to obtain channel attention features, extracts cross-spatial dependencies based on a spatial attention mechanism to obtain spatial attention features, fuses the channel attention features and spatial attention features to obtain fused features, and retrieves historical images matching the image to be retrieved from a pre-established historical image fused feature library based on the fused features. By enhancing the attention scores of channels with important information through the channel attention mechanism and enhancing the attention scores of regions with important information through the spatial attention mechanism, the method can effectively extract discriminative visual information from images, accurately distinguish between two highly sensitive images, and improve the efficiency and accuracy of image retrieval.

[0052] Example 2

[0053] Figure 2A This is a flowchart of an image retrieval method provided in Embodiment 2 of the present invention. Figure 2B This is a schematic diagram of an image retrieval model structure provided by an embodiment of the present invention. This embodiment is a refinement based on the above-described embodiment one, and describes in detail the specific implementation process of each step, such as... Figure 2A and Figure 2B As shown, the method includes:

[0054] S201. A feature extraction network is used to expand the channels of the image to be retrieved, and multiple features of different dimensions are extracted as image features.

[0055] In this embodiment of the invention, the image to be retrieved is a 3D medical image, such as a CT (Computed Tomography) scan. The size of the 3D medical image is C×D×H×W, where C is the number of channels in the 3D medical image, D is the 3D depth, H is the 3D height, and W is the 3D width. For example, a 3D medical image with a size of 1×80×80×80 is an image with 80 slices, each slice having one channel (i.e., a grayscale image), and a height and width of 80 pixels.

[0056] A feature extraction network is used to perform channel expansion on the image to be retrieved, extracting multiple features of different dimensions as image features. These multiple features of different dimensions may include texture, size, shape, etc., which are not limited to these aspects in this embodiment of the invention.

[0057] Figure 2C This is a schematic diagram of the structure of a feature extraction network provided in an embodiment of the present invention, such as... Figure 2C As shown, in this embodiment of the invention, the feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, and a fourth convolutional block. The feature extraction network performs channel expansion on the image to be retrieved, extracting multiple features of different dimensions as image features, including:

[0058] 1. The first convolutional block performs convolution processing on the image to be retrieved, resulting in a 64-channel first feature.

[0059] In this embodiment of the invention, the first convolutional block performs convolution processing on the image to be retrieved, outputting a 64-channel first feature, with each channel representing a different feature. This embodiment of the invention does not limit the internal structure of the first convolutional block, as long as it can output a 64-channel first feature.

[0060] For example, such as Figure 2CAs shown, the first convolutional block includes a first convolutional layer, a second convolutional layer, and a first max-pooling layer. The kernel size of the first convolutional layer, the second convolutional layer, and the first max-pooling layer is 3×3×3. The first convolutional layer has 32 convolutional kernels and outputs 32 features. The second convolutional layer has 64 convolutional kernels, each with 32 channels, and outputs 64 features. The first max-pooling layer performs max pooling on each channel of the features output by the second convolutional layer and outputs a first feature with 64 channels.

[0061] 2. The first feature is convolved in the second convolution block to obtain the second feature with 128 channels.

[0062] In this embodiment of the invention, the second convolutional block performs convolution processing on the first feature output by the first convolutional block, outputting a second feature with 128 channels, where each channel represents a different feature. This embodiment of the invention does not limit the internal structure of the second convolutional block, as long as it can output a second feature with 128 channels.

[0063] For example, such as Figure 2C As shown, the second convolutional block includes a third convolutional layer, a fourth convolutional layer, and a second max-pooling layer. The kernel size of the third, fourth, and second convolutional layers is 3×3×3. The third convolutional layer has 64 kernels, each with 64 channels, outputting 64 features. The fourth convolutional layer has 128 kernels, each with 64 channels, outputting 128 features. The second max-pooling layer max-pools each channel of the features output from the fourth convolutional layer and outputs a second feature with 128 channels.

[0064] 3. The second feature is convolved in the third convolution block to obtain the third feature with 256 channels.

[0065] In this embodiment of the invention, the third convolutional block performs convolution processing on the second feature output by the second convolutional block, outputting a third feature with 256 channels, where each channel represents a different feature. This embodiment of the invention does not limit the internal structure of the third convolutional block, as long as it can output a third feature with 256 channels.

[0066] For example, such as Figure 2CAs shown, the third convolutional block includes a fifth convolutional layer, a sixth convolutional layer, and a third max-pooling layer. The kernel size of all three convolutional layers is 3×3×3. The fifth convolutional layer has 128 kernels, each with 128 channels, outputting 128 features. The sixth convolutional layer has 256 kernels, each with 128 channels, outputting 256 features. The third max-pooling layer max-pools each channel of the features output from the sixth convolutional layer and outputs a third feature with 256 channels.

[0067] 4. The third feature is convolved in the fourth convolution block to obtain 512-channel image features.

[0068] In this embodiment of the invention, the fourth convolutional block performs convolution processing on the third feature output by the third convolutional block, outputting 512 channels of image features, with each channel representing a different feature. This embodiment of the invention does not limit the internal structure of the fourth convolutional block, as long as it can output 512 channels of image features.

[0069] For example, such as Figure 2C As shown, the fourth convolutional block includes a seventh and an eighth convolutional layer, both with 3×3×3 kernels. The seventh convolutional layer has 256 kernels, each with 256 channels, outputting 256 feature channels. The eighth convolutional layer has 512 kernels, each with 256 channels, outputting 512×10×10×10 image features. That is, the image features have 512 channels, with a depth, height, and width of 10 pixels each.

[0070] It should be noted that the structure and specific parameters of the above-described feature extraction network are exemplary descriptions of the present invention, and not specific limitations thereof. In other embodiments of the present invention, the feature extraction network may also have different structures and parameters, and the embodiments of the present invention are not limited herein.

[0071] S202. Local convolution is used to capture the local dependencies between K adjacent channels to obtain the weights of each channel.

[0072] For example, in one embodiment of the present invention, local convolution is used to capture the local dependencies between K adjacent channels, and the weight of each channel is determined based on the local dependencies between K adjacent channels.

[0073] Figure 2D A data processing flowchart of a channel attention mechanism provided in an embodiment of the present invention is shown below. Figure 2D As shown, this channel attention mechanism uses local convolution to capture the local dependencies between K adjacent channels. The specific processing procedure is as follows:

[0074] 1. Perform global average pooling on the image features to obtain pooled features.

[0075] For example, global average pooling (AvgPool) is performed on each channel of the image features obtained in the above feature extraction steps. For a 512×10×10×10 image feature, after global average pooling, a 1×512 one-dimensional vector is obtained as the pooling feature.

[0076] 2. A one-dimensional convolution with a kernel size of K is used to convolve the pooling features to obtain local relation features that represent the local dependencies between K adjacent channels.

[0077] For example, a one-dimensional convolution (Conv1) with a kernel size of 3 is used to convolve the pooling features, resulting in local relation features representing the local dependencies between three adjacent channels. The size of the local relation features is 1×512. It should be noted that, to maintain the feature dimension before and after the one-dimensional convolution, zero-padding is required on the pooling features during the convolution process; this will not be elaborated upon here. The one-dimensional convolution process captures the local dependencies between channels; a kernel size of 3 captures the dependencies between three adjacent channels.

[0078] It should be noted that the kernel size of the above one-dimensional convolution can be set according to the actual task requirements, and the embodiments of the present invention do not limit it here.

[0079] 3. Normalize the local relational features to obtain the weights of each channel.

[0080] In this embodiment of the invention, the obtained local relation features are normalized to obtain the weights of each channel. For example, the Sigmoid function is used to normalize the local relation features to obtain the weights of each channel in the feature image.

[0081] S203. Calculate the product of each channel and its corresponding weight in the image features to obtain the channel attention features.

[0082] like Figure 2D As shown, the channel attention features are obtained by multiplying the elements of each channel in the image features by the weights corresponding to that channel.

[0083] S204. Calculate the spatial correlation between each element in the image features and all other elements to obtain the correlation matrix.

[0084] In this embodiment of the invention, the spatial correlation of each element in the image features with all other elements is calculated to obtain a correlation matrix. Spatial correlation refers to the degree of closeness of the relationship between different elements in space.

[0085] Figure 2E A data processing flowchart of a spatial attention mechanism provided in an embodiment of the present invention is shown below. Figure 2E As shown, the processing flow of the spatial attention mechanism is as follows:

[0086] 1. Perform convolution operations in three directions on the image features to obtain the Q matrix, K matrix and V matrix respectively.

[0087] As mentioned earlier, the image features have 512 channels, and their depth, height, and width are all 10 pixels. For example, convolution operations are performed on the image features in three different directions: depth, height, and depth, yielding the Q matrix, K matrix, and V matrix, respectively. The sizes of the Q matrix, K matrix, and V matrix are 512×1000, 1000×512, and 512×1000, respectively.

[0088] 2. Multiply the Q matrix and the K matrix to obtain the correlation weight matrix of the Q matrix and the K matrix.

[0089] Calculate the product of the Q matrix and the K matrix to obtain the correlation weight matrix of the Q matrix and the K matrix, which represents the correlation degree of each element in the Q matrix and the K matrix. The correlation weight matrix is ​​1000×1000 in size, representing that there are 1000 elements in the space, and the correlation degree between the elements is modeled as a 1000×1000 matrix.

[0090] 3. Normalize the correlation weight matrix to obtain the correlation weight coefficient matrix of Q matrix and K matrix.

[0091] For example, the elements of each row or column of the correlation weight matrix are normalized using the softmax function to obtain a normalized correlation weight coefficient matrix. The element with the highest correlation weight coefficient in a given row or column is most relevant to that row or column.

[0092] 4. Multiply the correlation weight coefficient matrix with the V matrix to obtain the correlation matrix, which represents the spatial correlation between each element in the image feature and all other elements.

[0093] Multiplying the correlation weight coefficient matrix by the V matrix yields a correlation matrix representing the spatial correlation between each element in the image feature and all other elements. In this embodiment of the invention, to enable subsequent multiplication with the image feature, the channels of the correlation matrix need to be expanded. Typically, a one-dimensional convolution operation is performed on the correlation matrix to expand it to the same size as the image feature, i.e., 512×10×10×10.

[0094] S205. Calculate the product of the correlation matrix and the image features to obtain the spatial attention matrix.

[0095] like Figure 2E As shown, the product of the correlation matrix and the image features is calculated, and the spatial attention mechanism is applied to all elements of the image features. Essentially, each element of the output is a weighted average of all other elements, and the spatial attention matrix is ​​finally obtained.

[0096] S206. Merge channel attention features and spatial attention features to obtain fused features.

[0097] like Figure 2B As shown in this embodiment of the invention, after obtaining the channel attention features and spatial attention features, they are processed by a Generalized Mean Pooling Layer. The channel attention features, after processing by the Generalized Mean Pooling Layer, yield a 512×1 feature vector, and the spatial attention features, after processing by the Generalized Mean Pooling Layer, yield a 10×10×10 feature vector.

[0098] The two feature vectors mentioned above are concatenated in space through a concatenation layer to increase the dimension of the feature vectors, resulting in a 1×1512 fused feature.

[0099] S207. Calculate the similarity between the fusion feature and each historical fusion feature in the historical fusion feature database.

[0100] In this embodiment of the invention, historical images (i.e., images that already exist) are processed in the same way as described above to obtain the fusion features of the historical images (referred to as historical fusion features), and a large number of historical images' historical fusion features form a historical fusion feature library.

[0101] Calculate the similarity between the fused feature and each historical fused feature in the historical fused feature library. For example, the similarity can be represented by calculating the cosine distance, Euclidean distance, etc., between the fused feature and each historical fused feature in the historical fused feature library.

[0102] S208. Determine the similarity that is greater than the preset similarity threshold as the target similarity.

[0103] The calculated similarities are compared with a preset similarity threshold, and the similarities that are greater than the preset similarity threshold are determined as the target similarities.

[0104] S209. Use the historical image corresponding to the historical fusion feature corresponding to the target similarity as the matching image.

[0105] Exemplarily, the historical image corresponding to the historical fusion feature corresponding to the target similarity is used as the matching image and output for display. In some embodiments of the present invention, the historical images are sorted in descending order according to the target similarity, and the historical images with larger similarities are preferentially displayed at the front of the list for the user to view.

[0106] The image retrieval method provided by the embodiments of the present invention enhances the attention score of the channels with important information through the channel attention mechanism and enhances the attention score of the regions with important information through the spatial attention mechanism, which can effectively extract the discriminative visual information in the image, accurately distinguish two highly sensitive images, and improve the efficiency and accuracy of image retrieval.

[0107] The embodiments of the present invention also provide a training method for an image retrieval model. This method constructs triplets based on class labels and uses triplet loss to train the image retrieval model. The triplet samples <anchor, positive, negative> are collected and input into the model for training. Anchor represents the anchor sample, positive and anchor belong to the same class and serve as positive samples, and negative does not belong to the same class as anchor and serves as a negative sample.

[0108] Triplet loss aims to reduce the distance between positive (positive sample) and anchor, and increase the distance between negative (negative sample) and anchor. Based on the above triplets, a positive pair <a, p> and a negative pair <a, n> can be constructed. The purpose of triplet loss is to separate the positive pair and the negative pair at a certain distance (margin), that is, D(a, p) < D(a, n); at the same time, it satisfies at a certain distance (margin): D(a, p) + margin < D(a, n), where D represents the similarity distance calculation function, usually the cosine similarity function. Through training with triplet loss, the trained model increases the distance between positive and negative samples and improves the similarity between positive samples for retrieval and matching.

[0109] Triple samples can be sampled from all samples (n samples) before input, or they can be sampled from each batch of training samples. Assuming m samples are input into the model for training at one time, triple samples are sampled from these m samples, as in circle loss. n is the total number of samples, m is a batch of all samples, and m is a subset of n samples. These two triple sample sampling strategies can be chosen based on the task. Sampling from all samples before input means the model processes 3 samples each time, which requires significant computational power. Sampling from a batch of m samples after input is related to the selection of samples in each batch, and there may be batches that fail to collect valid triples, but it has the advantage of lower computational requirements. Therefore, a trade-off can be made between performance and efficiency.

[0110] Example 3

[0111] Example 3 provides another image retrieval method. The difference between this method and the image retrieval method described in Example 2 is that, in the channel attention mechanism processing, global convolution is used to capture the global dependencies between all channels. The parts that are the same as those in Example 2 will not be described again here. Figure 3A This is a flowchart of a channel attention mechanism processing method provided in an embodiment of the present invention. Figure 3B A data processing flowchart for another channel attention mechanism provided in an embodiment of the present invention is shown below. Figure 3A and Figure 3B As shown, this channel attention mechanism uses global convolution to capture the global dependencies between all channels. The specific processing procedure is as follows:

[0112] S301. Perform global average pooling on the image features to obtain pooled features.

[0113] For example, global average pooling (AvgPool) is performed on each channel of the image features obtained in the feature extraction step. For a 512×10×10×10 image feature, after global average pooling, a 1×512 one-dimensional vector is obtained as the pooling feature.

[0114] S302. A fully connected layer is used to perform a fully connected mapping on the pooling features to obtain global relation features that represent the global dependencies between all channels.

[0115] Specifically, a fully connected layer (FC layer) is used to perform fully connected mapping on the pooling features. The FC layer weights the contributions of all channels to obtain global relation features that represent the global dependencies between all channels.

[0116] S303. Normalize the global relational features to obtain the weights of each channel.

[0117] In this embodiment of the invention, the obtained global relation features are normalized to obtain the weights of each channel. For example, the Sigmoid function is used to normalize the global relation features to obtain the weights of each channel in the feature image.

[0118] The image retrieval method provided in this embodiment of the invention has the same effect as the image retrieval method provided in the foregoing embodiments, and will not be described again in this embodiment of the invention.

[0119] Example 4

[0120] Figure 4 This is a schematic diagram of the structure of an image retrieval device provided in Embodiment 4 of the present invention, as shown below. Figure 4 As shown, the device includes:

[0121] The feature extraction module 401 is used to extract image features from the image to be retrieved;

[0122] The channel attention module 402 is used to extract cross-channel dependencies in the image features based on the channel attention mechanism to obtain channel attention features;

[0123] The spatial attention module 403 is used to extract cross-space dependencies in the image features based on the spatial attention mechanism to obtain spatial attention features;

[0124] Feature fusion module 404 is used to fuse the channel attention features and the spatial attention features to obtain fused features;

[0125] The matching module 405 is used to retrieve historical images that match the image to be retrieved from a pre-established historical image historical fusion feature library based on the fusion features.

[0126] In some embodiments of the present invention, the feature extraction module 401 includes:

[0127] The feature extraction submodule is used to perform channel expansion on the image to be retrieved using a feature extraction network, and extract features of multiple different dimensions as image features.

[0128] In some embodiments of the present invention, the feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, and a fourth convolutional block, and the feature extraction submodule includes:

[0129] The first feature extraction unit is used to perform convolution processing on the image to be retrieved in the first convolution block to obtain a 64-channel first feature.

[0130] The second feature extraction unit is used to perform convolution processing on the first feature in the second convolution block to obtain a second feature with 128 channels.

[0131] The third feature extraction unit is used to perform convolution processing on the second feature in the third convolution block to obtain a 256-channel third feature.

[0132] An image feature extraction unit is used to perform convolution processing on the third feature in the fourth convolution block to obtain 512-channel image features.

[0133] In some embodiments of the present invention, the channel attention module 402 includes:

[0134] The local dependency capture submodule is used to capture the local dependencies between K adjacent channels using local convolution to obtain the weights of each channel;

[0135] The first channel attention feature calculation sub-module is used to calculate the product of each channel and its corresponding weight in the image features to obtain the channel attention features.

[0136] In some embodiments of the present invention, the local dependency capture submodule includes:

[0137] The first global average pooling processing unit is used to perform global average pooling processing on the image features to obtain pooled features.

[0138] A one-dimensional convolutional unit is used to perform convolution processing on the pooling features using a one-dimensional convolution with a kernel size of K, to obtain local relation features that characterize the local dependencies between K adjacent channels;

[0139] The first normalization unit is used to normalize the local relation features to obtain the weights of each channel.

[0140] In some embodiments of the present invention, the channel attention module 402 includes:

[0141] The global dependency capture submodule is used to capture the global dependencies between all channels using global convolution and obtain the weights of each channel;

[0142] The second channel attention feature calculation sub-module is used to calculate the product of each channel and its corresponding weight in the image features to obtain the channel attention features.

[0143] In some embodiments of the present invention, the global dependency capture submodule includes:

[0144] The second global average pooling processing unit is used to perform global average pooling processing on the image features to obtain pooled features.

[0145] A fully connected mapping unit is used to perform fully connected mapping on the pooling features using a fully connected layer to obtain global relation features that characterize the global dependencies between all channels;

[0146] The second normalization unit is used to normalize the global relation features to obtain the weights of each channel.

[0147] In some embodiments of the present invention, the spatial attention module 403 includes:

[0148] The spatial correlation calculation submodule is used to calculate the spatial correlation between each element in the image feature and all other elements to obtain a correlation matrix.

[0149] The spatial attention matrix calculation submodule is used to calculate the product of the correlation matrix and the image features to obtain the spatial attention matrix.

[0150] In some embodiments of the present invention, the spatial correlation calculation submodule includes:

[0151] The convolution unit is used to perform convolution operations on the image features in three directions to obtain the Q matrix, K matrix and V matrix respectively;

[0152] A matrix multiplication unit is used to multiply the Q matrix and the K matrix to obtain a correlation weight matrix between the Q matrix and the K matrix;

[0153] The third normalization unit is used to normalize the correlation weight matrix to obtain the correlation weight coefficient matrix of the Q matrix and the K matrix.

[0154] The correlation matrix calculation unit is used to multiply the correlation weight coefficient matrix with the V matrix to obtain a correlation matrix representing the spatial correlation between each element in the image feature and all other elements.

[0155] In some embodiments of the present invention, the base matching module 405 includes:

[0156] The similarity calculation submodule is used to calculate the similarity between the fused feature and each historical fused feature in the historical fused feature library;

[0157] The target similarity determination submodule is used to determine the similarity that is greater than a preset similarity threshold as the target similarity.

[0158] The matching image determination submodule is used to use the historical images corresponding to the historical fusion features of the target similarity as the matching images.

[0159] The image retrieval device described above can execute the image retrieval method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0160] Example 5

[0161] Embodiment 5 of the present invention provides a computer device, Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 5 of the present invention, as shown below. Figure 5 As shown, the computer device includes:

[0162] The mobile terminal includes a processor 501, a memory 502, a communication module 503, an input device 504, and an output device 505; the number of processors 501 in the mobile terminal can be one or more. Figure 5 Taking a processor 501 as an example; the processor 501, memory 502, communication module 503, input device 504, and output device 505 in the mobile terminal can be connected via a bus or other means. Figure 5 Taking a bus connection as an example, the processor 501, memory 502, communication module 503, input device 504, and output device 505 mentioned above can be integrated into a computer device.

[0163] The memory 502, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the module corresponding to the image retrieval method in the above embodiment. The processor 501 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 502, thereby implementing the image retrieval method described above.

[0164] Memory 502 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the microcomputer. Furthermore, memory 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, memory 502 may further include memory remotely located relative to processor 501, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0165] The communication module 503 is used to establish a connection with external devices (such as smart terminals) and to realize data interaction with external devices. The input device 504 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the computer device.

[0166] The computer device provided in this embodiment can execute the image retrieval method provided in any of the above embodiments of the present invention, and has corresponding functions and beneficial effects.

[0167] Example 6

[0168] Embodiment 6 of the present invention provides a storage medium containing computer-executable instructions, on which a computer program is stored. When executed by a processor, the program implements the image retrieval method provided in any of the above embodiments of the present invention, the method comprising:

[0169] Extract image features from the image to be retrieved;

[0170] Based on the channel attention mechanism, cross-channel dependencies in the image features are extracted to obtain channel attention features;

[0171] Based on the spatial attention mechanism, cross-space dependencies in the image features are extracted to obtain spatial attention features;

[0172] The channel attention features and the spatial attention features are fused to obtain the fused features;

[0173] Based on the fusion features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fusion feature library.

[0174] It should be noted that the embodiments of the apparatus, device, and storage medium are basically similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the description of the method embodiments.

[0175] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a robot, personal computer, server, or network device, etc.) to execute the image retrieval method described in any embodiment of the present invention.

[0176] It is worth noting that the various modules, sub-modules, and units included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as they can achieve the corresponding functions; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0177] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution device. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0178] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0179] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. An image retrieval method, characterized in that, include: Image features are extracted from the image to be retrieved, wherein the image to be retrieved is a medical image; Based on the channel attention mechanism, cross-channel dependencies in the image features are extracted to obtain channel attention features; Based on the spatial attention mechanism, cross-space dependencies in the image features are extracted to obtain spatial attention features; The channel attention features and the spatial attention features are fused to obtain the fused features; Based on the fusion features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fusion feature library. The step of extracting cross-channel dependencies in the image features based on the channel attention mechanism to obtain channel attention features includes: The image features are subjected to global average pooling to obtain pooled features; The pooling features are convolved using a one-dimensional convolution with a kernel size of K to obtain local relation features that characterize the local dependencies between K adjacent channels. The local relation features are normalized to obtain the weights of each channel; The channel attention features are obtained by calculating the product of each channel and its corresponding weight in the image features.

2. The image retrieval method according to claim 1, characterized in that, Image features are extracted from the image to be retrieved, including: A feature extraction network is used to perform channel expansion on the image to be retrieved, and multiple features of different dimensions are extracted as image features.

3. The image retrieval method according to claim 2, characterized in that, The feature extraction network includes a first convolutional block, a second convolutional block, a third convolutional block, and a fourth convolutional block. The feature extraction network performs channel expansion on the image to be retrieved, extracting multiple features of different dimensions as image features, including: The image to be retrieved is convolved in the first convolutional block to obtain a first feature with 64 channels. The first feature is convolved in the second convolution block to obtain a second feature with 128 channels. The second feature is convolved in the third convolution block to obtain a third feature with 256 channels; The third feature is convolved in the fourth convolution block to obtain a 512-channel image feature.

4. The image retrieval method according to any one of claims 1-3, characterized in that, Based on the channel attention mechanism, cross-channel dependencies in the image features are extracted to obtain channel attention features, which also include: Global convolution is used to capture the global dependencies between all channels and obtain the weights of each channel; The channel attention features are obtained by calculating the product of each channel and its corresponding weight in the image features.

5. The image retrieval method according to claim 4, characterized in that, Global convolution is used to capture the global dependencies between all channels, and the weights of each channel are obtained, including: The image features are subjected to global average pooling to obtain pooled features; A fully connected layer is used to perform a fully connected mapping on the pooling features to obtain global relation features that characterize the global dependencies between all channels; The global relational features are normalized to obtain the weights of each channel.

6. The image retrieval method according to any one of claims 1-3 and 5, characterized in that, Based on the spatial attention mechanism, cross-spatial dependencies in the image features are extracted to obtain spatial attention features, including: Calculate the spatial correlation between each element in the image feature and all other elements to obtain the correlation matrix; The spatial attention matrix is ​​obtained by multiplying the correlation matrix with the image features.

7. The image retrieval method according to claim 6, characterized in that, Calculate the spatial correlation between each element in the image features and all other elements to obtain a correlation matrix, including: The image features are convolved in three directions to obtain the Q matrix, K matrix and V matrix respectively; Multiply the Q matrix and the K matrix to obtain the correlation weight matrix between the Q matrix and the K matrix; The correlation weight matrix is ​​normalized to obtain the correlation weight coefficient matrix of the Q matrix and the K matrix; Multiplying the correlation weight coefficient matrix by the V matrix yields a correlation matrix representing the spatial correlation between each element in the image feature and all other elements.

8. The image retrieval method according to any one of claims 1-3, 5, and 7, characterized in that, Based on the fusion features, historical images matching the image to be retrieved are retrieved from a pre-established historical image fusion feature library, including: Calculate the similarity between the fusion feature and each historical fusion feature in the historical fusion feature library; The similarity score that is greater than a preset similarity threshold is defined as the target similarity score. The historical image corresponding to the historical fusion feature corresponding to the target similarity is used as the matching image.

9. An image retrieval device, characterized in that, include: The feature extraction module is used to extract image features from the image to be retrieved, wherein the image to be retrieved is a medical image; The channel attention module is used to extract cross-channel dependencies in the image features based on the channel attention mechanism to obtain channel attention features; The spatial attention module is used to extract cross-space dependencies in the image features based on the spatial attention mechanism to obtain spatial attention features; The feature fusion module is used to fuse the channel attention features and the spatial attention features to obtain fused features; The matching module is used to retrieve historical images that match the image to be retrieved from a pre-established historical image historical fusion feature library based on the fusion features; The channel attention module includes: The local dependency capture submodule includes: a first global average pooling processing unit, used to perform global average pooling processing on the image features to obtain pooled features; a one-dimensional convolution unit, used to perform convolution processing on the pooled features using a one-dimensional convolution with a kernel size of K to obtain local relation features representing the local dependencies between K adjacent channels; and a first normalization unit, used to normalize the local relation features to obtain the weights of each channel. The first channel attention feature calculation submodule is used to calculate the product of each channel and its corresponding weight in the image features to obtain the channel attention features.

10. A computer device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image retrieval method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the image retrieval method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Image feature extraction method based on joint attention mechanism

    CN112766279A

  • Convolutional neural network based on four-branch attention mechanism, and image segmentation method

    CN112949838A

  • Image super-resolution reconstruction method based on attention mechanism and dual-channel network

    CN113362223A