Image label generation method and device based on feature fusion, equipment and medium
Through multi-scale feature extraction, dynamic adjustment of category attention information, and fusion of global and local features, the problem of insufficient image segmentation accuracy and adaptability in the prior art is solved, and a more efficient image segmentation effect is achieved, which is suitable for multiple application fields.
Patent Information
- Application Number
- CN202510141900.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-08
AI Technical Summary
When processing images with large changes within classes and small changes between classes, the prior art lacks effective extraction of category differences information and fusion of global and local features, resulting in insufficient image segmentation accuracy and adaptability.
By acquiring image data, multi-scale feature extraction, extracting category attention information and building a category matrix, generating gated signals to adjust the feature map, extracting global and local category features, and fusing them to generate category attention guide feature maps, and finally generating category label maps through the feature fusion head module.
It improves the accuracy and adaptability of image segmentation, and can more effectively identify target areas with large changes within classes and small differences between classes. It is suitable for medical image analysis, disaster assessment and financial claims and other fields.
Smart Images

Figure CN119992272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and medical health, and in particular to a method, device, equipment and storage medium for generating image labels based on feature fusion. Background Art
[0002] With the continuous advancement of deep learning technology, the Transformer architecture has been widely used in semantic segmentation tasks due to its powerful self-attention mechanism. Semantic segmentation is an important task in computer vision, which aims to assign category labels to each pixel in the image, such as buildings, roads, vegetation, etc. This task can be regarded as a pixel-level classification problem. However, most of the existing semantic segmentation methods are designed for natural scene images, and it is difficult to effectively deal with the unique properties of remote sensing images, especially in scenes with large intra-class changes and small inter-class changes. The performance is poor, resulting in limited accuracy and generalization ability of image segmentation. Therefore, in practical applications, the existing technology has many shortcomings in image processing tasks in the fields of healthcare and finance.
[0003] In the field of medical health, semantic segmentation technology is often used for automatic recognition and analysis of medical images, such as the annotation of lesion areas in CT, MRI, X-ray and other images. Existing technologies mainly rely on convolutional neural networks (CNN) to extract image features, but due to the variable morphology and blurred boundaries of lesion areas in medical images, it is difficult to accurately identify lesion areas based on local features alone. Medical images often show large intra-class changes and small inter-class changes, that is, the same lesion may present a variety of forms, and the boundary between the lesion area and normal tissue is not obvious. In addition, the existing segmentation methods lack an effective combination of global information and local features, and cannot identify the scope of the lesion from a global perspective. It is also difficult to take into account the changes in details in the image. These deficiencies make it difficult for existing image segmentation technologies to meet the needs of clinical diagnosis for accurate recognition and real-time performance.
[0004] In the financial field, image segmentation technology is widely used in disaster assessment and asset claims scenarios. For example, in the process of claims settlement after natural disasters, insurance companies often use images of disaster sites taken by drones to quickly identify targets such as buildings, roads, and vegetation in the disaster area. However, the existing technology has problems of low recognition accuracy and large errors in practical applications, and it is difficult to accurately extract detailed features of the disaster area. Image data of disaster scenes are often affected by factors such as occlusion, lighting changes, and environmental interference, resulting in the inability of traditional segmentation models to accurately identify target categories in complex scenes. In addition, target objects in disaster scenes may present a multi-scale distribution, including both large-scale disaster areas and small-scale detail changes. Existing technologies are insufficient in processing these multi-scale features and cannot effectively extract complete disaster information, affecting the insurance company's claims efficiency and the accuracy of loss assessment. Summary of the invention
[0005] The main purpose of the present invention is to provide a method, device, equipment and storage medium for generating image labels based on feature fusion, aiming to solve the technical problem that the existing technology lacks effective extraction of category difference information and fusion of global and local features when processing images with large intra-class changes and small inter-class changes, resulting in insufficient image segmentation accuracy and adaptability.
[0006] To achieve the above object, the present invention provides an image label generation method based on feature fusion, comprising:
[0007] Acquire image data, and perform multi-scale local feature extraction on the image data through a multi-scale feature extraction module to generate a multi-scale local feature map;
[0008] Extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information;
[0009] Generate a gating signal through a gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map;
[0010] Extract global category features from the gated category feature map through a global category attention guidance module;
[0011] Extract local category features from the gated category feature map through the local category attention guidance module;
[0012] Fusion of the global category features with the local category features to generate a category attention guided feature map;
[0013] The category attention guided feature map is fused with the multi-scale local feature map, and the fused feature map is weighted through the feature fusion head module to generate the final fused feature map;
[0014] Generate a category label map according to the final fused feature map.
[0015] Furthermore, to achieve the above object, the present invention provides an image label generation device based on feature fusion, comprising:
[0016] A multi-scale feature extraction module is used to obtain image data, perform multi-scale local feature extraction on the image data through the multi-scale feature extraction module, and generate a multi-scale local feature map;
[0017] A category attention extraction module, used to extract category attention information from the multi-scale local feature map, and construct a category matrix based on category weights and category indexes in the category attention information;
[0018] A gating unit module, used to generate a gating signal through the gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map;
[0019] A global category attention guidance module, used for extracting global category features from the gated category feature map through the global category attention guidance module;
[0020] A local category attention guidance module, used for extracting local category features from the gated category feature map through the local category attention guidance module;
[0021] A category attention fusion module, used for fusing the global category feature with the local category feature to generate a category attention guided feature map;
[0022] The feature fusion module is used to fuse the category attention guided feature map with the multi-scale local feature map, and perform weighted processing on the fused feature map through the feature fusion head module to generate the final fused feature map;
[0023] The label generation module is used to generate a category label map according to the final fusion feature map.
[0024] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a feature fusion-based image label generation program stored in the memory and executable on the processor, wherein the feature fusion-based image label generation program, when executed by the processor, implements the steps of the feature fusion-based image label generation method as described above.
[0025] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which is stored a feature fusion-based image label generation program, and when the feature fusion-based image label generation program is executed by a processor, the steps of the feature fusion-based image label generation method as described above are implemented.
[0026] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and medical health, and discloses a method for generating image labels based on feature fusion, including: acquiring image data, and generating a multi-scale local feature map through multi-scale feature extraction, extracting category attention information and constructing a category matrix; generating a gating signal through a gating unit, adjusting the feature map, extracting global category features and local category features, and fusing the two to generate a category attention guided feature map; weighting the fused feature map to finally generate a category label map. The present invention improves the accuracy and adaptability of image segmentation through multi-scale feature extraction, dynamic adjustment of category attention information, and fusion of global and local features, and can more effectively identify target areas with large intra-class changes and small inter-class differences, and is suitable for fields such as medical image analysis, disaster assessment, and financial claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0028] Figure 1 A schematic diagram of an application environment of a method for generating image labels based on feature fusion in an embodiment of the present invention;
[0029] Figure 2 It is a flow chart of an embodiment of a method for generating image labels based on feature fusion according to the present invention;
[0030] Figure 3 This is a functional module diagram of a preferred embodiment of the image label generation device based on feature fusion of the present invention;
[0031] Figure 4 A schematic diagram of the structure of a computer device in one embodiment of the present invention;
[0032] Figure 5 FIG. 4 is another schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0033] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0034] The image label generation method based on feature fusion provided by the embodiment of the present invention can be applied in the following aspects: Figure 1In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain image data through the user terminal, and generate a multi-scale local feature map through multi-scale feature extraction, extract category attention information and construct a category matrix; generate a gating signal through a gating unit, adjust the feature map, extract global category features and local category features, and fuse the two to generate a category attention guidance feature map; weight the fused feature map to finally generate a category label map. The present invention improves the accuracy and adaptability of image segmentation through multi-scale feature extraction, dynamic adjustment of category attention information and fusion of global and local features, and can more effectively identify target areas with large intra-class changes and small inter-class differences, and is suitable for medical image analysis, disaster assessment and financial claims and other fields. Among them, the user terminal can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0035] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a method for generating image tags based on feature fusion provided by the present invention. It should be noted that although a logical sequence is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0036] like Figure 2 As shown, the image label generation method based on feature fusion proposed by the present invention includes the following steps:
[0037] S10, acquiring image data, and performing multi-scale local feature extraction on the image data through a multi-scale feature extraction module to generate a multi-scale local feature map
[0038] In this embodiment, the image data can come from a variety of sources, including but not limited to satellite remote sensing, drone photography, medical imaging equipment (such as CT, MRI), surveillance cameras, etc. During the implementation process, appropriate equipment can be selected according to the actual application scenario. For example, in a disaster assessment scenario, a drone can quickly capture high-resolution image data of the disaster-stricken area; in the field of medical health, medical imaging equipment can obtain image data of the lesion area.
[0039] The acquired image data is usually in RGB three-channel or grayscale. When implemented, the device's own API interface can be used to acquire image data and store the data in common formats such as JPEG, PNG or TIFF. For medical imaging, the data may be stored in DICOM format, which requires special software tools to decode.
[0040] After acquiring the image, basic preprocessing of the data is required to ensure the quality of subsequent analysis. Preprocessing steps may include image denoising, grayscale, resizing, and format conversion. In the implementation process, open source image processing libraries (such as OpenCV and Pillow) can be used to quickly batch process images.
[0041] The collected image data needs to be stored on a local storage device or cloud storage platform. To ensure data security, hierarchical storage and encryption technology can be used. In the field of healthcare, the storage of image data needs to comply with data privacy regulations.
[0042] For example, drones can be used to fly over disaster-affected areas and collect high-resolution surface images. The drones are equipped with camera modules that automatically record and store image data of the disaster area. Parameters such as the drone's flight altitude and camera focal length can be adjusted according to the size of the target area to be assessed. The collected images are wirelessly transmitted to a ground workstation for further processing.
[0043] In medical imaging, CT or MRI equipment will obtain image data of internal tissues or organs of the human body according to preset scanning parameters (such as layer thickness and resolution). It is stored in DICOM format and analyzed in subsequent image processing software. In order to improve image quality, denoising and artifact elimination algorithms can be used to optimize the imaging effect.
[0044] Example description: In the agricultural insurance scenario, insurance companies can use drones to collect image data of damaged farmland. They can quickly extract the areas where crops are damaged, and use the category attention network to identify the extent of damage to different crops, thereby providing insurance companies with more accurate disaster assessment data and assisting in formulating claims settlement plans. Through automated image processing and analysis technology, insurance companies can significantly shorten the claims settlement cycle and improve customer satisfaction.
[0045] In the diagnosis of lung diseases, CT scanners are used to obtain lung imaging data of patients. Abnormal areas in lung images can be quickly extracted, and the type of lesions, such as lung nodules and inflammation, can be identified based on the category attention network. Combining imaging data with other diagnostic and treatment information of patients, doctors can more accurately judge the patient's condition and develop targeted treatment plans, thereby improving treatment outcomes and reducing misdiagnosis rates.
[0046] By acquiring high-quality image data, a reliable data foundation is provided for subsequent multi-scale feature extraction and category attention analysis. It can flexibly adapt to a variety of scenarios, not only can it quickly obtain the image data required for disaster assessment, but also can meet the needs of high-precision imaging in the medical and health field, and improve the accuracy and timeliness of data analysis.
[0047] The core of the multi-scale feature extraction module is to build a pre-trained basic convolutional network model. This network can extract the basic features of image data through convolution operations based on mainstream deep learning models such as ResNet and VGG. The basic convolutional network mainly recognizes the edges, textures, and color distribution of the image to generate a basic feature map. This feature map contains low-level information of the image and provides input for subsequent multi-scale feature extraction.
[0048] Select a suitable basic network architecture, such as ResNet-50, ResNet-101, or VGG-16. Initialize the network using the weights of the pre-trained model. Input the image data into the basic convolutional network to extract the basic feature map.
[0049] The key step of the multi-scale feature extraction module is to perform multi-scale convolution operations on the basic feature map. Multi-scale convolution refers to the use of convolution kernels of different sizes (such as 1×1, 3×3, 5×5, etc.) to extract features from different areas of the image. Small convolution kernels can capture detailed information, and large convolution kernels can capture global context information.
[0050] Convolution operations are performed on the basic feature maps using convolution kernels of different sizes. Convolution kernels of different sizes can be implemented in the form of parallel branches, that is, a convolution branch is set up separately for each convolution kernel. The feature map output by each convolution operation needs to be processed by batch normalization and activation functions (such as ReLU, LeakyReLU) to improve the expressiveness of the features.
[0051] Multi-scale feature extraction can also enhance the recognition of long strip targets by introducing striped convolution. Striped convolution uses convolution kernels with different aspect ratios (such as 1×7, 7×1, 1×11, 11×1, etc.) to extract the features of linear targets. This is of great significance for identifying roads, boundary lines, etc. in images.
[0052] The output feature map of the multi-scale convolution operation is convolved using a long strip convolution kernel. The aspect ratio of the strip convolution kernel can be adjusted according to the application scenario. For example, in remote sensing images, a longer convolution kernel can be used to extract slender targets such as rivers and roads.
[0053] After completing the multi-scale convolution and strip convolution operations, the output feature maps of these convolution operations need to be aggregated to generate the final multi-scale local feature map. The aggregation operation can be done by concatenation or addition to fuse feature maps of different scales together to form a feature map containing information of multiple scales.
[0054] The output feature maps of convolution operations at different scales are concatenated to form a multi-channel feature map. Channel weighting is performed on the concatenated feature maps to adjust the weight of each scale feature map, thereby improving the expression of key features. The final multi-scale local feature map can be further compressed through the convolution layer to reduce the dimension of the feature map and improve computational efficiency.
[0055] For example, local features can be extracted through a CNN-based remote sensing (RS) image multi-scale local feature extraction module (CNN). CNN can be divided into scale and strip feature extraction. CNN introduces the pre-trained deep learning model ResNet in the scale feature extraction step, receives the ResNet features of the corresponding stage, and then enters a branch consisting of 3 convolutional layers (1×1, 3×3 and 5×5) to obtain contextual information of different scales. Batch normalization (BN) and LeakyReLU activation functions appear after each convolutional layer. After extracting the local features at three scales, the local features are aggregated by summing them. The formula can be:
[0056]
[0057] The strip feature extraction step involves 3 sets of strip convolutions (1×7 and 7×1, 1×11 and 11×1, 1×21 and 21×1). Strip convolution has fewer parameters and can provide better strip target recognition results, such as roads and water areas. The above local features are aggregated through the convolution layer (3×3) and fused with multi-scale features using residual connections. ⊕ represents the addition of corresponding positions of the matrix to obtain the CNN output
[0058] X 7×7 =BN(Conv 7×1 (Conv 7×1 (X scale ))
[0059] X 11×11 =BN(Conv 7×1 (Conv 7×1 (X scale ))
[0060] X 21×21 =BN(Conv 7×1 (Conv 7×1 (X scale ))
[0061]
[0062] Among them, X Ris the input feature map, usually a preliminary feature map extracted by a basic convolutional network. Its dimension is X B×C×H×W , B represents the batch size, C represents the number of channels, H represents the height of the image, and W represents the width of the image; Conv 1×1 、Conv 3×3 、Conv 5×5 They represent convolution kernels of different sizes respectively; BN() stands for BatchNormalization, which normalizes the convolution output to speed up the model convergence and improve stability; Indicates concatenation or summation by channel, which is used to integrate multi-scale feature maps; X 7×7 , X 11×11 , X 21×21 They represent feature maps of different scales after strip convolution operations. These feature maps are generated by convolution operations on input feature maps using convolution kernels of different sizes. scale Represents the generated multi-scale feature map, which contains multi-scale local features obtained through 1×1, 3×3 and 5×5 convolution operations.
[0063] Example description: In the financial field, insurance companies can use drones to obtain image data of disaster-stricken areas. Through the multi-scale feature extraction module, the damage of different targets such as houses, roads, and farmland can be identified at the same time. For example, in a flood scene, the multi-scale convolution operation can extract the damage features of houses and farmland, and the strip convolution operation can extract the morphological features of roads or rivers destroyed by floods. The final multi-scale local feature map provides accurate data support for insurance companies to generate damage assessment reports.
[0064] The multi-scale feature extraction module extracts multi-scale local features from image data, which can extract global and local features from the image and enhance the recognition of linear and multi-scale targets. It can improve the accuracy and robustness of image segmentation and has significant application effects in asset damage assessment in the financial field and lesion area identification in the medical field. By extracting multi-scale local feature maps, high-quality input features can be provided for the subsequent category attention network, thereby improving the overall performance of the entire system.
[0065] S20, extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information;
[0066] In this embodiment, the category attention information is the core information extracted from the multi-scale local feature map, which aims to obtain the probability of each pixel belonging to different categories. The multi-scale local feature map contains multi-level feature information of the image, and by analyzing these features, the category probability distribution of each pixel is generated.
[0067] The generation process of class attention information can be based on convolutional neural networks or multi-layer perceptrons (MLPs) to map the input feature map into the probability value of each class. Class attention information contains two types of information: class weight and class index. Class weight is used to indicate the degree of belonging of each pixel to each class, and class index is used to identify the most likely class of each pixel.
[0068] The multi-scale local feature map is input into the classification network, and the category attention information is generated by calculating the category probability distribution of each pixel. In the implementation process, the Softmax function can be used to normalize the category probability to ensure that the sum of the probabilities of each pixel in all categories is 1.
[0069] The category matrix is an important data structure for storing category features. By extracting the category weight and category index from the category attention information, a two-dimensional matrix containing category features can be constructed. The category weight is the belonging probability value of each pixel, indicating the possibility that each pixel belongs to each category.
[0070] The category index is the predicted category label for each pixel, indicating the most likely category for each pixel. By combining the category weight map and the category index map, a category matrix can be constructed for subsequent gating operations or attention mechanism calculations.
[0071] The class weight and class index are extracted from the class attention information to generate a class weight map and a class index map respectively. The class weight map and the class index map are combined to generate a class matrix. The class matrix is used to generate a gating signal in the subsequent steps to focus on and optimize the class features of the feature map.
[0072] For example, the feature map X output from CNN CNN The process of extracting category attention information includes:
[0073]
[0074] Where B is the batch size, C is the number of channels, H is the height, and W is the width. The category information is enhanced through batch normalization (BN) and multi-layer perceptron (MLP) layers.
[0075] Then, a 1×1 convolutional layer is used to change the dimension of the feature map to the dataset category number N. c Finally, the category matrix is obtained through the Softmax and Argmax functions. This process is similar to the output stage of the semantic segmentation task.
[0076] XC =Argmax(Softmax(Conv 1×1 (X MLP )))
[0077] The formula represents the application of a 1×1 convolutional layer, a Softmax function, and an Argmax operation to extract the category weight and category index of each pixel from the feature map. Softmax calculates the category weight, Argmax extracts the category index, and finally generates the category matrix X C .
[0078] Example description: In the construction of customer portraits in the financial field, the behavior data of each customer can be regarded as a multi-scale feature map of the image, and the behavior features of different categories can be extracted through category attention information. For example, the behavior patterns of "high-risk", "medium-risk" and "low-risk" customers can be extracted, and through the category matrix construction, attention and management optimization of different customer categories can be achieved.
[0079] In medical image analysis, local feature maps of different anatomical structures can be obtained through multi-scale feature extraction modules. The extraction of category attention information can be used to identify the category labels of lesion areas (such as tumors, normal tissues, inflammation, etc.). By constructing a category matrix, category information can be effectively integrated to provide doctors with accurate lesion localization and classification results, thereby improving the accuracy of diagnosis.
[0080] By extracting category attention information from multi-scale local feature maps and constructing a category matrix based on category weights and category indices, effective representation of category features is achieved. This category matrix construction method can more accurately capture the category differences of each pixel in the image, thereby providing support for subsequent feature adjustment and attention calculation, and improving the model's attention to category features and recognition accuracy.
[0081] S30, generating a gating signal through a gating unit, adjusting the multi-scale local feature map through the gating signal, and performing channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map;
[0082] In this embodiment, the function of the gating unit is to generate a dynamic gating signal, which is used to control the transmission of information flow and the dynamic adjustment of feature weights. The input of the gating unit is the feature map of the current layer and the feature map of the previous layer. For this step, the feature map of the current layer is a multi-scale local feature map, and the feature map of the previous layer is the output of the previous stage of the network. The internal structure of the gating unit includes two convolution operations, which perform convolution calculations on the feature map of the current layer and the feature map of the previous layer respectively to obtain a preliminary adjustment signal. After adding these two adjustment signals, a nonlinear transformation is performed through an activation function (such as a Sigmoid function) to generate a dynamic gating signal.
[0083] In the actual implementation process, the input multi-scale local feature map is first processed through the convolution layer to generate a preliminary feature map; secondly, the previous layer feature map is also processed through the convolution operation to obtain another feature map. After adding the two feature maps element by element, the activation function is applied to perform a nonlinear transformation so that the generated gating signal varies between 0 and 1. This nonlinear processing method enables the network to control the transmission strength of the information flow.
[0084] After the gating signal is generated, the multi-scale local feature map of the current layer needs to be adjusted. The adjustment process includes two parts: channel weighting and spatial weighting. The gating signal acts as a weight coefficient, directly acting on the feature value of each channel, and weighted adjustment is performed according to the feature weight information of different pixel positions, so that the network can better focus on pixel areas with feature differences when processing different categories of information.
[0085] In the specific implementation process, channel weighting processing is to multiply the weight of each channel of the gated signal by the corresponding feature map channel value, and spatial weighting processing is to perform pixel-by-pixel multiplication on the weight information of each pixel position. The adjusted multi-scale local feature map can more accurately represent the category information in the subsequent network calculations and improve the classification performance.
[0086] The adjusted multi-scale local feature map needs to be further combined with the category weight information in the category matrix for channel-by-channel weighting. The category matrix is constructed from the category attention information, which contains the category weight corresponding to each pixel. The process of channel-by-channel weighting is to perform a weighted operation on the value of each channel of the adjusted feature map according to the weight values of different categories in the category matrix, thereby enhancing the ability to distinguish category features.
[0087] In the implementation process, the category weight information of each channel is first extracted from the category matrix, and then these weight values are element-by-element multiplied with the corresponding channel values of the adjusted multi-scale local feature map to generate a new feature map. This channel-by-channel weighted operation enables the network to strengthen or suppress different channel features according to the category information, thereby more accurately representing the category features.
[0088] After the dynamic gating signal is adjusted and the channel-by-channel weighted processing of the category matrix is performed, the gated category feature map is finally generated. This feature map contains more accurate category feature information and can provide input for the subsequent global category attention guidance module and local category attention guidance module.
[0089] The generation process of the gated category feature map includes two main steps: the first step is to adjust the feature map using the dynamic gating signal so that the network can dynamically adjust the transmission of feature information; the second step is to use the category weight information of the category matrix to weight the feature map channel by channel so that the network can better focus on important category features.
[0090] For example, in order to make more precise category feature adjustments to multi-scale local feature maps, a gate unit needs to be introduced. The core function of this unit is to generate dynamic gating signals through the information of the category matrix to adapt to the weighted adjustment of different category features.
[0091] Specifically, the core calculation process of the gate control unit is:
[0092] G = σ(Conv(X R )+Conv(X prev ))
[0093] Where: G represents the generated gated signal matrix; σ is the Sigmoid function, and Conv represents the convolution operation. Assume X prev is the output feature of the previous layer, and its dimension is the same as X R same.
[0094] Use the gating signal G to adjust the class matrix CM:
[0095]
[0096] The gating signal G is a dynamic weight signal generated from the feature map by the gating unit. The category matrix CM is constructed after extracting the category attention information from the multi-scale local feature map. When the gating signal G is applied to the category matrix CM, the category weights can be dynamically adjusted, thereby enhancing the category differences in the feature map.
[0097] Where CM' represents the adjusted category matrix, G is the gating signal, and ⊙ represents the element-by-element multiplication operation. Each element in the gating signal G is multiplied pixel by pixel with the corresponding element in the category matrix CM, thereby adjusting the weight value in the category matrix according to the dynamic change of the category weight.
[0098] X′ R =CM′·X R
[0099] Among them, X' R represents the adjusted feature map, that is, the gated category feature map, CM' is the adjusted category matrix, X R is the original multi-scale local feature map. This step is completed through matrix multiplication. According to the adjusted category weights, the category features in the multi-scale local feature map are weighted channel by channel to generate a more accurate category feature map.
[0100] By introducing the gating unit and the channel-by-channel weighted processing of the category matrix, the dynamic adjustment of the feature map and the refined representation of the category features are realized. It can more accurately extract and strengthen the category feature information, especially when processing image data with large intra-class differences and small inter-class differences, thereby improving the accuracy and robustness of image classification, and is suitable for many fields such as medical health, finance, etc.
[0101] S40, extracting global category features from the gated category feature map through a global category attention guidance module;
[0102] In this embodiment, in the global category attention guidance module, the gated category feature map is first windowed. The window division operation is intended to decompose the entire feature map into multiple small window feature blocks, each of which contains information about a local area in the feature map. The size of the window division can be set according to specific application requirements, such as using a fixed-size k×k window or dynamically adjusting the window size.
[0103] In the implementation process, a sliding window method is used to slide the feature map in rows and columns to extract each local area in the feature map to form non-overlapping window feature blocks. This operation helps to reduce the computational complexity of global feature extraction and retain local structural information.
[0104] In order to better adapt to the subsequent global category attention mechanism, each window feature block needs to be expanded into a one-dimensional vector form through a linear transformation operation. This process can be achieved through matrix flattening and a fully connected layer.
[0105] The flattening operation converts the two-dimensional matrix form of each window feature block into a one-dimensional vector, which constitutes the input sequence of the attention mechanism. Subsequently, the dimension of the feature is adjusted through linear transformation to match the dimension of the query vector, key vector, and value vector, which is convenient for subsequent attention calculation.
[0106] When constructing global category attention, the linearly expanded input sequence needs to be mapped into query vectors, key vectors, and value vectors. The query vector is used to determine the focus direction of the feature block in the current window, the key vector is used to represent the association between the feature blocks, and the value vector contains the feature representation information of each feature block.
[0107] The mapping process is completed using convolution operations or linear layers to transform each input sequence into a query vector, a key vector, and a value vector. This operation ensures that the model can calculate the correlation between window feature blocks through different feature representations.
[0108] The attention score function is constructed using the scaled dot product attention mechanism of the query vector Q, the key vector K, and the value vector V. By calculating the attention score function, the weight matrix of category attention can be generated for the extraction of global category features. First, the attention score function is calculated:
[0109]
[0110] Among them, ⊕ represents the element-level addition operation, and CA represents the category weight map.
[0111] Finally, the calculated attention output vector is concatenated to restore it to the same dimension as the original feature map. After concatenation, the concatenated feature vector is linearly transformed to generate the final global category feature.
[0112] The linear transformation process can be implemented through convolution operations or fully connected layers to ensure that the output global category features maintain the same shape and dimension as the input feature map. The generated global category features contain the category information and global context information of different regions in the image, which provides a richer global feature representation for subsequent feature fusion.
[0113] For example, applying the attention weight matrix to the value vector V to extract global category features:
[0114]
[0115] in, represents the element-wise product operation, X GCA Represents the global category feature.
[0116] The global category features are extracted from the gated category feature map through the global category attention guidance module, realizing the global modeling capability of category information in the image. The model's attention to and ability to distinguish global category information are effectively improved. By incorporating category weight information into the attention scoring function, the module can dynamically adjust the weights of each window feature block, making the model more robust and accurate in identifying different categories.
[0117] S50, extracting local category features from the gated category feature map through a local category attention guidance module;
[0118] In this embodiment, the gated category feature map is generated by the gated unit in the previous step, and contains the fusion information of global and category features. In order to better extract local category features of different scales, a multi-scale convolution operation is required, that is, applying convolution kernels of different sizes to the same feature map, such as 3×3, 5×5, 7×7, etc. This multi-scale convolution operation can capture feature information of different levels and scales in the image, thereby improving the perception of local category features in the image.
[0119] Construct a multi-scale convolution module and use convolution kernels of different sizes to perform convolution operations on the gated category feature map. Perform batch normalization and activation processing on the output results of each convolution operation to ensure that the feature value is within a reasonable range and avoid gradient explosion or disappearance. Finally, splice or superimpose the results of the multi-scale convolution operation to generate a preliminary local category feature map.
[0120] In remote sensing images, linear targets (such as roads, rivers, etc.) are common category features, but these linear targets may not be effectively extracted by traditional square convolution kernels. Therefore, the strip convolution operation is introduced to extract linear target features specifically through convolution kernels of different lengths and widths. For example, 1×7, 1×11, and 1×21 convolution kernels are used to extract horizontal linear features, and 7×1, 11×1, and 21×1 convolution kernels are used to extract vertical linear features.
[0121] A strip convolution operation is applied to the local category feature map to extract linear target features in the horizontal and vertical directions respectively; the convolution results in the horizontal and vertical directions are superimposed to generate a feature map containing linear target information.
[0122] In order to further enhance the local category features, it is necessary to perform channel weighting on the results of the convolution operation. The core idea of channel weighting is to dynamically adjust the weights of each channel in the feature map according to the distribution of each category in the image, emphasizing the features of important categories and suppressing irrelevant information.
[0123] Calculate the weight value of each channel and weight it according to the importance of the category feature; multiply each channel of the feature map by the corresponding weight value to obtain the weighted feature map.
[0124] Depthwise separable convolution is an efficient convolution operation that can significantly reduce the number of parameters of the convolution operation without losing the ability to extract features. It decomposes the standard convolution into depthwise convolution and pointwise convolution, and extracts features in the spatial dimension and channel dimension respectively.
[0125] A depth-wise separable convolution operation is applied to the local category feature map to extract more refined local category features; the feature information of different channels is fused together through point convolution to generate the final local category feature map.
[0126] In order to ensure that the value range of the feature map is within a reasonable range and enhance the nonlinear expression ability of the network, it is necessary to activate and normalize the local category features. The activation function can be ReLU or Leaky ReLU, and the normalization operation can be batch normalization.
[0127] An activation function is applied to the local category feature map to enhance the nonlinear expression ability of the network; the activated feature map is normalized to avoid gradient explosion or gradient disappearance; the final output local category feature map contains rich local category information, providing important input for subsequent feature fusion.
[0128] For example, the Local Category Attention Guided Module (LCAG) is a core module for extracting local category features. It mainly achieves accurate extraction of local category information by performing dot product attention calculation and convolution operation on the gated category feature map. The whole process can be divided into the following steps:
[0129]
[0130] This formula represents the local attention weight matrix Atten calculated by the local category attention mechanism. LCA , where: CA is the category feature in the category matrix; V is the value vector, indicating the actual value of the category feature; Represents the dot product operation of a tensor; softmax is a normalization function used to normalize the weight value to the range [0,1].
[0131] Compute local category features:
[0132]
[0133] This formula represents the local attention weight matrix Atten LCA And the gated category feature map X CNN The local category feature calculation process, where: X CNN is the gated category feature map of the input; Represents element-wise product operation; Represents a residual connection (i.e., adding the input feature map directly to the output feature map).
[0134] Convolution and normalization of local category features:
[0135] X LCA =BN(Conv 3×3 (X′ LCA ))
[0136] This formula represents the local category feature map X' LCA Perform a 3×3 convolution operation and normalize the convolution result through batch normalization (BN), where: Conv 3×3 Indicates that a convolution operation is performed using a 3×3 convolution kernel; BN indicates a batch normalization operation.
[0137] Example description: In medical image analysis, the local category attention guidance module can be used to accurately extract local features of the lesion area. For example, in lung CT images, the local category attention mechanism can effectively distinguish the boundaries between normal tissue and lesion tissue, helping doctors to more accurately identify the location and type of lesions, thereby improving the accuracy and efficiency of diagnosis.
[0138] In financial risk assessment, the local category attention guidance module can be used to accurately identify key features in image data. For example, in drone image analysis, the local category attention mechanism can accurately identify the types and growth conditions of crops in farmland, thereby providing agricultural insurance companies with a more accurate basis for risk assessment. At the same time, the module can also help banks identify specific information on assets such as real estate and land when reviewing mortgage assets, thereby improving the accuracy and efficiency of the review.
[0139] By processing the gated category feature map through the local category attention guidance module, the local category features in the image can be effectively extracted, especially for category information extraction with significant boundaries or linear targets. The network can dynamically adjust the feature weights during the local feature extraction process, thereby improving the perception of local category information and classification accuracy.
[0140] S60, fusing the global category feature with the local category feature to generate a category attention guiding feature map;
[0141] In this embodiment, after the global category features and the local category features are generated, they need to be fused to generate a more comprehensive category attention guiding feature map.
[0142] First, feature fusion is completed by weighted summing the global category feature and the local category feature matrix element by element. This process ensures that the category feature map contains both global context information and local detail information, thereby improving the category representation ability and classification accuracy.
[0143] The mathematical formula for the fusion operation is as follows:
[0144]
[0145] Where: X CAGM represents the fused category attention guided feature map; X GCArepresents the global category feature map generated by the global category attention guidance module; X LCA represents the local category feature map generated by the local category attention guidance module; Represents an element-wise addition operation (also known as Hadamard addition).
[0146] After the fusion operation is completed, the generated category attention guided feature map has the following characteristics:
[0147] It can capture global information and local detail information in the image; it improves the model's recognition accuracy of category features, especially under complex backgrounds or uneven lighting conditions.
[0148] By fusing global and local category features, a category attention guided feature map is generated, thereby improving the feature representation ability and classification performance of the model. In practical applications, this scheme can effectively improve the recognition accuracy in complex image scenes and has wide applicability.
[0149] S70, fusing the category attention guided feature map with the multi-scale local feature map, performing weighted processing on the fused feature map through a feature fusion head module, and generating a final fused feature map;
[0150] In this embodiment, after the category attention guided feature map and the multi-scale local feature map are generated, the two feature maps need to be fused, and the fused feature map is weighted by the feature fusion head module to generate the final fused feature map.
[0151] Category attention guided feature map (denoted as X CAGM ) contains the feature representation of global and local information of different categories in the image, and the multi-scale local feature map (denoted as X CNN ) contains local detail features of different scales in the image.
[0152] In order to better utilize the advantages of the two feature maps, the two feature maps are fused. The fusion operation is completed by element-by-element addition to ensure that the feature value of each pixel contains category information and local detail information. The mathematical representation of fusion is as follows:
[0153] X fusion =X CAGM ⊕X CNN
[0154] Where: X fusion Represents the fused feature map; X CAGM represents the category attention guided feature map; X CNN Represents a multi-scale local feature map; ⊕ represents an element-by-element addition operation.
[0155] This operation ensures that the fused feature map has both global category information and local detail information, providing high-quality input data for subsequent weighted processing.
[0156] Fusion feature map X fusion After generation, it needs to be weighted by the feature fusion head module. The feature fusion head module is a convolutional network structure that can assign different weights to different channels and spatial positions of the fused feature map, thereby achieving a more refined feature representation.
[0157] The mathematical representation of weighted processing is:
[0158] X FFH =Conv 1×1 (BN(Conv 3×3 (X fusion )))
[0159] Where: X FFH Represents the final fused feature map after weighted processing; Conv 3×3 represents a 3×3 convolution operation, which is used to extract local context information; BN represents a batch normalization operation, which is used to stabilize the training process; Conv 1×1 Represents a 1×1 convolution operation, which is used for channel weighted processing.
[0160] Through the weighted processing of the feature fusion head module, the model can dynamically adjust each channel of the feature map according to the importance of different feature channels, thereby generating a more discriminative feature map.
[0161] After being processed by the feature fusion head module, the final fusion feature map X FFH It contains a fusion representation of global, local and category information. This feature map can more comprehensively describe the target category characteristics in the image and provide high-quality input data for the subsequent category label map generation step.
[0162] In generating the final fusion feature map X FFH Finally, in order to optimize the performance of the model, it is necessary to introduce a loss function for supervision and optimization. The design of the loss function covers primary loss, scale auxiliary loss, category auxiliary loss, and disaster-related loss. These loss terms can work together in the network training process to improve the model's expressiveness and prediction accuracy.
[0163] Primary loss
[0164] The primary loss is combined with the cross entropy loss and Dice loss For the final fusion feature map X FFH The generated category label map is optimized:
[0165]
[0166] in:
[0167] Dice loss:
[0168]
[0169] Cross Entropy Loss:
[0170]
[0171] N represents the number of samples, C represents the number of categories, represents the predicted probability that sample n belongs to category c, Indicates the true label that sample n belongs to category c.
[0172] Scale-assisted loss
[0173] The scale-assisted loss is achieved by fusing multi-scale features X scale , constraining the expressive power of multi-scale features:
[0174]
[0175] X scale A feature map representing the fusion of multi-scale features; Represents the k-th order feature map of the category attention guidance module; UP 2× UP 4× Represent the upsampling operations of two and four times, respectively, which are used to adjust the feature maps to the same resolution.
[0176] Class-assisted loss
[0177] In the calculation of the category auxiliary loss, in order to improve the supervision effect of the category features, a multi-stage fusion strategy is adopted. First, the feature maps of each stage output by the category attention module are To ensure that these feature maps have consistent resolution, an upsampling operation UP is used. 2× , U.P. 4× The feature maps at different stages are adjusted to a unified resolution.
[0178] Then, the adjusted category feature maps are fused element by element to generate the final category auxiliary feature X category The fusion feature contains multi-level category information, which can more comprehensively describe the category distribution in the image and provide high-quality input for the calculation of category-assisted loss. The formula is:
[0179]
[0180] Disaster-related losses
[0181]
[0182] represents disaster-related losses, measuring the difference between predicted losses and actual losses; Represents the predicted value of the nth sample for category c; Represents the true value of the nth sample for category c.
[0183] Overall loss function (L):
[0184]
[0185] Represents the overall loss function, which is used for comprehensive optimization of the model; α, β, and γ represent the weight parameters of the loss term, which can be adjusted according to the actual scenario.
[0186] Example description: Disaster scene images taken by drones usually contain multiple categories of features such as buildings, roads, and vegetation. The feature fusion head module can be used to effectively combine local features of different scales in the image data collected by drones (such as cracks and collapse details of damaged buildings) with global features (such as the degree of damage and scope of the overall disaster area). For example, in a flood disaster, images taken by drones may include collapsed houses, flooded roads, and other disaster-stricken areas. These local details can be extracted to identify the scope and extent of various types of damaged areas, and then fused and analyzed with the overall disaster scope to generate a disaster map with high accuracy. These data can not only quickly help insurance companies assess the amount of losses and shorten the claim settlement time, but also provide scientific rescue decision-making basis for governments and rescue departments.
[0187] The final fused feature map is generated by fusing the category attention guided feature map with the multi-scale local feature map and weighting the fused feature map through the feature fusion head module.
[0188] S80, generating a category label map according to the final fused feature map.
[0189] In this embodiment, the final fused feature map integrates global, local and category feature information and has high-dimensional contextual semantic expression capabilities. In order to realize the generation of the category label map, the final fused feature map needs to be further decoded. The main purpose of the decoding operation is to restore the high-dimensional features to the same resolution as the original input image and predict a category label for each pixel.
[0190] The spatial resolution of the final fused feature map is expanded using upsampling operations (such as deconvolution or interpolation). The upsampled feature map will be aligned with the spatial dimensions of the input image while keeping the semantic information of the final fused feature intact.
[0191] The upsampled feature map is fed into a classification layer (such as a 1×1 convolutional layer) to calculate the probability distribution of each pixel corresponding to each category. The core of this step is to generate a category probability distribution map through a pixel-by-pixel multi-category classification operation. The output of the classification layer is normalized using the Softmax function to obtain a category probability distribution map. The Softmax function ensures that the sum of all category probabilities for each pixel is 1, thereby achieving pixel-level category prediction.
[0192] In the category probability distribution map, the category of each pixel is determined by the category with the maximum probability value. The core of this step is to map the continuous probability distribution into a discrete category label map. The Argmax operation is used to extract the category index with the largest probability from the category probability distribution to generate the final category label map. Each pixel in the category label map represents the category corresponding to that position.
[0193] Example description: In the field of medical health, it can be applied to medical image analysis, such as lesion segmentation and tissue recognition tasks. Taking lung CT images as an example, for different lung structures (such as airways, alveoli, blood vessels, etc.) and lesion areas (such as lung nodules, tumors, inflammation, etc.), it can achieve pixel-by-pixel category segmentation, thereby accurately identifying different types of tissues and lesion areas.
[0194] In the specific application process, doctors do not need to manually annotate the image data. Instead, the multi-scale feature extraction module automatically extracts different scale features in the medical image, and focuses on the category features of the specific lesion area through the global category attention guidance module and the local category attention guidance module. For example, when identifying lung nodules, the global feature guidance module can focus on the overall structure of the lungs, while the local feature guidance module can finely identify the boundaries, size and position of the nodules, thereby improving segmentation accuracy and recognition effect.
[0195] In actual deployment, drones or telemedicine equipment can be used to collect CT images of patients. After the image data is input, the system dynamically adjusts the multi-scale local feature map through the gating unit to generate a fused feature map containing global and local category information. Finally, the category label map generated based on the fused feature map can accurately identify different lung structures and lesion areas. For example, the red area identifies lung nodules, the yellow area identifies normal lung tissue, and the blue area identifies vascular pathways. The following beneficial effects can be achieved:
[0196] Improve the accuracy of lesion identification: By fusing global and local category information, the boundaries and internal structures of lung nodules can be identified more accurately.
[0197] Reduce misdiagnosis and missed diagnosis rates: The category attention mechanism effectively distinguishes different tissues and lesion areas, avoiding recognition errors caused by similar categories.
[0198] Automated medical image analysis: From image data input to category label map output, the entire process of lesion area identification is automated, providing doctors with efficient and accurate auxiliary diagnostic tools.
[0199] For example, for a patient at risk of lung nodules, doctors can use automatically generated category label maps to quickly identify suspicious areas in the lungs, and further analyze the nature, size, and change trends of the nodules to develop a more accurate treatment plan.
[0200] Similarly, in the financial field, it can be applied to disaster assessment and insurance claims scenarios, especially for geographic remote sensing image analysis tasks in agricultural insurance and property insurance. For example, after a disaster occurs (such as floods, fires, earthquakes, etc.), insurance companies usually need to quickly assess the losses in the disaster-stricken areas and provide data support for insurance claims. Through the semantic segmentation analysis of remote sensing images, accurate identification of different categories such as buildings, roads, farmland, and forests in the disaster area is achieved, and the category label map of the damaged area is automatically generated to provide a basis for disaster loss assessment.
[0201] In specific applications, geographic remote sensing images of the disaster-stricken area are obtained through drones or satellites, and the image data is input into the model. First, the multi-scale feature extraction module is used to extract different scale features in the remote sensing image, such as identifying the boundaries of buildings, the layout of roads, the distribution of farmland, etc. Then, the feature map is dynamically adjusted and optimized through the gating unit and the category attention guidance module to generate a gated category feature map. This feature map can effectively distinguish different categories in the disaster-stricken area, such as damaged buildings, destroyed roads, flood-covered areas, and unaffected farmland.
[0202] In the process of generating the category label map, the system automatically performs global and local weighted fusion of the category features to ensure that the segmentation results have higher accuracy and robustness. The final output category label map can accurately identify the category and degree of damage of each area. For example, red areas represent completely destroyed buildings, yellow areas represent partially damaged roads, green areas represent unaffected farmland, and blue areas represent areas covered by floods. The following beneficial effects can be achieved:
[0203] Improve the accuracy of disaster loss assessment: The automatically generated category label map covers multi-category information of the disaster area, ensuring that the identification of the damaged area is more comprehensive and accurate.
[0204] Shorten claims settlement time: The automated disaster assessment process significantly reduces manual verification time, allowing insurance companies to complete disaster claims more quickly.
[0205] Reduce assessment costs: There is no need to manually check the disaster situation one by one, which reduces the manpower and time costs of disaster assessment.
[0206] For example, in a flood disaster, an insurance company used drones to obtain remote sensing image data of the affected area, and the system automatically generated a category label map that identified the area of buildings and farmland covered by the flood. The insurance company can quickly assess the area and severity of the damaged area based on the label map, and then provide quick claims services to affected customers.
[0207] By upsampling, classifying and probability mapping the final fused feature map, not only can a category label map consistent with the input image resolution be generated, but each category area can also be accurately divided in pixels.
[0208] The present invention relates to the fields of artificial intelligence technology and medical health, and discloses a method for generating image labels based on feature fusion, including: acquiring image data, and generating a multi-scale local feature map through multi-scale feature extraction, extracting category attention information and constructing a category matrix; generating a gating signal through a gating unit, adjusting the feature map, extracting global category features and local category features, and fusing the two to generate a category attention guidance feature map; weighting the fused feature map to finally generate a category label map. The present invention improves the accuracy and adaptability of image segmentation through multi-scale feature extraction, dynamic adjustment of category attention information, and fusion of global and local features, and can more effectively identify target areas with large intra-class changes and small inter-class differences, and is suitable for fields such as medical image analysis, disaster assessment, and financial claims.
[0209] In one embodiment, after the above S20, the method further includes:
[0210] S301, performing window division on the category label of each pixel in the category matrix to obtain a plurality of window feature matrices;
[0211] S302, extracting category labels from each window feature matrix respectively, and combining the category labels into a category label sequence according to a preset order;
[0212] S303, performing a subtraction operation on the category label sequence corresponding to each window feature matrix to obtain a category difference matrix;
[0213] S304, adding the position index to the category difference matrix to obtain a category difference matrix containing position information;
[0214] S305, flattening the category difference matrix containing the position information and offsetting the negative values to construct a relative category index;
[0215] S306: Adjust the category matrix based on the relative category index to optimize the feature representation of the category matrix.
[0216] In this embodiment, after extracting the category attention information from the multi-scale local feature map and constructing the category matrix based on the category weight and category index in the category attention information, in order to further optimize the feature representation of the category matrix, it is necessary to perform window division on the category label of each pixel in the category matrix to obtain more refined category feature information. The window division operation is intended to divide the category matrix into multiple small feature regions, each region being a window feature matrix. The division process can be set according to a preset window size and sliding step size to ensure that the category information of each pixel in the category matrix is locally analyzed. This window division process can effectively capture the spatial locality of the category feature, making the representation of the category feature more accurate and complete.
[0217] After completing the window division, it is necessary to extract the category labels from each window feature matrix and combine these category labels into a category label sequence in a preset order. The core of this process is to arrange the category labels in each window area into a linear sequence to facilitate subsequent serialization processing and feature analysis. In the process of generating the category label sequence, it is necessary to ensure the accurate recording of the order and position information of each category label to ensure that the correct mapping of the category features can be maintained in subsequent processing.
[0218] For the category label sequence generated by each window feature matrix, a subtraction operation needs to be performed to obtain a category difference matrix. The generation process of the category difference matrix aims to measure the degree of change of the category features in each window by calculating the difference between the category labels. During the subtraction operation, the value of each category label is usually compared with the mean or median in the window to obtain the difference value between the categories. The result of the category difference matrix can reflect the change trend of the category features in each window area, which helps to identify the boundary information and category transition areas in the image.
[0219] In order to further improve the expressive power of the category difference matrix, it is necessary to add the position index to the category difference matrix to generate a category difference matrix containing position information. The addition of the position index allows each category difference value to contain not only the difference information between categories, but also the specific location of the difference value, thereby preserving the spatial distribution information of the category features. The position index is usually expressed in the form of two-dimensional coordinates and can be generated by marking the pixel positions of the category matrix.
[0220] The category difference matrix containing position information is flattened, and negative values are offset to construct a relative category index. The purpose of this step is to flatten the multidimensional data of the category difference matrix into a one-dimensional data sequence to facilitate vectorization in subsequent calculations. In order to avoid the influence of negative values on feature representation during the flattening process, negative values need to be offset to convert them into positive values or zero values, thereby ensuring the validity and accuracy of the category difference data. The implementation methods of flattening and negative value offset can be adjusted according to the needs of different application scenarios, such as optimizing the representation effect of category features through normalization, standardization or threshold operations.
[0221] After the relative category index is constructed, the category matrix is adjusted based on the relative category index to optimize the feature representation of the category matrix. The core of this adjustment process is to correct the weight of each pixel in the category matrix through the relative category index, thereby improving the expressive power of the category feature. The adjustment of the category matrix can be achieved through operations such as matrix multiplication and weighted average. The specific adjustment strategy needs to be dynamically optimized according to the distribution of the category features and the weight value of the relative category index.
[0222] This embodiment achieves refined expression and differentiated recognition capabilities of category features by optimizing and adjusting the category matrix. Through steps such as window division, category label sequence generation, category difference matrix calculation, position index addition and planarization processing, the feature representation effect of the category matrix can be effectively improved, and the spatial distribution information and category difference information of the category features can be enhanced. The optimized category matrix can more accurately reflect the category features in the image, providing high-quality input data for subsequent feature fusion and category label map generation.
[0223] In one embodiment, the above S40 includes:
[0224] S401, performing window segmentation on the gated category feature map to generate a plurality of window feature blocks;
[0225] S402, linearly expand each window feature block to construct an input sequence of multi-head attention;
[0226] S403, mapping the query vector, key vector and value vector of the input sequence to form an initial vector of multi-head attention;
[0227] S404, performing a scaled dot product attention calculation on the initial vector, and incorporating the category attention information into an attention scoring function as a summation term;
[0228] S405, concatenating and linearly transforming the output vectors of the attention scoring function to generate the global category feature.
[0229] In this embodiment, the process of extracting global category features from the gated category feature map through the global category attention guidance module aims to enhance the representation ability of category features by using global context information. This process is mainly achieved through steps such as window segmentation, multi-head attention mechanism, scaled dot product attention calculation, splicing and linear transformation to ensure that more comprehensive category features are extracted from the overall level.
[0230] First, the gated category feature map is window segmented to divide the feature map into multiple window feature blocks. The purpose of window segmentation is to split the large-size feature map into smaller area blocks without losing global information, so as to facilitate subsequent attention calculations. Each window feature block contains the feature information of the local area in the image while maintaining the overall spatial structure relationship. During the window segmentation process, the size and overlap ratio of the window can be adjusted according to the specific task requirements to balance the computational overhead and information retention effect.
[0231] After completing the window segmentation, each window feature block needs to be linearly expanded to construct the input sequence of multi-head attention. The linear expansion process converts the two-dimensional feature map into a one-dimensional feature sequence, so that the pixel information in each feature block can be processed one by one. This conversion operation not only retains the spatial structure information of each feature block, but also provides input data format conversion for subsequent query, key, and value mapping.
[0232] The input sequence is mapped into query vector, key vector and value vector, aiming to provide basic feature representation for the calculation of multi-head attention mechanism. The query vector represents the target information to be obtained from the feature map, the key vector represents the feature information of each position in the feature map, and the value vector represents the output information of the corresponding position in the feature map. By mapping the input sequence, the original feature sequence can be converted into the three core vectors of the multi-head attention mechanism, thereby realizing the association modeling between features.
[0233] After obtaining the query vector, key vector and value vector, the scaled dot product attention calculation is performed on the initial vector, and the category attention information is incorporated into the attention scoring function as a sum term. The core of this step is to generate an attention weight matrix by calculating the dot product of the query vector and the key vector to measure the correlation between the input features. In order to avoid the numerical instability caused by large eigenvalues, the dot product results need to be scaled. In the calculation of the attention scoring function, the category attention information is introduced as a sum term, which further enhances the recognition ability of category features and enables the model to pay more attention to the difference information between categories.
[0234] Finally, the output vector of the attention scoring function is concatenated and linearly transformed to generate the global category feature. The concatenation operation merges the output features of different attention heads to form a comprehensive representation containing multiple feature perspectives. The linear transformation adjusts the concatenated feature vector to the same dimension as the original feature map through the multiplication operation of the weight matrix, thereby ensuring that the generated global category feature can maintain the same shape as the input feature map. The global category feature processed by linear transformation contains category information from different window areas and integrates global context information, thereby providing high-quality input data for subsequent feature fusion and category label map generation.
[0235] This embodiment effectively solves the problem of insufficient global context information in the prior art by extracting global category features through the global category attention guidance module. Through the fusion of multi-head attention mechanism and category attention information, the model can focus on the category association between different regions in the image, thereby improving the representation ability of category features.
[0236] In one embodiment, the above S50 includes:
[0237] S501, performing branch convolution operations on the gated category feature map using convolution kernels of different sizes to extract local features of different scales;
[0238] S502, concatenating or adding the output results of the convolution operations of each branch to obtain a multi-scale local response;
[0239] S503, performing a depthwise separable convolution operation on the multi-scale local response, and performing activation and normalization operations on an output feature map of the depthwise separable convolution operation to generate the local category feature.
[0240] In this embodiment, the process of extracting local category features from the gated category feature map through the local category attention guidance module is to effectively capture the detailed features of the local area in the image, especially when dealing with objects with scale changes, local feature extraction is crucial. This process uses convolution kernels of different sizes to perform branch convolution operations, multi-scale feature fusion, and depth-separable convolution operations to achieve efficient local feature extraction, thereby generating local category features with stronger discriminative ability.
[0241] First, branch convolution operations are performed on the gated category feature map using convolution kernels of different sizes to extract local features of different scales. In actual implementation, common combinations of convolution kernel sizes can be selected, such as 1×1, 3×3, 5×5, etc. Each convolution kernel can focus on the feature extraction capabilities of different receptive fields. The 1×1 convolution kernel is used to extract pixel-level feature information, the 3×3 convolution kernel focuses on smaller local areas, and the 5×5 convolution kernel can capture a wider range of contextual information. This design of multi-branch convolution operations can effectively solve the problem of representing target objects in images at different scales, ensuring that feature information of both small and large targets can be extracted.
[0242] After completing the branch convolution operation, the output results of each branch convolution operation are spliced or added to obtain a multi-scale local response. The splicing operation can retain the feature integrity of each branch convolution and directly integrate feature maps of different scales into a larger feature map; the addition operation fuses the features of different branch convolutions to reduce the redundant information of the feature map. These two operation methods can be selected according to the actual application scenario and the limitations of computing resources. For example, in scenarios with fewer computing resources, the addition operation can be used to reduce the computational complexity; in scenarios that pay more attention to accuracy, the splicing operation can be used to retain more feature details.
[0243] Next, a depth-separable convolution operation is performed on the multi-scale local response, and the output feature map of the depth-separable convolution operation is activated and normalized to generate the local category feature. The depth-separable convolution operation decomposes the standard convolution into two steps: depth convolution and point convolution, thereby greatly reducing the complexity of the convolution calculation. In the depth convolution process, each convolution kernel only acts on one input channel, so it can extract local feature information more efficiently; and the point convolution integrates the outputs of each channel to generate an output feature map with comprehensive information. In order to ensure that the feature map after the depth-separable convolution operation can maintain numerical stability and fast convergence during the training process, the output feature map needs to be activated and normalized. Commonly used activation functions include ReLU, Leaky ReLU, etc., which can effectively introduce nonlinear changes and improve the expressiveness of the model; batch normalization can make the distribution of the feature map more stable, thereby accelerating the convergence of the model.
[0244] After being processed by the local category attention guidance module, the generated local category feature map can more accurately capture the category feature information of the local area in the image. Both small target objects and boundary details of the local area can be effectively extracted, providing high-quality input data for subsequent feature fusion and category label map generation.
[0245] This embodiment effectively solves the problem of the insufficiency of the prior art in extracting local detail features of images by extracting local category features from the gated category feature map through the process of local category attention guidance module. The use of convolution kernels of different scales and depth-separable convolution technology ensures that the model can accurately capture the category features of local areas when processing target objects with large scale changes. In addition, the splicing or addition operation of multi-scale local responses realizes the fusion of features of different scales, so that the model can capture local detail information while retaining the overall contextual information. Activation and normalization processing ensure the numerical stability of the feature map during the training process, thereby improving the convergence speed and classification accuracy of the model.
[0246] In one embodiment, in the above S10, performing multi-scale local feature extraction on the image data by a multi-scale feature extraction module to generate a multi-scale local feature map includes:
[0247] S101, constructing a pre-trained basic convolutional network, and performing preliminary feature extraction on the image data through the basic convolutional network to obtain a basic feature map;
[0248] S102, performing a multi-scale convolution operation on the basic feature map using convolution kernels of different sizes;
[0249] S103, performing a strip convolution operation on the output result of the multi-scale convolution operation to extract feature information of a long strip or a boundary shape;
[0250] S104, performing aggregation processing on the output results of the strip convolution operation to generate the multi-scale local feature map.
[0251] In this embodiment, the basic convolutional network usually uses a pre-trained model, such as ResNet, VGG or EfficientNet, which are pre-trained on large-scale data sets and have strong feature extraction capabilities. The input of the basic convolutional network is image data, and a feature map containing preliminary features is generated through a series of convolution, activation and pooling operations. This step is mainly responsible for extracting the underlying features of the image, including edge, texture, color and other information. These underlying features provide basic data for subsequent multi-scale feature extraction.
[0252] The convolution layer in the basic convolutional network extracts features from the input image by sliding windows. The size and step size of the convolution kernel determine the receptive field range of the extracted features, while the activation function (such as ReLU) introduces nonlinear changes, so that the feature extraction results can better describe the complex structure in the image.
[0253] On the basis of the basic feature map, multi-scale convolution operations are performed using convolution kernels of different sizes to extract feature information of different scales in the image. The core of this step is to use convolution kernels with multiple receptive fields to capture multi-level information in the image. For example, a 1×1 convolution kernel is used for pixel-level feature extraction, a 3×3 convolution kernel is used to capture local area features, and a 5×5 or larger convolution kernel is used to obtain contextual information in a larger range. Through multi-scale convolution operations, the scale change problem in the image can be effectively dealt with, so that both small and large targets can be accurately recognized by the model.
[0254] The output of the multi-scale convolution operation contains multiple feature maps of different scales, each of which corresponds to the feature information extracted by convolution kernels of different sizes. These feature maps together form a multi-scale feature space, which provides input data for subsequent strip convolution operations.
[0255] Strip convolution is a special convolution operation that aims to capture the characteristic information of long strips or borders in the image. Unlike the traditional square convolution kernel, the strip convolution kernel is more slender and can more effectively focus on features such as lines, borders, and textures in the image. This operation is particularly important for the recognition of long strips of target objects such as roads, rivers, and borders in remote sensing images.
[0256] The strip convolution operation can use convolution kernels in different directions, such as horizontal strip convolution, vertical strip convolution or diagonal strip convolution, to capture the long strip feature information in different directions. Each strip convolution kernel can focus on the boundary features in a specific direction, so that the model has stronger boundary recognition ability.
[0257] Aggregation processing is to integrate the output results of strip convolution operations into a unified feature map. Common aggregation methods include feature concatenation, feature addition, or feature pooling operations. Through aggregation processing, strip features in different directions and multi-scale convolution features can be integrated into a multi-scale local feature map containing rich contextual information.
[0258] In actual implementation, multiple feature maps can be aggregated using convolution or pooling operations. Convolution operations are used to generate feature maps containing multi-level information, while pooling operations can compress the size of feature maps and reduce computational complexity. The generated multi-scale local feature maps can more comprehensively describe the category features in the image, providing high-quality input data for subsequent category attention information extraction and label map generation.
[0259] This embodiment uses a multi-scale feature extraction module to extract multi-scale local features from the image data, and the generated multi-scale local feature map contains rich multi-scale information and boundary feature information, which can effectively solve the deficiencies in the image feature extraction process in the prior art. Through multi-scale convolution operations and strip convolution operations, it is ensured that the model can capture the feature information of target objects of different scales and directions, thereby improving the robustness and recognition accuracy of the model.
[0260] In one embodiment, the above S20 includes:
[0261] S201, extracting category attention information from the multi-scale local feature map, where the category attention information includes a category probability distribution of each pixel;
[0262] S202, applying a Softmax function to the category attention information to analyze the category weight of each pixel and generate a category weight map;
[0263] S203, applying the Argmax function to the category attention information, extracting the maximum probability category index of each pixel from the category probability distribution, and generating a category index map;
[0264] S204: Combine the category weight map and the category index map to generate the category matrix.
[0265] In this embodiment, extracting the category attention information from the multi-scale local feature map and constructing a category matrix based on the category weight and category index in the category attention information is an important step to achieve category feature enhancement. By extracting the category probability distribution of each pixel, calculating the category weight and category index, and finally generating a category matrix for subsequent feature guidance and adjustment. The category matrix contains the category information of each pixel in the image and is the basis for adjusting the category feature weight in the subsequent steps.
[0266] The multi-scale local feature map contains the local feature representation of each pixel in the image. In order to further obtain the category information at the pixel level, the multi-scale local feature map needs to be subjected to category probability analysis. The process of extracting category attention information usually uses a classifier or a fully connected layer to map each pixel in the feature map to a category probability distribution. This means that the feature vector of each pixel will be converted into a category probability vector whose length is equal to the total number of categories, and each value represents the probability that the pixel belongs to the corresponding category.
[0267] In the implementation process, the mapping of the category probability distribution can be completed through a simple linear layer or a multi-layer perceptron (MLP). The feature vector X of each pixel point pixel is input into the linear layer and outputs a class probability vector P pixel , where each element Indicates the probability value that the pixel belongs to category c.
[0268] In order to standardize the class probability distribution into comparable weight values, it is necessary to apply the Softmax function to the class probability distribution. The Softmax function converts any real-valued vector into a probability distribution so that the sum of the output values is 1. Through the Softmax function, the weight value of each class can represent the relative possibility that the pixel belongs to that class.
[0269] For each pixel, the class probability vector P pixel , after applying the Softmax function, we get the category weight vector W pixel , where each value represents the weight of the pixel to the corresponding category. These category weight values can be regarded as the importance of the category and are used for subsequent feature weight adjustment and category matrix construction.
[0270] In addition to the category weight map, it is also necessary to extract the maximum probability category index of each pixel to indicate the category to which the pixel is most likely to belong. By applying the Argmax function to the category probability distribution, the category index corresponding to the maximum probability value of each pixel can be extracted to generate a category index map.
[0271] The function of Argmax is to find the index position of the maximum value from the input vector. pixel , after applying the Argmax function, we get the category index I pixel Each element of the category index map represents the category index value to which the corresponding pixel belongs. This index value can be used in the subsequent category feature guidance process, enabling the model to adjust the feature map according to the category information.
[0272] The process of constructing the category matrix is completed by combining the category weight map and the category index map. The category weight map contains the weight value of each pixel for different categories, and the category index map contains the category index information to which each pixel belongs. Combining these two images can generate a category matrix containing category weights and category indexes.
[0273] Each element of the category matrix M pixel It can be expressed as a two-tuple (W pixel , I pixel ), where W pixel is the category weight value, I pixel is the category index value. The category matrix is used in the subsequent category feature adjustment step to optimize the representation ability of the feature map by dynamically adjusting the weights of the category features.
[0274] This embodiment can accurately capture the category information of each pixel in the image by extracting category attention information from the multi-scale local feature map and constructing a category matrix based on category weights and category indexes. The generated category matrix contains category weights and category index information, so that the model can better understand the category distribution in the image. This category matrix construction process can not only effectively reduce the confusion between categories, but also enhance the model's ability to pay attention to detail features, thereby improving the overall classification accuracy and segmentation effect of the model.
[0275] In one embodiment, the above S30 includes:
[0276] S307, using the multi-scale local feature map as the current layer feature map, and obtaining the previous layer feature map of the current layer feature map;
[0277] S308, performing convolution operations on the feature map of the current layer and the feature map of the previous layer respectively to generate a preliminary adjustment signal;
[0278] S309, processing the preliminary adjustment signal through an activation function to generate a gating signal;
[0279] S310, performing channel weighting and spatial weighting adjustment on the multi-scale local feature map through the gating signal;
[0280] S311, multiplying the adjusted multi-scale local feature map by the category weight map of the category matrix pixel by pixel to generate a gated category feature map.
[0281] In this embodiment, a gating signal is generated by a gating unit, and the multi-scale local feature map is adjusted by the gating signal, and the adjusted multi-scale local feature map is weighted channel by channel by the category matrix to generate a gated category feature map. In this process, a more accurate category feature map can be generated by dynamically controlling the weight values of different feature channels and adjusting the weights of each pixel in the feature map according to the category information. This operation not only enhances the expressive power of the category information, but also effectively reduces the redundant information in the feature map, thereby improving the recognition performance of the overall model.
[0282] First, the construction of the gating unit requires the feature map of the current layer and the feature map of the previous layer as input. The feature map of the current layer is usually a multi-scale local feature map, while the feature map of the previous layer is the feature map generated in the previous stage. These two feature maps provide different image information: the feature map of the current layer contains rich local detail features, while the feature map of the previous layer contains higher-level abstract information. The gating unit generates a gating signal by fusing these two features to dynamically control the weight distribution of the feature map.
[0283] In order to generate the gating signal, it is necessary to perform convolution operations on the feature map of the current layer and the feature map of the previous layer respectively. The purpose of the convolution operation is to extract the key information in the feature map and reduce the dimension of the feature map to reduce the computational complexity. Through the convolution operation, a preliminary adjustment signal can be obtained. Next, the preliminary adjustment signal is processed by the activation function to obtain the gating signal. The activation function usually selects the Sigmoid function or the ReLU function to ensure that the output value of the gating signal is within a reasonable range. Each value of the gating signal represents the importance of the corresponding feature channel. The larger the value, the more important the information of the channel.
[0284] After acquiring the gated signal, the multi-scale local feature map is adjusted by channel weighting and spatial weighting, so that the important features in the feature map are enhanced and the unimportant features are suppressed. Channel weighting adjustment refers to adjusting the weight of each channel of the feature map separately, thereby dynamically adjusting the contribution value of each channel. Spatial weighting adjustment refers to adjusting the weight of each pixel point in the feature map according to the importance of features at different positions. Through this adjustment process, the feature map can more accurately reflect the key information in the image.
[0285] The adjusted multi-scale local feature map needs to be further processed in combination with the category weight information of the category matrix. The category matrix contains a category weight map and a category index map, where the category weight map contains the category weight value of each pixel. These weight values reflect the relevance of each pixel to different categories. Multiplying the adjusted feature map with the category weight map pixel by pixel can further enhance the category information expression ability of the feature map, thereby generating the final gated category feature map. This category feature map not only contains the detailed information in the original feature map, but also strengthens the features related to the category, significantly improving the model's ability to distinguish different categories.
[0286] This embodiment dynamically adjusts the multi-scale local feature map through the gating unit, and performs pixel-by-pixel weighting processing in combination with the category weight information of the category matrix, thereby effectively enhancing the features of different categories in the feature map. The introduction of the gating unit allows the model to dynamically adjust the weight value of each feature channel according to the feature information of the current layer and the previous layer, so as to better capture the key features in the image. Combined with the category weight information of the category matrix, the adjusted feature map is weighted pixel by pixel, which can further strengthen the expression of category features in the feature map and reduce the confusion of category features.
[0287] In one embodiment, a device for generating image labels based on feature fusion is provided, and the device for generating image labels based on feature fusion corresponds to the method for generating image labels based on feature fusion in the above embodiment. Figure 3 , Figure 3This is a functional module diagram of a preferred embodiment of the image label generation device based on feature fusion of the present invention. Multi-scale feature extraction module 10, category attention extraction module 20, gate control unit module 30, global category attention guidance module 40, local category attention guidance module 50, category attention fusion module 60, feature fusion module 70 and label generation module 80. The functional modules are described in detail as follows:
[0288] A multi-scale feature extraction module 10 is used to obtain image data, perform multi-scale local feature extraction on the image data through the multi-scale feature extraction module, and generate a multi-scale local feature map;
[0289] A category attention extraction module 20, configured to extract category attention information from the multi-scale local feature map, and construct a category matrix based on category weights and category indexes in the category attention information;
[0290] A gating unit module 30, configured to generate a gating signal based on the category matrix through a gating unit, and adjust the multi-scale local feature map through the gating signal to generate a gated category feature map;
[0291] A global category attention guidance module 40, configured to extract global category features from the gated category feature map through the global category attention guidance module;
[0292] A local category attention guidance module 50, used to extract local category features from the gated category feature map through the local category attention guidance module;
[0293] A category attention fusion module 60, used to fuse the global category feature with the local category feature to generate a category attention guidance feature map;
[0294] A feature fusion module 70 is used to fuse the category attention guidance feature map with the multi-scale local feature map, and perform weighted processing on the fused feature map through the feature fusion head module to generate a final fused feature map;
[0295] The label generation module 80 is used to generate a category label map according to the final fused feature map.
[0296] In one embodiment, the category attention extraction module 20 is specifically used to:
[0297] Performing window division on the category label of each pixel in the category matrix to obtain multiple window feature matrices;
[0298] Extracting category labels from each window feature matrix respectively, and combining the category labels into a category label sequence according to a preset order;
[0299] Subtract the category label sequence corresponding to each window feature matrix to obtain the category difference matrix;
[0300] Adding the position index to the category difference matrix to obtain a category difference matrix containing position information;
[0301] The category difference matrix containing position information is flattened and negative values are offset to construct a relative category index;
[0302] The category matrix is adjusted based on the relative category index to optimize the feature representation of the category matrix.
[0303] In one embodiment, the global category attention guidance module 40 is specifically used to:
[0304] Performing window segmentation on the gated category feature map to generate a plurality of window feature blocks;
[0305] Expand each window feature block linearly to construct the input sequence of multi-head attention;
[0306] Mapping the input sequence into a query vector, a key vector, and a value vector to form an initial vector of multi-head attention;
[0307] Performing a scaled dot product attention calculation on the initial vector, and incorporating the category attention information as a summation term into an attention scoring function;
[0308] The output vectors of the attention scoring function are concatenated and linearly transformed to generate the global category feature.
[0309] In one embodiment, the local category attention guidance module 50 is specifically used to:
[0310] Using convolution kernels of different sizes to perform branch convolution operations on the gated category feature map to extract local features of different scales;
[0311] The output results of the convolution operations of each branch are concatenated or added to obtain multi-scale local responses;
[0312] A depthwise separable convolution operation is performed on the multi-scale local response, and an activation and normalization operation is performed on an output feature map of the depthwise separable convolution operation to generate the local category feature.
[0313] In one embodiment, the multi-scale feature extraction module 10 is specifically used for:
[0314] Constructing a pre-trained basic convolutional network, and performing preliminary feature extraction on the image data through the basic convolutional network to obtain a basic feature map;
[0315] Performing a multi-scale convolution operation on the basic feature map using convolution kernels of different sizes;
[0316] Performing a strip convolution operation on the output result of the multi-scale convolution operation to extract feature information of a long strip or boundary morphology;
[0317] Aggregate the output results of the strip convolution operation to generate the multi-scale local feature map.
[0318] In one embodiment, the category attention extraction module 20 is specifically used to:
[0319] Extracting category attention information from the multi-scale local feature map, the category attention information including category probability distribution of each pixel;
[0320] Applying a Softmax function to the category attention information to analyze the category weight of each pixel and generate a category weight map;
[0321] Applying the Argmax function to the category attention information, extracting the maximum probability category index of each pixel from the category probability distribution, and generating a category index map;
[0322] The category weight map and the category index map are combined to generate the category matrix.
[0323] In one embodiment, the gate control unit module 30 is specifically used for:
[0324] Use the multi-scale local feature map as the current layer feature map, and obtain the previous layer feature map of the current layer feature map;
[0325] Perform convolution operations on the feature map of the current layer and the feature map of the previous layer respectively to generate a preliminary adjustment signal;
[0326] Processing the preliminary adjustment signal through an activation function to generate a gating signal;
[0327] Performing channel weighting and spatial weighting adjustment on the multi-scale local feature map through the gating signal;
[0328] The adjusted multi-scale local feature map is multiplied pixel by pixel with the category weight map of the category matrix to generate a gated category feature map.
[0329] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a service-side method for generating image tags based on feature fusion.
[0330] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side method for generating image tags based on feature fusion
[0331] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0332] Acquire image data, and perform multi-scale local feature extraction on the image data through a multi-scale feature extraction module to generate a multi-scale local feature map;
[0333] Extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information;
[0334] Generate a gating signal through a gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map;
[0335] Extract global category features from the gated category feature map through a global category attention guidance module;
[0336] Extract local category features from the gated category feature map through the local category attention guidance module;
[0337] Fusion of the global category features with the local category features to generate a category attention guided feature map;
[0338] The category attention guided feature map is fused with the multi-scale local feature map, and the fused feature map is weighted through the feature fusion head module to generate the final fused feature map;
[0339] Generate a category label map according to the final fused feature map.
[0340] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0341] Acquire image data, and perform multi-scale local feature extraction on the image data through a multi-scale feature extraction module to generate a multi-scale local feature map;
[0342] Extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information;
[0343] Generate a gating signal through a gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map;
[0344] Extract global category features from the gated category feature map through a global category attention guidance module;
[0345] Extract local category features from the gated category feature map through the local category attention guidance module;
[0346] Fusion of the global category features with the local category features to generate a category attention guided feature map;
[0347] The category attention guided feature map is fused with the multi-scale local feature map, and the fused feature map is weighted through the feature fusion head module to generate the final fused feature map;
[0348] Generate a category label map according to the final fused feature map.
[0349] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0350] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0351] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0352] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for generating image labels based on feature fusion, characterized in that: The following steps are involved: Acquire image data, and perform multi-scale local feature extraction on the image data through a multi-scale feature extraction module to generate a multi-scale local feature map; Extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information; Generate a gating signal through a gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map; Extract global category features from the gated category feature map through a global category attention guidance module; Extract local category features from the gated category feature map through the local category attention guidance module; Fusion of the global category features with the local category features to generate a category attention guided feature map; The category attention guided feature map is fused with the multi-scale local feature map, and the fused feature map is weighted through the feature fusion head module to generate the final fused feature map; Generate a category label map according to the final fused feature map.
2. The method for generating image labels based on feature fusion according to claim 1, characterized in that: After extracting the category attention information from the multi-scale local feature map and constructing a category matrix based on the category weights and category indexes in the category attention information, the method further includes: Performing window division on the category label of each pixel in the category matrix to obtain multiple window feature matrices; Extracting category labels from each window feature matrix respectively, and combining the category labels into a category label sequence according to a preset order; Subtract the category label sequence corresponding to each window feature matrix to obtain the category difference matrix; Adding the position index to the category difference matrix to obtain a category difference matrix containing position information; The category difference matrix containing position information is flattened and negative values are offset to construct a relative category index; The category matrix is adjusted based on the relative category index to optimize the feature representation of the category matrix.
3. The image label generation method based on feature fusion according to claim 1, characterized in that: The global category features are extracted from the gated category feature map through the global category attention guidance module, including: Performing window segmentation on the gated category feature map to generate a plurality of window feature blocks; Expand each window feature block linearly to construct the input sequence of multi-head attention; Mapping the input sequence into a query vector, a key vector, and a value vector to form an initial vector of multi-head attention; Performing a scaled dot product attention calculation on the initial vector, and incorporating the category attention information as a summation term into an attention scoring function; The output vectors of the attention scoring function are concatenated and linearly transformed to generate the global category feature.
4. The method for generating image labels based on feature fusion according to claim 1, characterized in that: The local category features are extracted from the gated category feature map through the local category attention guidance module, including: Using convolution kernels of different sizes to perform branch convolution operations on the gated category feature map to extract local features of different scales; The output results of the convolution operations of each branch are concatenated or added to obtain multi-scale local responses; A depthwise separable convolution operation is performed on the multi-scale local response, and an activation and normalization operation is performed on an output feature map of the depthwise separable convolution operation to generate the local category feature.
5. The method for generating image labels based on feature fusion according to claim 1, characterized in that: The multi-scale local feature extraction module is used to extract the multi-scale local features of the image data to generate a multi-scale local feature map, including: Constructing a pre-trained basic convolutional network, and performing preliminary feature extraction on the image data through the basic convolutional network to obtain a basic feature map; Performing a multi-scale convolution operation on the basic feature map using convolution kernels of different sizes; Performing a strip convolution operation on the output result of the multi-scale convolution operation to extract feature information of a long strip or boundary morphology; Aggregate the output results of the strip convolution operation to generate the multi-scale local feature map.
6. The method for generating image labels based on feature fusion according to claim 1, characterized in that: Extracting category attention information from the multi-scale local feature map, and constructing a category matrix based on category weights and category indexes in the category attention information, including: Extracting category attention information from the multi-scale local feature map, the category attention information including category probability distribution of each pixel; Applying a Softmax function to the category attention information to analyze the category weight of each pixel and generate a category weight map; Applying the Argmax function to the category attention information, extracting the maximum probability category index of each pixel from the category probability distribution, and generating a category index map; The category weight map and the category index map are combined to generate the category matrix.
7. The method for generating image labels based on feature fusion according to claim 1, characterized in that: The method generates a gating signal by a gating unit, adjusts the multi-scale local feature map by the gating signal, and performs channel-by-channel weighted processing on the adjusted multi-scale local feature map by the category matrix to generate a gated category feature map, including: Use the multi-scale local feature map as the current layer feature map, and obtain the previous layer feature map of the current layer feature map; Perform convolution operations on the feature map of the current layer and the feature map of the previous layer respectively to generate a preliminary adjustment signal; Processing the preliminary adjustment signal through an activation function to generate a gating signal; Performing channel weighting and spatial weighting adjustment on the multi-scale local feature map through the gating signal; The adjusted multi-scale local feature map is multiplied pixel by pixel with the category weight map of the category matrix to generate a gated category feature map.
8. A device for generating image labels based on feature fusion, characterized in that: The image label generation device based on feature fusion includes: A multi-scale feature extraction module is used to obtain image data, perform multi-scale local feature extraction on the image data through the multi-scale feature extraction module, and generate a multi-scale local feature map; A category attention extraction module, used to extract category attention information from the multi-scale local feature map, and construct a category matrix based on category weights and category indexes in the category attention information; A gating unit module, used to generate a gating signal through the gating unit, adjust the multi-scale local feature map through the gating signal, and perform channel-by-channel weighted processing on the adjusted multi-scale local feature map through the category matrix to generate a gated category feature map; A global category attention guidance module, used for extracting global category features from the gated category feature map through the global category attention guidance module; A local category attention guidance module, used for extracting local category features from the gated category feature map through the local category attention guidance module; A category attention fusion module, used to fuse the global category feature with the local category feature to generate a category attention guidance feature map; The feature fusion module is used to fuse the category attention guided feature map with the multi-scale local feature map, and perform weighted processing on the fused feature map through the feature fusion head module to generate the final fused feature map; The label generation module is used to generate a category label map according to the final fusion feature map.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a feature fusion-based image label generation program stored in the memory and executable on the processor. When the feature fusion-based image label generation program is executed by the processor, the steps of the feature fusion-based image label generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program for generating an image label based on feature fusion, and when the program for generating an image label based on feature fusion is executed by a processor, the steps of the method for generating an image label based on feature fusion according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-focus joint segmentation method in retina OCT image based on hybrid network
CN118657800A
Image super-resolution processing method and device, computer equipment and medium
CN119168867A
Counterfactual context-aware texture learning for camouflaged object detection
WO2024187334A1
Cited By
User tag fusion method and system based on artificial intelligence
CN120470536A
Automatic channel system identification method and system based on artificial intelligence
CN120673263A
Image interpretable classification method and device, computer equipment and storage medium
CN120783103A
Unmanned aerial vehicle performance intelligent analysis method and system based on cloud platform
CN121388928A
Osteoporosis assessment method and system based on spine CT image and application method of system
CN122200080A