A multi-scale edge information multi-object detection method fusing text
Patent Information
- Application Number
- CN202511733717.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-08-07
- Estimated Expiration
- 2045-11-24
AI Technical Summary
[0005]本发明为了解决现有的文本融合方法在面对复杂背景时,存在的难以高效处理图像与文本之间的语义差异,同时,图文结合对于部分特征的理解仍不充分的问题,提供了一种融合文本的多尺度边缘信息多目标检测方法
[0056]1、通过引入MSES模块,结合多分支并行处理和边缘增强机制,有效提取多尺度上下文信息和边缘特征,显著提升了模型对物体边缘的敏感性,尤其在复杂背景和小目标检测中表现出更强的鲁棒性和准确性。
Smart Images

Figure CN121582546B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and deep learning, specifically to a multi-scale edge information multi-target detection method that integrates text. Background Technology
[0002] With the rapid development of deep learning, traditional image content-based object detection methods have made continuous progress. Their typical structure relies on predefined fixed target category vocabulary to complete the detection task. Against this backdrop, object detection methods that integrate text information have gradually become a research hotspot. By combining image and text information, the detection system is given greater flexibility, enabling it to better adapt to handling open vocabulary or objects of unknown categories. Text descriptions can provide semantic supplements to objects in the image, helping the system to understand the image content more deeply.
[0003] However, despite significant progress, existing text fusion methods still face a series of challenges: they struggle to efficiently handle semantic differences between images and text when faced with complex backgrounds, and the understanding of some features through image-text fusion remains insufficient. These issues are the main obstacles to further improving detection accuracy.
[0004] Therefore, it is necessary to propose a multi-scale edge information multi-target detection method that integrates text to solve the above problems. Summary of the Invention
[0005] To address the shortcomings of existing text fusion methods in efficiently handling semantic differences between images and text when faced with complex backgrounds, and the insufficient understanding of certain features in image-text fusion, this invention provides a multi-scale edge information multi-target detection method that integrates text.
[0006] This invention is achieved using the following technical solution:
[0007] A multi-scale edge information multi-target detection method integrating text includes the following steps:
[0008] S1: Acquire scene images and preprocess them, construct a dataset based on the preprocessed images, and divide it into training set and test set;
[0009] S2: Construct a basic YOLO-World model that includes a backbone network, a neck network, and a head network;
[0010] S3: Perform structural optimization on the basic YOLO-World model, including:
[0011] In the backbone network, the original C2f module is replaced by the MSES module;
[0012] At the end of the backbone network, an AFCAtion attention module is introduced;
[0013] In the neck network, the SSFF module is used to improve the feature fusion method of the original feature pyramid network;
[0014] S4: Iteratively train the optimized YOLO-World model using the training set, and evaluate the model's performance using the test set to obtain the object detection model;
[0015] S5: Use the trained object detection model to detect objects in the input image and output the object detection results.
[0016] Furthermore, the backbone network in the optimized YOLO-World model includes:
[0017] The initial feature extraction layer performs two convolutional downsampling operations on the input image to obtain the first feature map;
[0018] The feature extraction backbone is formed by connecting multiple layers in sequence. Except for the last layer, each layer contains an MSES module and a downsampling convolutional layer located thereafter. The last layer contains an MSES module. The feature extraction backbone is used to process the first feature map to output feature maps P3, P4 and P5 of different scales.
[0019] The AFCAtion attention module is used to perform channel attention enhancement on the P5 feature map output by the feature extraction backbone, and output the enhanced feature map to the neck network.
[0020] Furthermore, the MSES module includes:
[0021] The multi-branch parallel processing unit is used to extract and upsample multiple sets of feature maps of the same scale by processing the feature maps of the input MSES module through multiple adaptive average pooling with different pooling kernels, 1×1 convolution dimensionality reduction and depthwise separable convolution.
[0022] Multiple Edge modules are connected one-to-one with each branch of the multi-branch parallel processing unit to enhance the edge information of the feature map output by each branch.
[0023] The feature concatenation layer is used to concatenate the feature map enhanced by all Edge modules with the feature map of the input MSES module through channels.
[0024] The DSM module is used to optimize the spliced feature map output by the feature splicing layer through frequency domain decomposition and feature enhancement operations;
[0025] The convolutional processing layer is used to perform 1×1 convolution processing on the feature map output by the DSM module to generate the final output feature map of the MSES module.
[0026] Furthermore, the specific process by which the Edge module performs edge information enhancement is as follows:
[0027] The feature map of the input Edge module is smoothed by average pooling to extract its low-frequency information;
[0028] Subtract the smoothed feature map from the input Edge module's feature map to obtain high-frequency edge information;
[0029] The high-frequency edge information is subjected to convolution processing;
[0030] The edge information after convolution is added to the feature map of the input Edge module to output the final output feature map of the Edge module.
[0031] Furthermore, the specific process of the DSM module performing optimization is as follows:
[0032] The feature map of the input DSM module is subjected to max pooling and average pooling along the channel dimension, and the pooling results are concatenated and then processed by convolution to generate a general feature map.
[0033] Perform channel separation transformation on the general feature map;
[0034] A mean filter is applied to the feature map after channel separation transformation to generate low-frequency features;
[0035] The low-frequency features are subtracted from the feature map of the input DSM module to obtain complementary high-frequency features;
[0036] The final output feature map of the DSM module is obtained by using the high-frequency features and the feature map after channel separation and transformation through element-wise multiplication and residual connection.
[0037] Furthermore, the specific process by which the AFCAtion attention module performs channel attention enhancement is as follows:
[0038] Global average pooling is performed on the feature map of the input AFCAtion attention module to obtain channel descriptors;
[0039] The channel descriptors are processed using a strip matrix to capture local information that characterizes the interaction relationships between local channels;
[0040] The channel descriptors are operated on using a diagonal matrix to capture global information representing global channel dependencies;
[0041] The obtained local information and global information are combined by matrix transpose and element-wise multiplication to generate a correlation matrix;
[0042] The Sigmoid function is applied to the correlation matrix, and a learnable parameter is introduced to dynamically adjust the feature weights in order to generate attention weights corresponding to the local and global information.
[0043] The local and global information are weighted and summed using the attention weights respectively to obtain the fused channel attention features;
[0044] The fused channel attention features are mapped and multiplied with the feature map of the input AFCAtion attention module to output the final output feature map of the AFCAtion attention module.
[0045] Furthermore, the SSFF module is used to perform scale sequence fusion processing on feature maps of different scales output by the backbone network. The specific process is as follows:
[0046] The feature maps at different scales of the input SSFF module are convolved with a set of Gaussian kernels with increasing standard deviations to generate multi-scale smooth feature maps.
[0047] The size of the multi-scale smooth feature map is adjusted to be uniform using the nearest neighbor interpolation method;
[0048] The multi-scale smooth feature maps with uniform size are stacked, and the stacked tensor is converted to a specified dimension using a depth transformation method.
[0049] The transformed feature map is processed sequentially by convolution, batch normalization, and activation function to complete the fusion and extraction of multi-scale features, and output the final feature map of the SSFF module.
[0050] Furthermore, the working process of the head network in the optimized YOLO-World model includes:
[0051] The feature map output by the neck network is convolved to extract visual features;
[0052] The input text is encoded using a CLIP encoder to obtain text feature embeddings;
[0053] The extracted visual features are embedded and compared with the embedded text features to calculate their similarity.
[0054] The initial detection results based on visual features are combined with the calculated similarity results to generate the final target detection result.
[0055] This invention provides a multi-scale edge information multi-target detection method that integrates text, which has the following advantages compared with the prior art:
[0056] 1. By introducing the MSES module and combining multi-branch parallel processing and edge enhancement mechanism, multi-scale contextual information and edge features are effectively extracted, which significantly improves the model's sensitivity to object edges, especially showing stronger robustness and accuracy in complex background and small object detection.
[0057] 2. The AFCAtion attention module is adopted, which dynamically adjusts feature weights through interaction between local and global channels, enhancing feature representation capabilities and enabling the model to focus more accurately on key regions, reducing semantic differences between images and text.
[0058] 3. The SSFF module is used for scale sequence fusion. Through Gaussian smoothing and multi-scale feature alignment, more efficient feature fusion is achieved, which improves the detection performance of multi-scale targets.
[0059] 4. After overall model optimization, when fusing text information, the similarity between visual and text features is calculated by comparative learning. This not only improves the flexibility of open vocabulary object detection, but also enhances the model's adaptability to unknown categories, making it more practical in real-world application scenarios.
[0060] In summary, these improvements collectively address the problems of existing text fusion methods in handling semantic differences and insufficient feature understanding in complex contexts, providing a more advanced solution for object detection technology. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of the optimized YOLO-World model in this invention.
[0062] Figure 2 This is a schematic diagram of the MSES module in this invention.
[0063] Figure 3 This is a schematic diagram of the AFCAtion attention module in this invention.
[0064] Figure 4 This is a schematic diagram of the SSFF module in this invention. Detailed Implementation
[0065] The relevant technical solutions will be clearly and completely described below. Obviously, the described embodiments are only some embodiments, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. Example
[0066] A multi-scale edge information multi-target detection method integrating text includes the following steps:
[0067] S1: Acquire scene images and preprocess them. Construct a dataset based on the preprocessed images and divide it into training and test sets.
[0068] S2: Construct a basic YOLO-World model that includes a backbone network, a neck network, and a head network.
[0069] S3: Optimize the structure of the basic YOLO-World model, as shown in the attached figure. Figure 1 As shown, it includes:
[0070] In the backbone network, the original C2f module is replaced by the MSES module.
[0071] At the end of the backbone network, an AFCAtion attention module is introduced.
[0072] In the neck network, the SSFF module is used to improve the feature fusion method of the original feature pyramid network.
[0073] The optimized YOLO-World model's backbone network includes: a preliminary feature extraction layer, a feature extraction backbone formed by multiple layers connected sequentially, and an AFCAtion attention module.
[0074] In the feature extraction backbone, except for the last layer, each layer contains an MSES module and a downsampling convolutional layer thereafter, and the last layer contains an MSES module; the feature extraction backbone is used to process the first feature map to output feature maps P3, P4 and P5 of different scales.
[0075] As attached Figure 2 As shown, the MSES module includes a multi-branch parallel processing unit, multiple Edge modules, a feature concatenation layer, a DSM module, and a convolutional processing layer.
[0076] S4: Iteratively train the optimized YOLO-World model using the training set, and evaluate the model's performance using the test set to obtain the object detection model.
[0077] S5: Use the trained object detection model to detect objects in the input image and output the object detection results.
[0078] In this embodiment, the input image input to the backbone network of the optimized YOLO-World model is labeled as I.
[0079] The input image I undergoes two convolutional downsampling operations in the initial feature extraction layer of the backbone network to obtain the first feature map F1, as shown in the formula:
[0080] ;
[0081] In the formula, It is represented as a standard 3×3 convolution.
[0082] The first feature map F1 is processed by the MSES module, and the processing procedure is as follows:
[0083] First, the first feature map F1 is input into the four parallel branches of the multi-branch parallel processing unit. These four branches are processed using adaptive average pooling layers with kernel sizes of 3, 6, 9, and 12, respectively, resulting in multiple sets of feature maps at different scales. This captures multi-scale contextual information in a lightweight manner. Then, each feature map is reduced in dimensionality using a 1×1 convolution to decrease the number of channels and reduce subsequent computation. Finally, features are further extracted using depthwise separable convolution, as shown in the following formula:
[0084] ;
[0085] In the formula, i represents different branches, and its values are 3, 6, 9, and 12. Depth-separable convolutions with a 3×3 kernel size are performed; finally, upsampling is used to align features at different scales to the same scale.
[0086] F output by each parallel branch i The input is fed into the corresponding Edge module, which is an edge enhancement module specifically designed to extract edge information to enhance the sensitivity of the target detection model to edges. The specific process of edge information enhancement performed by the Edge module is as follows:
[0087] First, the feature map F of the input Edge module... i Average pooling smoothing is performed to extract its low-frequency information;
[0088] Then, the feature map F of the input Edge module is... i Subtracting the smoothed feature map from the high-frequency edge information yields the high-frequency edge information.
[0089] Next, the high-frequency edge information is convolved to obtain F', as shown in the following formula:
[0090] ;
[0091] In the formula, AvgPool represents the average pooling smoothing operation, and Conv represents the convolution operation;
[0092] Finally, the edge information after convolution is added to the feature map of the input Edge module to output the final output feature map F* of the Edge module, as shown in the following formula:
[0093] .
[0094] In this invention, the low-frequency information refers to features that characterize regions of gradual change in an image, obtained through smoothing operations such as average pooling; the high-frequency edge information refers to features that characterize regions of abrupt intensity changes (such as object edges) in an image, obtained by performing a difference operation between the original feature map and the smoothed feature map.
[0095] The feature map F* output by the Edge module is input into the DSM module. The specific optimization process performed by the DSM module is as follows:
[0096] First, the feature map F* of the input DSM module is subjected to max pooling and average pooling along the channel dimension, and the pooling results are concatenated and then processed by convolution to generate a general feature map F. t The formula is as follows:
[0097] ;
[0098] In the formula, AvgPool represents the average pooling operation, MaxPool represents the max pooling operation, and Conv3 represents the convolution operation with a kernel size of 3×3.
[0099] Then, the general feature map F t Perform channel separation transformation to obtain F s The formula is as follows:
[0100] ;
[0101] In the formula, This represents a cascaded depthwise convolutional layer with kernel sizes of 5×5 and 7×7. This represents a depthwise convolution with a 3×3 kernel size. This represents element-wise multiplication. It is a copy The tiling function along the second-order channel dimension is extended to... H and W are the spatial height and width of the feature map, respectively, and C is the number of channels in the feature map;
[0102] Finally, the transformed feature F s A mean filter is applied to generate low-frequency features. These low-frequency features are subtracted from the feature map of the input DSM module to obtain complementary high-frequency features. The high-frequency features are then combined with the channel-separated feature map through element-wise multiplication and residual concatenation to obtain the final output feature map of the DSM module. The formula is as follows:
[0103] ;
[0104] In the formula, It is a mean filter. This indicates element-wise multiplication.
[0105] Subsequently, the final output feature map of the DSM module is processed by a 1×1 convolution layer of the MSES module to generate the final output feature map P3 of the MSES module at this level. By repeating the processing steps involving the MSES module and downsampling in subsequent layers of the feature extraction backbone, feature maps P4 and P5 are obtained sequentially. At this point, the feature extraction backbone has completed multi-scale feature extraction of the input image and finally outputs the multi-scale feature maps P3, P4, and P5.
[0106] The P5 feature map, which has the lowest resolution and richest semantic information, is input into the AFCAtion attention module located at the end of the backbone network for channel attention enhancement. (See attached image) Figure 3 As shown, the specific process of the AFCAtion attention module performing channel attention enhancement is as follows:
[0107] First, feature map P5 is transformed from a feature map containing spatial information into a one-dimensional channel descriptor through global average pooling. The formula is as follows:
[0108] ;
[0109] In the formula, GAP represents the global average pooling operation, which transforms the shape of feature map P5 from C×H×W to C×1×1, where C, H, and W are the number of channels, height, and width of the input feature map P5, respectively. This represents the feature value of the nth channel at spatial location (i, j), where n is the channel index, n∈[1,2,...,C]; U n This is the value of the obtained channel descriptor on the nth channel.
[0110] Then, local channel interaction is performed using the strip matrix B, as shown below:
[0111] ;
[0112] In the formula, k represents the number of adjacent channels, U i This is local information.
[0113] In this embodiment, the operation of using a strip matrix for local channel interaction is implemented through a one-dimensional convolutional layer.
[0114] The operation of capturing global dependencies using a diagonal matrix is implemented through a two-dimensional convolutional layer.
[0115] Next, the dependencies between channels are captured as global information using a diagonal matrix D, as shown below:
[0116] ;
[0117] In the formula, c represents the number of channels, U g This is global information.
[0118] In this embodiment, the operation of capturing global dependencies using a diagonal matrix is implemented through a two-dimensional convolutional layer.
[0119] Then, the local information obtained through the strip matrix and the global information obtained through the diagonal matrix are combined, and the correlation at different fine-grained levels is obtained by using matrix transpose and element-wise multiplication, as shown in the following formula:
[0120] ;
[0121] In the formula, M is the correlation matrix, and an adaptive fusion strategy is adopted to accurately allocate feature weights and reduce computational complexity.
[0122] By introducing learnable parameters To dynamically adjust feature weights for dynamic fusion, the formula is as follows:
[0123] ;
[0124] ;
[0125] ;
[0126] In the formula, , These represent the weights of the fused global channel information and the weights of the local channel information, respectively, where c is the number of channels. This is represented by the sigmoid activation function.
[0127] Finally, the obtained weight W is mapped and multiplied with the feature map P5 of the input AFCAtion attention module to output the final output feature map of the AFCAtion attention module.
[0128] The SSFF module is used to perform scale sequence fusion processing on feature maps of different scales output by the backbone network, as shown in the attached figure. Figure 4 As shown, the specific process is as follows:
[0129] First, P3, P4, and P5 are each compared with a set of increasing standard deviations. < < Two-dimensional Gaussian kernel Perform convolution to generate multi-scale smooth feature maps P3 smooth P4 smooth P5 smooth The formula for a two-dimensional Gaussian kernel is as follows:
[0130] ;
[0131] In the formula, (u,v) are the coordinates within the convolution kernel window. This is the standard deviation parameter.
[0132] Then, the nearest neighbor interpolation method is used to smooth the P4 using Gaussian. smooth P5 smooth Size adjusted to P3 smooth Size. The adjusted P3 will follow. smooth P4 smooth P5 smooth By stacking, each feature layer is transformed from a 3D tensor (height, width, channels) to a 4D tensor (depth, height, width, channels) by adding a depth dimension.
[0133] Finally, the spliced feature maps are subjected to 3D convolution, 3D batch normalization, and SiLU activation function to complete the extraction of scale sequences.
[0134] The working process of the head network in the optimized YOLO-World model includes:
[0135] The feature map output by the neck network is convolved to extract visual features;
[0136] The input text is encoded using a CLIP encoder to obtain text feature embeddings;
[0137] The extracted visual features are embedded with the text features for comparative learning, and their similarity is calculated. This process is implemented through a comparative learning head, which may contain a branch with batch normalization or a standard similarity calculation branch, and finally outputs the similarity score between each visual region and the text description.
[0138] The initial detection results based on visual features are concatenated with the calculated semantic similarity results, and the final target detection result is generated through subsequent processing. This result includes both the spatial location of the object and its matching confidence with the text description.
[0139] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-scale edge information multi-target detection method that integrates text, characterized in that: Includes the following steps: S1: Acquire scene images and preprocess them, construct a dataset based on the preprocessed images, and divide it into training set and test set; S2: Construct a basic YOLO-World model that includes a backbone network, a neck network, and a head network; S3: Perform structural optimization on the basic YOLO-World model, including: In the backbone network, the original C2f module is replaced by the MSES module; At the end of the backbone network, an AFCAtion attention module is introduced; In the neck network, the SSFF module is used to improve the feature fusion method of the original feature pyramid network; The MSES module includes a multi-branch parallel processing unit, multiple edge modules, a feature concatenation layer, a DSM module, and a convolutional processing layer. The multiple edge modules are connected one-to-one with each branch of the multi-branch parallel processing unit and are used to enhance the edge information of the feature map output by each branch. The DSM module is used to optimize the concatenated feature map output by the feature concatenation layer through frequency domain decomposition and feature enhancement operations. The specific process of optimization performed by the DSM module is as follows: The feature map of the input DSM module is subjected to max pooling and average pooling along the channel dimension, and the pooling results are concatenated and then processed by convolution to generate a general feature map. Perform channel separation transformation on the general feature map; A mean filter is applied to the feature map after channel separation transformation to generate low-frequency features; The low-frequency features are subtracted from the feature map of the input DSM module to obtain complementary high-frequency features; The final output feature map of the DSM module is obtained by using the high-frequency features and the feature map after channel separation and transformation through element-wise multiplication and residual connection. The specific process by which the AFCAtion attention module performs channel attention enhancement is as follows: Global average pooling is performed on the feature map of the input AFCAtion attention module to obtain channel descriptors; The channel descriptors are processed using a strip matrix to capture local information that characterizes the interaction relationships between local channels; The channel descriptors are operated on using a diagonal matrix to capture global information representing global channel dependencies; The obtained local information and global information are combined by matrix transpose and element-wise multiplication to generate a correlation matrix; The Sigmoid function is applied to the correlation matrix, and a learnable parameter is introduced to dynamically adjust the feature weights in order to generate attention weights corresponding to the local and global information. The local and global information are weighted and summed using the attention weights respectively to obtain the fused channel attention features; The fused channel attention features are mapped and multiplied with the feature map of the input AFCAtion attention module to output the final output feature map of the AFCAtion attention module. S4: Iteratively train the optimized YOLO-World model using the training set, and evaluate the model's performance using the test set to obtain the object detection model; S5: Use the trained object detection model to detect objects in the input image and output the object detection results.
2. The multi-scale edge information multi-target detection method for fused text according to claim 1, characterized in that: The backbone network in the optimized YOLO-World model includes: The initial feature extraction layer performs two convolutional downsampling operations on the input image to obtain the first feature map; The feature extraction backbone is formed by connecting multiple layers in sequence. Except for the last layer, each layer contains an MSES module and a downsampling convolutional layer located thereafter. The last layer contains an MSES module. The feature extraction backbone is used to process the first feature map to output feature maps P3, P4 and P5 of different scales. The AFCAtion attention module is used to perform channel attention enhancement on the P5 feature map output by the feature extraction backbone, and output the enhanced feature map to the neck network.
3. The multi-scale edge information multi-target detection method for fused text according to claim 1, characterized in that: The multi-branch parallel processing unit is used to process the feature map of the input MSES module through multiple adaptive average pooling with different pooling kernels, 1×1 convolution dimensionality reduction and depthwise separable convolution, and extract and upsample multiple sets of feature maps to the same scale. The feature stitching layer is used to perform channel stitching of the feature map enhanced by all Edge modules with the feature map of the input MSES module; The convolutional processing layer is used to perform 1×1 convolution processing on the feature map output by the DSM module to generate the final output feature map of the MSES module.
4. The multi-scale edge information multi-target detection method for fused text according to claim 1, characterized in that: The specific process by which the Edge module performs edge information enhancement is as follows: The feature map of the input Edge module is smoothed by average pooling to extract its low-frequency information; Subtract the smoothed feature map from the input Edge module's feature map to obtain high-frequency edge information; The high-frequency edge information is subjected to convolution processing; The edge information after convolution is added to the feature map of the input Edge module to output the final output feature map of the Edge module.
5. The multi-scale edge information multi-target detection method for fused text according to claim 2, characterized in that: The SSFF module is used to perform scale sequence fusion processing on feature maps of different scales output by the backbone network. The specific process is as follows: The feature maps at different scales of the input SSFF module are convolved with a set of Gaussian kernels with increasing standard deviations to generate multi-scale smooth feature maps. The size of the multi-scale smooth feature map is adjusted to be uniform using the nearest neighbor interpolation method; The multi-scale smooth feature maps with uniform size are stacked, and the stacked tensor is converted to a specified dimension using a depth transformation method. The transformed feature map is processed sequentially by convolution, batch normalization, and activation function to complete the fusion and extraction of multi-scale features, and output the final feature map of the SSFF module.
6. The multi-scale edge information multi-target detection method for fused text according to claim 1, characterized in that: The working process of the head network in the optimized YOLO-World model includes: The feature map output by the neck network is convolved to extract visual features; The input text is encoded using a CLIP encoder to obtain text feature embeddings; The extracted visual features are embedded and compared with the embedded text features to calculate their similarity. The initial detection results based on visual features are combined with the calculated similarity results to generate the final target detection result.
Citation Information
Patent Citations
Lightweight night target detection method and system based on deep learning
CN119723045A
Tile product surface defect detection method based on SAPC-Net network
CN119887727A