A multi-domain modeling and semantic embedding enhanced lightweight small target detection method

A lightweight small target detection method enhanced by multi-domain modeling and semantic embedding solves the problems of high-frequency detail information loss and insufficient feature fusion in lightweight detection networks, and achieves high-precision detection of small targets.

CN122090229BActive Publication Date: 2026-07-14CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-04-23
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing lightweight detection networks suffer from problems such as loss of high-frequency detail information and insufficient fusion of spatial and frequency features during feature extraction, resulting in low detection accuracy for small targets.

Method used

A lightweight small target detection method with multi-domain modeling and semantic embedding enhancement is adopted. By using high-frequency residual enhancement spatial depth transformation convolutional units, spatial-frequency context enhancement layer aggregation units, and frequency-guided semantic injection units, high-frequency detail information is preserved and spatial and frequency domain context information is fully integrated to improve the discriminative ability of features.

Benefits of technology

It significantly improves the detection accuracy of the lightweight model for small targets, ensuring high-efficiency detection performance in scenarios with limited computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090229B_ABST
    Figure CN122090229B_ABST
Patent Text Reader

Abstract

The application discloses a multi-domain modeling and semantic embedding enhanced lightweight small target detection method. The multi-domain modeling and semantic embedding enhanced lightweight small target detection method comprises the following steps: acquiring a to-be-detected image; inputting the to-be-detected image into a trained lightweight detection model, so as to determine a backbone feature map of the to-be-detected image through a plurality of high-frequency residual enhancement spatial depth conversion convolution units and a plurality of spatial-frequency context enhancement layer aggregation units; fusing the backbone feature map through a feature pyramid fusion method to obtain a fused feature map; determining a neck feature map based on the fused feature through a plurality of frequency-guided semantic injection units; and detecting the neck feature map through a plurality of detection heads to obtain a detection result of the to-be-detected image, thereby improving the detection precision of small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of lightweight small target detection with multi-domain modeling and semantic embedding enhancement, and in particular to a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement. Background Technology

[0002] With the widespread application of target detection technology in computing-constrained scenarios such as mobile devices and embedded devices, lightweight detection models have become a key focus of research and application. Existing lightweight detection networks often reduce computational overhead by simplifying the network structure and using depthwise separable convolutions. However, during feature extraction, they often suffer from the loss of high-frequency detail information and insufficient fusion of spatial and frequency features, resulting in low detection accuracy for small targets. Summary of the Invention

[0003] This application aims to at least address the technical problems existing in the prior art. To this end, this application proposes a lightweight small target detection method based on multi-domain modeling and semantic embedding enhancement, which can improve the detection accuracy of the model for small targets.

[0004] The first aspect of this application provides a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement, comprising the following steps:

[0005] Acquire the image to be detected;

[0006] The image to be detected is input into a trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model; wherein, the trained lightweight detection model includes several high-frequency residual enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads, and the process of the trained lightweight detection model outputting the detection result of the image to be detected includes:

[0007] The backbone feature map of the image to be detected is determined by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units.

[0008] The backbone feature map is fused using the feature pyramid fusion method to obtain the fused feature map;

[0009] Based on the fused features, the neck feature map is determined by the several frequency-guided semantic injection units;

[0010] The detection results of the image to be detected are obtained by detecting the neck feature map using the aforementioned detection heads.

[0011] The lightweight small target detection method based on multi-domain modeling and semantic embedding enhancement according to the embodiments of this application has at least the following beneficial effects:

[0012] This application first determines the backbone feature map of the image to be detected through several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units. By introducing high-frequency residual-enhanced spatial depth transformation convolutional units, high-frequency detail information in the image is effectively preserved, avoiding the loss of small target features. Furthermore, by introducing spatial-frequency context enhancement layer aggregation units, spatial and frequency domain context information is fully integrated, improving the discriminative ability of features. Based on the fused features, several frequency-guided semantic injection units are used to determine the neck feature map, thereby enhancing the semantic representation of small targets. Finally, several detection heads are used to detect the neck feature map to obtain the detection result of the image to be detected. Thus, this application significantly improves the detection accuracy of small targets under a lightweight model architecture.

[0013] A second aspect of this application provides a lightweight small target detection system with multi-domain modeling and semantic embedding enhancement, the lightweight small target detection system with multi-domain modeling and semantic embedding enhancement comprising:

[0014] The data construction module is used to acquire the image to be detected;

[0015] An image detection module is used to input the image to be detected into a trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model; wherein, the trained lightweight detection model includes several high-frequency residual enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads, and the process of the trained lightweight detection model outputting the detection result of the image to be detected includes:

[0016] The backbone feature map of the image to be detected is determined by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units.

[0017] The backbone feature map is fused using the feature pyramid fusion method to obtain the fused feature map;

[0018] Based on the fused features, the neck feature map is determined by the several frequency-guided semantic injection units;

[0019] The detection results of the image to be detected are obtained by detecting the neck feature map using the aforementioned detection heads.

[0020] This system first determines the backbone feature map of the image to be detected through several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units. By introducing high-frequency residual-enhanced spatial depth transformation convolutional units, high-frequency detail information in the image is effectively preserved, avoiding the loss of small target features. Furthermore, by introducing spatial-frequency context enhancement layer aggregation units, spatial and frequency domain context information is fully integrated, improving the discriminative ability of features. Based on the fused features, several frequency-guided semantic injection units determine the neck feature map, thereby strengthening the semantic representation of small targets. Finally, several detection heads detect the neck feature map to obtain the detection result of the image to be detected. Thus, this application significantly improves the detection accuracy of small targets under a lightweight model architecture.

[0021] A third aspect of this application provides an electronic device including at least one controller and a memory for communicatively connecting to the controller; the memory stores instructions executable by the at least one controller, the instructions being executed by the at least one controller to cause the at least one controller to perform a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement as described in the first aspect of this application.

[0022] A fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement as described in the first aspect of this application.

[0023] It should be noted that the beneficial effects of the third and fourth aspects of this application compared with the prior art are the same as the beneficial effects of the aforementioned lightweight small target detection method with multi-domain modeling and semantic embedding enhancement compared with the prior art, and will not be elaborated here.

[0024] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0025] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0026] Figure 1 This is a flowchart illustrating an embodiment of the lightweight small target detection method with multi-domain modeling and semantic embedding enhancement provided in this application;

[0027] Figure 2This is a schematic diagram of the overall structure of a trained lightweight detection model, representing an embodiment of the lightweight small target detection method with multi-domain modeling and semantic embedding enhancement provided in this application.

[0028] Figure 3 This is a schematic diagram of the structure of an embodiment of the lightweight small target detection system with multi-domain modeling and semantic embedding enhancement provided in this application;

[0029] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0030] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0031] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.

[0032] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0033] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0034] With the widespread application of target detection technology in computing-constrained scenarios such as mobile devices and embedded devices, lightweight detection models have become a key focus of research and application. Existing lightweight detection networks often reduce computational overhead by simplifying the network structure and using depthwise separable convolutions. However, during feature extraction, they often suffer from the loss of high-frequency detail information and insufficient fusion of spatial and frequency features, resulting in low detection accuracy for small targets.

[0035] To address the aforementioned technical deficiencies, this application provides a lightweight small target detection method that combines multi-domain modeling and semantic embedding enhancement.

[0036] Please see Figure 1 and Figure 2 This is a flowchart illustrating a lightweight small target detection method based on multi-domain modeling and semantic embedding enhancement provided in an embodiment of this application. This method is applied to electronic devices, such as servers. Figure 1 As shown, this lightweight small object detection method with multi-domain modeling and semantic embedding enhancement includes:

[0037] Step S101: Obtain the image to be detected;

[0038] In step S101, the above-mentioned acquisition of the image to be detected can be a digital image file stored in a local storage device manually selected by the user as the image to be detected.

[0039] Step S102: Input the image to be detected into the trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model; wherein, the trained lightweight detection model includes several high-frequency residual enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads, and the process of the trained lightweight detection model outputting the detection result of the image to be detected includes:

[0040] Step S1021: Determine the backbone feature map of the image to be detected by using several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units;

[0041] The trained lightweight detection model can be a lightweight detection model based on the YOLOv10 architecture. The trained lightweight detection model also includes a first C2f unit, a first Conv (convolutional) unit, a first SCDown unit, a second SCDown unit, a C2fCIB unit, an SPPF unit, and an AIFI unit. The trained lightweight detection model can also include a second Conv (convolutional) unit, a third SCDown unit, a second C2f unit, and a third C2f unit. The trained lightweight detection model also includes a fourth C2f unit, a fifth C2f unit, and a sixth C2f unit.

[0042] Reference Figure 2 The C2f unit mentioned above can be the core feature extraction unit of the YOLOv10 backbone network and the YOLOv10 neck network. The YOLOv10 backbone network is used to extract basic visual features from the original image in layers. The YOLOv10 neck network is located between the backbone network of the model and the detection head and is used for bidirectional fusion and enhancement of multi-scale features.

[0043] The SCDown unit mentioned above can be a downsampling unit of the YOLOv10 backbone network and the YOLOv10 neck network, used to perform efficient spatial and channel feature reconstruction operations and achieve efficient feature dimensionality reduction.

[0044] The aforementioned C2fCIB unit (cross-stage partially compact inverted bottleneck unit) can be an improved module for the backbone network in the YOLOv10 architecture, which replaces the Bottleneck unit in the traditional C2f structure to improve efficiency and performance.

[0045] The SPPF unit mentioned above can be the fast spatial pyramid pooling unit used in the backbone network of the YOLOv10 architecture.

[0046] The AIFI unit mentioned above can be an attention-based intra-scale feature interaction unit of the YOLOv10 backbone network.

[0047] Reference Figure 2 , Figure 2 In this context, Concat means concatenation. Figure 2 In this context, "Upsample" indicates upsampling, referring to the aforementioned high-frequency residual-enhanced spatial depth transformation convolutional units (for...). Figure 2 The HRSPD-Conv in the model may include a first high-frequency residual enhanced spatial depth transformation convolutional unit and a second high-frequency residual enhanced spatial depth transformation convolutional unit.

[0048] The above-mentioned space-frequency context enhancement layer aggregation units (for) Figure 2 The SFELAN in the model may include a first spatial-frequency context enhancement layer aggregation unit and a second spatial-frequency context enhancement layer aggregation unit; the aforementioned spatial-frequency context enhancement layer aggregation units may also include a third spatial-frequency context enhancement layer aggregation unit and a fourth spatial-frequency context enhancement layer aggregation unit. The aforementioned first spatial-frequency context enhancement layer aggregation unit may include an efficient channel attention mechanism block and several phantom convolution blocks.

[0049] The aforementioned backbone feature map may include a first backbone feature map, a second backbone feature map, and a third backbone feature map.

[0050] Step S1022: The backbone feature map is fused using the feature pyramid fusion method to obtain the fused feature map;

[0051] The fused feature map can include a first fused feature map, a second fused feature map, and a third fused feature map with progressively decreasing resolutions. For example, if the resolution of the first fused feature map is 80×80, the resolution of the second fused feature map is 40×40, and the resolution of the third fused feature map is 20×20.

[0052] Step S1023: Based on the fused features, determine the neck feature map by guiding semantic injection units through several frequencies;

[0053] The aforementioned neck feature maps may include high-resolution neck feature maps, medium-resolution neck feature maps, and low-resolution neck feature maps.

[0054] The above-mentioned frequency-guided semantic injection units (for) Figure 2 The FG-SIU in the text includes a first frequency-guided semantic injection unit, a second frequency-guided semantic injection unit, and a third frequency-guided semantic injection unit.

[0055] Step S1024: Detect the neck feature map using several detection heads to obtain the detection result of the image to be detected.

[0056] The detection results of the above-mentioned image to be detected may include the category of the image to be detected (such as human or animal), the target location coordinates of the image to be detected, and the target confidence value of the image to be detected.

[0057] This application first determines the backbone feature map of the image to be detected through several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units. By introducing high-frequency residual-enhanced spatial depth transformation convolutional units, high-frequency detail information in the image is effectively preserved, avoiding the loss of small target features. Furthermore, by introducing spatial-frequency context enhancement layer aggregation units, spatial and frequency domain context information is fully integrated, improving the discriminative ability of features. Based on the fused features, several frequency-guided semantic injection units are used to determine the neck feature map, thereby enhancing the semantic representation of small targets. Finally, several detection heads are used to detect the neck feature map to obtain the detection result of the image to be detected. Thus, this application significantly improves the detection accuracy of small targets under a lightweight model architecture.

[0058] In some embodiments, step S1021 may include steps S201 to S211:

[0059] Step S201: Extract features from the image to be detected using the first high-frequency residual enhanced spatial depth transformation convolutional unit to obtain the first downsampled feature map;

[0060] Step S202: Extract features from the first downsampled feature map using the second high-frequency residual enhanced spatial depth transformation convolutional unit to obtain the second downsampled feature map;

[0061] Step S203: Extract features from the second downsampled feature map using the first C2f unit to obtain the first transition feature map;

[0062] Step S204: Extract features from the first transition feature map using the first Conv unit to obtain the third downsampled feature map;

[0063] Step S205: Extract features from the third downsampled feature map through the first spatial-frequency context enhancement layer aggregation unit to obtain the first backbone feature map;

[0064] Step S206: Downsample the first backbone feature map using the first SCDown unit to obtain the fourth downsampled feature map;

[0065] Step S207: Extract features from the fourth downsampled feature map through the second spatial-frequency context enhancement layer aggregation unit to obtain the second backbone feature map;

[0066] Step S208: Downsample the second backbone feature map using the second SCDown unit to obtain the fifth downsampled feature map;

[0067] Step S209: Extract features from the fifth downsampled feature map using the C2fCIB unit to obtain the second transition feature map;

[0068] Step S210: Pool the second transition feature map using SPPF units to obtain the third transition feature map;

[0069] Step S211: Perform feature interaction on the third transition feature map through the AIFI unit to obtain the third main feature map.

[0070] This application introduces a first high-frequency residual enhanced spatial depth transformation convolutional unit and a second high-frequency residual enhanced spatial depth transformation convolutional unit, which can continuously capture and enhance high-frequency detail information from the image to be detected and intermediate feature maps, effectively making up for the shortcomings of traditional lightweight models in preserving high-frequency information. At the same time, through the first and second spatial-frequency context enhancement layer aggregation units, deep fusion of spatial context information and frequency structure information is achieved, so that the extracted features contain both rich local details and global semantic and frequency characteristics, significantly improving the expressive power of the features. In addition, the backbone feature map is divided into the first backbone feature map. The first, second, and third backbone feature maps correspond to features at different resolutions and abstraction levels. Together with the synergistic effects of the first C2f unit, first Conv unit, first SCDown unit, second SCDown unit, C2fCIB unit, SPPF unit, and AIFI unit, a multi-scale, multi-domain feature extraction backbone network is constructed. This ensures that information flow can be efficiently transmitted and integrated throughout the entire backbone feature map extraction process, avoids the loss of key information, and provides more accurate data for subsequent feature pyramid fusion. As a result, the accuracy and robustness of the trained lightweight detection model for small targets are significantly improved.

[0071] In some embodiments, step S201 may include steps S301 to S305:

[0072] Step S301: Rearrange the pixel blocks at adjacent spatial positions in the image to be detected into the channel to obtain the rearranged feature map;

[0073] Step S302: Extract the basic downsampled feature map from the rearranged feature map through a 3x3 convolution;

[0074] Step S303: Apply the Laplacian operator to the rearranged feature map to obtain the high-frequency response feature map;

[0075] Step S304: Extract the high-frequency enhanced feature map from the high-frequency response features using a 1x1 convolution;

[0076] Step S305: Add the basic downsampled feature map and the high-frequency enhanced feature map element by element to obtain the first downsampled feature map.

[0077] In some embodiments, the calculation process of extracting features from the first downsampled feature map by the second high-frequency residual enhanced spatial depth transformation convolution unit in step S202 to obtain the second downsampled feature map is similar to the calculation process of extracting features from the image to be detected by the first high-frequency residual enhanced spatial depth transformation convolution unit in step S201 to obtain the first downsampled feature map, and will not be described again here.

[0078] This application provides a structured foundation for subsequent high-frequency feature extraction by rearranging pixel blocks of the image to be detected into channels. It extracts basic spatial features through 3x3 convolution, while using the Laplacian operator to directly capture high-frequency responses from the rearranged feature map and refines them into high-frequency enhanced features through 1x1 convolution. Finally, the basic downsampled feature map and the high-frequency enhanced feature map are added element-wise, thereby achieving deep fusion of spatial features and high-frequency detail features. This provides a more accurate feature representation for subsequent detection and effectively improves the detection accuracy of the trained lightweight detection model for small targets.

[0079] In some embodiments, step S205 may include steps S401 to S412:

[0080] Step S401: Extract the first intermediate feature map from the third downsampled feature map using a 1x1 convolution;

[0081] Step S402: Perform reparameterized convolution processing on the first intermediate feature map to obtain the second intermediate feature map;

[0082] Step S403: After dividing the second intermediate feature map along the channel dimension to obtain the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map, the local spatial feature map is extracted from the first intermediate sub-feature map by a 3x3 convolution.

[0083] In step S403, the above-mentioned division of the second intermediate feature map along the channel dimension to obtain the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map can be achieved by uniformly dividing the second intermediate feature map along the channel dimension into four equal parts to obtain the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map. Specifically, when the second intermediate feature map is a 128-channel feature map, the 128 channels are uniformly divided into four equal parts, and each part is divided into 32 channels.

[0084] Step S404: Extract the context expansion feature map from the second intermediate sub-feature map using a 3x3 dilated convolution with a dilation rate of 2;

[0085] Step S405: Extract the frequency structure feature map from the third intermediate sub-feature map using wavelet convolution;

[0086] Step S406: Extract the channel supplementary feature map from the fourth intermediate sub-feature map through a 1x1 convolution;

[0087] Step S407: Concatenate the local spatial feature map, context extended feature map, frequency structure feature map, and channel supplementary feature map to obtain the concatenated feature map;

[0088] Step S408: Extract the high-efficiency channel attention feature map from the stitched feature map using the high-efficiency channel attention mechanism block;

[0089] Step S409: Extract the spatial-frequency fusion feature map from the efficient channel attention feature map through 1x1 convolution;

[0090] Step S410: After extracting the corresponding first spatial-frequency phantom convolutional feature map from the spatial-frequency fusion feature map through the first phantom convolutional block, extract the corresponding second spatial-frequency phantom convolutional feature map from the first spatial-frequency phantom convolutional feature map through the second phantom convolutional block, extract the corresponding third spatial-frequency phantom convolutional feature map from the second spatial-frequency phantom convolutional feature map through the third phantom convolutional block, and so on, until the corresponding (n-1)th spatial-frequency phantom convolutional feature map is extracted from the (n-2)th spatial-frequency phantom convolutional feature map through the (n-n-1)th phantom convolutional block, and extract the channel adjustment feature map from the (n-1)th spatial-frequency phantom convolutional feature map through a 1x1 convolution, where n-1 is the total number of phantom convolutional blocks;

[0091] Step S411: Concatenate the first intermediate feature map, the second intermediate feature map, the spatial-frequency fusion feature map, all spatial-frequency phantom convolution feature maps, and the channel adjustment feature map to obtain the aggregated feature map;

[0092] Step S412: Extract the first backbone feature map from the aggregated feature map through 1x1 convolution.

[0093] In some embodiments, the calculation process of extracting features from the fourth downsampled feature map through the second spatial-frequency context enhancement layer aggregation unit to obtain the second backbone feature map in step S207 is similar to the calculation process of extracting features from the third downsampled feature map through the first spatial-frequency context enhancement layer aggregation unit to obtain the first backbone feature map in step S205, and will not be described again here.

[0094] This application employs multi-channel parallel extraction of local spatial feature maps, context-extended feature maps, frequency structure feature maps, and channel supplementary feature maps, ensuring comprehensive capture of image information from different dimensions and avoiding information loss that may occur with single feature extraction methods. Subsequently, these multi-source features are initially integrated through a concatenation operation, and an efficient channel attention mechanism block is introduced to dynamically learn and strengthen channel features crucial for small target detection, effectively suppressing redundant information and making the fused features more discriminative. The efficient channel attention feature map is then processed through 1x1 convolution to obtain a spatial-frequency fusion feature map, and diverse spatial-frequency phantom convolution feature maps are generated using phantom convolution blocks. Finally, the first intermediate feature map, the second intermediate feature map, the spatial-frequency fusion feature map, all spatial-frequency phantom convolution feature maps, and the channel adjustment feature map are concatenated to obtain an aggregated feature map. The first backbone feature map is extracted from the aggregated feature map through 1x1 convolution, providing more accurate data for subsequent feature pyramid fusion and detection head, thereby significantly improving the detection accuracy of the trained lightweight detection model for small targets.

[0095] In some embodiments, step S1022 may include steps S501 to S512:

[0096] Step S501: Upsample the third backbone feature map to obtain the upsampled third backbone feature map;

[0097] Step S502: Combine the upsampled third backbone feature map with the second backbone feature map to obtain the first combined backbone feature map;

[0098] Step S503: Extract features from the first spliced ​​backbone feature map using the second C2f unit to obtain the intermediate fused feature map;

[0099] Step S504: Upsample the intermediate fused feature map to obtain the upsampled intermediate fused feature map;

[0100] Step S505: The upsampled intermediate fusion feature map is spliced ​​with the first backbone feature map to obtain the second spliced ​​backbone feature map;

[0101] Step S506: Extract features from the second spliced ​​backbone feature map using the third C2f unit to obtain the first fused feature map;

[0102] Step S507: Downsample the first fused feature map using the second Conv unit to obtain the downsampled first fused feature map;

[0103] Step S508: After downsampling, the first fused feature map and the intermediate fused feature map are spliced ​​together to obtain the third spliced ​​backbone feature map;

[0104] Step S509: Extract features from the third spliced ​​backbone feature map through the third space-frequency context enhancement layer aggregation unit to obtain the second fused feature map;

[0105] The calculation process of step S509 above, in which the third spliced ​​backbone feature map is extracted by the third spatial-frequency context enhancement layer aggregation unit to obtain the second fused feature map, is similar to the calculation process of step S205 above, in which the third downsampled feature map is extracted by the first spatial-frequency context enhancement layer aggregation unit to obtain the first backbone feature map, and will not be described again here.

[0106] Step S510: Downsample the second fused feature map using the third SCDown unit to obtain the downsampled second fused feature map;

[0107] Step S511: Combine the second fused feature map after downsampling with the third backbone feature map to obtain the fourth combined backbone feature map;

[0108] Step S512: Extract features from the fourth spliced ​​backbone feature map through the fourth spatial-frequency context enhancement layer aggregation unit to obtain the third fused feature map.

[0109] The calculation process of extracting features from the fourth spliced ​​backbone feature map through the fourth spatial-frequency context enhancement layer aggregation unit to obtain the third fused feature map in step S512 is similar to the calculation process of extracting features from the third downsampled feature map through the first spatial-frequency context enhancement layer aggregation unit to obtain the first backbone feature map in step S205, and will not be repeated here.

[0110] This application constructs a multi-scale fusion structure comprising a first fusion feature map, a second fusion feature map, and a third fusion feature map, enabling the model to comprehensively capture target information at different scales. Furthermore, in the key path of feature fusion, a third space-frequency context enhancement layer aggregation unit and a fourth space-frequency context enhancement layer aggregation unit are introduced. Simultaneously, by combining the second Conv unit, the third SCDown unit, the second C2f unit, and the third C2f unit, efficient feature transformation, downsampling, and feature extraction with low information loss are achieved, thereby significantly improving the detection accuracy of the trained lightweight detection model for small targets.

[0111] In some embodiments, step S1023 may include steps S601 to S606:

[0112] Step S601: Based on the second fused feature map and the third fused feature map, the semantic injection unit is guided by the first frequency to determine the intermediate enhanced feature map;

[0113] Step S601 above may include step S6011:

[0114] Step S6011: When using the second fused feature map as the first current scale feature map and the third fused feature map as the first higher-level semantic feature map, the intermediate enhanced feature map corresponding to the first current scale feature map and the first higher-level semantic feature map is determined by the first frequency-guided semantic injection unit. The specific process of determining the intermediate enhanced feature map corresponding to the first current scale feature map and the first higher-level semantic feature map may include:

[0115] By performing a 1-level Haar DWT (first-level Haar discrete wavelet transform) on the first current scale feature, the first high-frequency sub-band, the second high-frequency sub-band, and the third high-frequency sub-band are obtained.

[0116] In calculating the absolute values ​​of the first, second, and third high-frequency sub-bands, the absolute values ​​of the first, second, and third high-frequency sub-bands are added element by element to obtain the high-frequency energy map.

[0117] Calculate the average value of the high-frequency energy map along the channel dimension to obtain the single-channel high-frequency response map;

[0118] Upsample the single-channel high-frequency response graph to obtain the upsampled single-channel high-frequency response graph.

[0119] The upsampled single-channel high-frequency response map is smoothed using DWConv (depth-separable convolution) to obtain a smoothed response map;

[0120] Applying the Sigmoid function to the smooth response map yields a frequency-guided gating map;

[0121] Upsample the first higher-level semantic feature map to obtain an upsampled first higher-level semantic feature map with the same spatial size as the first current-scale feature map;

[0122] The first high-level semantic feature map after upsampling is processed by 1×1 convolution to extract features, resulting in an aligned semantic feature map.

[0123] The frequency-guided gating map is multiplied element-wise with the aligned semantic feature map to obtain the controlled semantic feature map;

[0124] The controlled semantic feature map is injected into the first current-scale feature map using the following formula in a residual manner to obtain the intermediate enhanced feature map:

[0125] ;

[0126] in, For intermediate enhanced feature maps, This is the first feature map at the current scale. This is a scaling factor pre-set according to actual needs. This is a controlled semantic feature map.

[0127] Step S602: Based on the first fused feature map and the second fused feature map, the semantic injection unit is guided by the second frequency to determine the top-level intermediate feature map;

[0128] Step S602 above may include step S6021:

[0129] Step S6021: When the first fused feature map is used as the second current scale feature map and the second fused feature map is used as the second higher-level semantic feature map, the top-level intermediate feature map corresponding to the second current scale feature map and the second higher-level semantic feature map is determined by the second frequency-guided semantic injection unit.

[0130] The calculation process of determining the top-level intermediate feature map corresponding to the second current scale feature map and the second higher-level semantic feature map through the second frequency-guided semantic injection unit in step S6021 is similar to the calculation process of determining the intermediate enhanced feature map corresponding to the first current scale feature map and the first higher-level semantic feature map through the first frequency-guided semantic injection unit in step S6011, and will not be repeated here.

[0131] Step S603: Based on the intermediate enhanced feature map and the top intermediate feature map, the top enhanced feature map is determined by guiding the semantic injection unit through the third frequency;

[0132] Step S603 above may include step S6031:

[0133] Step S6031: When the top intermediate feature map is used as the third current scale feature map and the intermediate enhanced feature map is used as the third higher level semantic feature map, the top enhanced feature map corresponding to the third current scale feature map and the third higher level semantic feature map is determined by the third frequency-guided semantic injection unit.

[0134] The calculation process of determining the top-level enhanced feature map corresponding to the third current scale feature map and the third higher-level semantic feature map through the third frequency-guided semantic injection unit in step S6031 is similar to the calculation process of determining the intermediate enhanced feature map corresponding to the first current scale feature map and the first higher-level semantic feature map through the first frequency-guided semantic injection unit in step S6011, and will not be repeated here.

[0135] Step S604: Extract features from the top-level enhanced feature map using the fourth C2f unit to obtain a high-resolution neck feature map;

[0136] Step S605: Extract features from the intermediate enhanced feature map using the fifth C2f unit to obtain a medium-resolution neck feature map;

[0137] Step S606: Extract features from the third fused feature map using the sixth C2f unit to obtain a low-resolution neck feature map.

[0138] This application ensures the consistency of semantic information in the feature fusion process at different levels by using a frequency-guided semantic injection unit. Furthermore, the use of the C2f unit further refines and extracts features at each level, enhancing their discriminative power and making them more suitable for detecting small targets at different resolutions. It also deepens the semantic expression of intermediate layer features while effectively preserving the global contextual information of low-resolution features. Finally, it generates multi-scale neck feature maps, thereby significantly improving the detection accuracy of the trained lightweight detection model for small targets.

[0139] In some embodiments, the training process of the trained lightweight detection model may include steps S701 to S709:

[0140] Step S701: Obtain the training dataset, which includes training images and a set of real target bounding boxes;

[0141] In step S701, the training dataset can be a historical object detection dataset obtained through manual annotation.

[0142] Step S702: Construct a teacher model and an initial student model. The teacher model includes several first initial high-frequency residual enhancement spatial depth transformation convolutional units, several first initial spatial-frequency context enhancement layer aggregation units, several first initial frequency-guided semantic injection units, and several first initial detection heads. The initial student model includes several second initial high-frequency residual enhancement spatial depth transformation convolutional units, several second initial spatial-frequency context enhancement layer aggregation units, several second initial frequency-guided semantic injection units, and several second initial detection heads.

[0143] The teacher model mentioned above can be an improved YOLOv10-X model, and the initial student model mentioned above can be an improved YOLOv10-N model.

[0144] Step S703: Input the training dataset into the teacher model to determine the first backbone feature map of the training dataset through several first initial high-frequency residual enhanced spatial depth transformation convolutional units and several first initial spatial-frequency context enhancement layer aggregation units of the teacher model.

[0145] In step S703, the above-mentioned process of inputting the training dataset into the teacher model to determine the first backbone feature map of the training dataset through several first initial high-frequency residual enhanced spatial depth transformation convolutional units and several first initial spatial-frequency context enhancement layer aggregation units is similar to the process of determining the backbone feature map of the image to be detected through several high-frequency residual enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units in step S1021, and will not be repeated here.

[0146] Step S704: Input the training dataset into the initial student model to determine the second backbone feature map of the training dataset through several second initial high-frequency residual enhanced spatial depth transformation convolutional units and several second initial spatial-frequency context enhancement layer aggregation units of the initial student model.

[0147] In step S704, the above-mentioned process of inputting the training dataset into the initial student model to determine the second backbone feature map of the training dataset through several second initial high-frequency residual enhanced spatial depth transformation convolutional units and several second initial spatial-frequency context enhancement layer aggregation units is similar to the process of determining the backbone feature map of the image to be detected through several high-frequency residual enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units in step S1021, and will not be repeated here.

[0148] Step S705: Based on the set of real target bounding boxes, the first backbone feature map, and the second backbone feature map, update the backbone network parameters of the initial student model to obtain the first student model. The first student model includes several third initial high-frequency residual enhancement spatial depth transformation convolutional units, several third initial spatial-frequency context enhancement layer aggregation units, several third initial frequency-guided semantic injection units, and several third initial detection heads.

[0149] The aforementioned first backbone feature map may include the first backbone feature map output by the teacher model, the second backbone feature map output by the teacher model, and the third backbone feature map output by the teacher model.

[0150] The aforementioned second backbone feature map may include the first backbone feature map output by the initial student model, the second backbone feature map output by the initial student model, and the third backbone feature map output by the initial student model.

[0151] In step S705, the backbone network parameters of the initial student model are updated based on the set of real target bounding boxes, the first backbone feature map, and the second backbone feature map to obtain the first student model, which may include:

[0152] According to the feature layer step size preset according to actual needs, the set of real target bounding boxes in the image coordinate system is mapped to the coordinate system corresponding to the current teacher feature layer, so as to obtain the set of bounding boxes of the current teacher layer corresponding to each feature layer. The current teacher feature layer includes the first backbone feature map output by the teacher model, the second backbone feature map output by the teacher model, and the third backbone feature map output by the teacher model.

[0153] By calculating the center coordinates of each current teacher layer annotation box in the current teacher layer annotation box set, a two-dimensional Gaussian heatmap corresponding to the current teacher feature layer is generated with the center coordinates of each current teacher layer annotation box as the mean and the width and height of each current teacher layer annotation box as the diffusion range.

[0154] Calculate the standard deviation of each two-dimensional Gaussian heatmap along the horizontal axis; calculate the standard deviation of each two-dimensional Gaussian heatmap along the vertical axis.

[0155] Based on the center coordinates of each current teacher layer's bounding box, the standard deviation of each 2D Gaussian heatmap along the horizontal axis, and the standard deviation of each 2D Gaussian heatmap along the vertical axis, the following formula is used to synthesize all 2D Gaussian heatmaps corresponding to all bounding boxes in the current teacher feature layer, resulting in the foreground probability mask map corresponding to each current teacher layer:

[0156] ;

[0157] in, The pixel position coordinates in the current teacher feature layer are... The values ​​of the foreground probability mask at that time. The x-coordinate of the pixel position value. The ordinate of the pixel position coordinate value. The pixel position coordinates are The x-coordinate of the center coordinate of the current teacher layer annotation box. The pixel position coordinates are The vertical coordinate of the center of the current teacher layer annotation box at that time. Pixel position coordinates are The standard deviation of the corresponding two-dimensional Gaussian heatmap along the horizontal axis. Pixel position coordinates are The standard deviation of the corresponding two-dimensional Gaussian heatmap along the vertical axis;

[0158] The high-frequency edge response value of each current teacher feature layer is calculated using the following formula:

[0159] ;

[0160] in, The pixel position coordinates in the current teacher feature layer are... High-frequency edge response value at time, For the Laplace operator, For the current teacher characteristic layer The corresponding feature map output by the teacher model at that time (the current teacher feature layer is...) The corresponding feature map output by the teacher model can be the first backbone feature map, the second backbone feature map, or the third backbone feature map output by the teacher model.

[0161] All high-frequency edge response values ​​of the current teacher feature layer are merged in the channel dimension to obtain an edge response map with the same spatial size as each current teacher feature layer;

[0162] Based on the edge response map and the foreground probability mask map corresponding to each current teacher feature layer, the spatial distillation weights of each current teacher feature layer are calculated using the following formula:

[0163] ;

[0164] in, The pixel position coordinates in the current teacher feature layer are... Spatial distillation weights over time This refers to all regions corresponding to the current set of annotation boxes for the teacher level. The edge enhancement coefficient is preset according to actual needs. A foreground enhancement coefficient pre-set according to actual needs;

[0165] After extracting the number of channels, height, and width of each current teacher feature layer, the hierarchical distillation weights of each current teacher feature layer are calculated using the following formula:

[0166] ;

[0167] ;

[0168] ;

[0169] in, For the current teacher characteristic layer The corresponding high-frequency response diagram, For the absolute value operation, For the current teacher characteristic layer The corresponding high-frequency intensity index, The high-frequency response diagram is at the 1st Each channel and the pixel position coordinates are High-frequency response value at time, For the current teacher characteristic layer The corresponding number of channels, For the current teacher characteristic layer The corresponding height at that time For the current teacher characteristic layer The corresponding width, For the current teacher characteristic layer The corresponding stratified distillation weights, This is the current set of teacher feature layers (including the first backbone feature map output by the teacher model, the second backbone feature map output by the teacher model, and the third backbone feature map output by the teacher model). For the current teacher characteristic layer The corresponding high-frequency intensity index, The temperature coefficient is preset according to actual needs;

[0170] Extract the pixel values ​​of the first backbone feature map, the second backbone feature map, and the third backbone feature map output by the initial student model; extract the pixel values ​​of the first backbone feature map, the second backbone feature map, and the third backbone feature map output by the teacher model.

[0171] When the number of channels in the current student feature layer (which includes the first backbone feature map, the second backbone feature map, and the third backbone feature map output by the initial student model) is the same as the number of channels in the corresponding current teacher feature layer, the double adaptive distillation loss value is calculated using the following formula based on spatial distillation weights and hierarchical distillation weights:

[0172] ;

[0173] in, This is a dual adaptive distillation loss value. It is the square norm 2. The current student feature layer output by the initial student model is The corresponding feature map at the pixel position coordinates is The pixel values ​​at that time (the current student feature layer includes the first backbone feature map output by the initial student model, the second backbone feature map output by the initial student model, and the third backbone feature map output by the initial student model; it should be noted that...) (This can be a pre-set index value based on actual needs) For the current teacher characteristic layer The pixel position coordinates are The pixel value at that time;

[0174] Specifically, when the number of channels in the current student feature layer (which includes the first backbone feature map, the second backbone feature map, and the third backbone feature map output by the initial student model) is different from the number of channels in the corresponding current teacher feature layer, a 1x1 convolution operation is applied to the current student feature layer to complete channel alignment, resulting in an aligned current student feature layer. Then, the double adaptive distillation loss value is calculated when the aligned current student feature layer replaces the current student feature layer output by the initial student model.

[0175] If the double adaptive distillation loss value is less than the preset loss threshold set according to actual needs, the initial student model is used as the first student model; if the double adaptive distillation loss value is greater than or equal to the preset loss threshold, the initial student model is iteratively updated according to the double adaptive distillation loss value through backpropagation until the number of iterations reaches the preset maximum number of iterations set according to actual needs, and the first student model is obtained.

[0176] Step S706: After fusing the first backbone feature map using the feature pyramid fusion method through the teacher model to obtain the first fused feature map, the first neck feature map is determined based on the first fused feature map and guided by several first initial frequencies through semantic injection units.

[0177] In step S706, the calculation process of fusing the first backbone feature map using the feature pyramid fusion method through the teacher model to obtain the first fused feature map is similar to the calculation process of fusing the backbone feature map using the feature pyramid fusion method in step S1022 to obtain the fused feature map, and will not be repeated here.

[0178] In step S706, the calculation process of determining the first neck feature map based on the first fused features and guided by several first initial frequencies through semantic injection units is similar to the calculation process of determining the neck feature map based on the fused features and guided by several frequencies through semantic injection units in step S1023, and will not be repeated here.

[0179] Step S707: Input the training dataset into the first student model to determine the third backbone feature map of the training dataset through several third initial high-frequency residual enhanced spatial depth transformation convolutional units and several third initial spatial-frequency context enhancement layer aggregation units of the first student model.

[0180] In step S707, the above-mentioned process of inputting the training dataset into the first student model to determine the third backbone feature map of the training dataset through several third initial high-frequency residual enhanced spatial depth transformation convolutional units and several third initial spatial-frequency context enhancement layer aggregation units of the first student model is similar to the process of determining the backbone feature map of the image to be detected through several high-frequency residual enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units in step S1021, and will not be repeated here.

[0181] Step S708: After fusing the third backbone feature map using the feature pyramid fusion method through the first student model to obtain the second fused feature map, the second neck feature map is determined based on the second fused feature map by guiding semantic injection units through several third initial frequencies.

[0182] In step S708, the calculation process of fusing the third backbone feature map using the feature pyramid fusion method through the first student model to obtain the second fused feature map is similar to the calculation process of fusing the backbone feature map using the feature pyramid fusion method in step S1022 to obtain the fused feature map, and will not be repeated here.

[0183] In step S708, the calculation process of determining the second neck feature map based on the second fused features and guided by several third initial frequencies through semantic injection units is similar to the calculation process of determining the neck feature map based on the fused features and guided by several frequencies through semantic injection units in step S1023, and will not be repeated here.

[0184] Step S709: Based on the set of real target bounding boxes, the first neck feature map, and the second neck feature map, update the neck network parameters and detection head parameters of the first student model to obtain the trained lightweight detection model.

[0185] In step S709, the calculation process of updating the neck network parameters and detection head parameters of the first student model based on the real target bounding box set, the first neck feature map, and the second neck feature map to obtain the trained lightweight detection model is similar to the calculation process of updating the backbone network parameters of the initial student model based on the real target bounding box set, the first backbone feature map, and the second backbone feature map to obtain the first student model in step S705, and will not be repeated here.

[0186] This application focuses on transferring edge and texture structure information of small targets through the distillation of features in the backbone network. Through the distillation of the neck network, the student model is guided to learn the semantic expression and detection decision-making ability of the teacher model. By coordinating the spatial distillation weights and hierarchical distillation weights, the distillation signal is accurately focused at key spatial locations and key feature levels. After completing the dual adaptive distillation training of the backbone network, neck network and detection head, the student model can achieve accurate detection of small targets in complex backgrounds while maintaining a low number of parameters and computational complexity. This significantly improves the detection accuracy of the trained lightweight detection model for small targets.

[0187] Additionally, refer to Figure 3 One embodiment of this application provides a lightweight small target detection system with multi-domain modeling and semantic embedding enhancement, including a data construction module 1100 and an image detection module 1200, wherein:

[0188] Data construction module 1100 is used to acquire the image to be detected;

[0189] The image detection module 1200 is used to input the image to be detected into a trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model. The trained lightweight detection model includes several high-frequency residual-enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads. The process of the trained lightweight detection model outputting the detection result of the image to be detected includes:

[0190] The backbone feature map of the image to be detected is determined by using several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context-enhanced layer aggregation units.

[0191] The backbone feature maps are fused using the feature pyramid fusion method to obtain the fused feature map.

[0192] Based on the fused features, the neck feature map is determined by guiding the semantic injection unit through several frequencies.

[0193] The detection results of the image to be detected are obtained by detecting the neck feature map using several detection heads.

[0194] This system first determines the backbone feature map of the image to be detected through several high-frequency residual-enhanced spatial depth transformation convolutional units and several spatial-frequency context enhancement layer aggregation units. By introducing high-frequency residual-enhanced spatial depth transformation convolutional units, high-frequency detail information in the image is effectively preserved, avoiding the loss of small target features. Furthermore, by introducing spatial-frequency context enhancement layer aggregation units, spatial and frequency domain context information is fully integrated, improving the discriminative ability of features. Based on the fused features, several frequency-guided semantic injection units determine the neck feature map, thereby strengthening the semantic representation of small targets. Finally, several detection heads detect the neck feature map to obtain the detection result of the image to be detected. Thus, this application significantly improves the detection accuracy of small targets under a lightweight model architecture.

[0195] It should be noted that the system embodiments described above are based on the same inventive concept as the method embodiments described above. Therefore, the relevant content of the method embodiments described above is also applicable to the system embodiments described above, and will not be repeated here.

[0196] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations. The acquisition, storage, use and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0197] like Figure 4 One embodiment of this application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned lightweight small target detection method with multi-domain modeling and semantic embedding enhancement. The electronic device includes:

[0198] At least one battery;

[0199] At least one memory;

[0200] At least one processor;

[0201] At least one program;

[0202] The program is stored in memory, and the processor executes at least one program to implement a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement according to the above embodiments of the present disclosure.

[0203] Electronic devices can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0204] The electronic devices according to embodiments of this application will now be described in detail.

[0205] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.

[0206] The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to implement a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement according to an embodiment of this disclosure.

[0207] The input / output interface 1800 is used to implement information input and output.

[0208] The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0209] Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900);

[0210] The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0211] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the detection method of the pressurized water reactor containment pressure control system described above.

[0212] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0213] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.

Claims

1. A lightweight small target detection method using multi-domain modeling and semantic embedding enhancement, characterized in that, The lightweight small target detection method with multi-domain modeling and semantic embedding enhancement includes: Acquire the image to be detected; The image to be detected is input into a trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model; wherein, the trained lightweight detection model includes several high-frequency residual enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads, and the process of the trained lightweight detection model outputting the detection result of the image to be detected includes: The backbone feature map of the image to be detected is determined by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units. The backbone feature map includes a first backbone feature map, a second backbone feature map, and a third backbone feature map. The plurality of high-frequency residual enhanced spatial depth transformation convolutional units include a first high-frequency residual enhanced spatial depth transformation convolutional unit and a second high-frequency residual enhanced spatial depth transformation convolutional unit. The plurality of spatial-frequency context enhancement layer aggregation units include a first spatial-frequency context enhancement layer aggregation unit and a second spatial-frequency context enhancement layer aggregation unit. The trained lightweight detection model further includes a first C2f unit, a first Conv unit, a first SCDown unit, a second SCDown unit, a C2fCIB unit, a SPPF unit, and an AIFI unit. The determination of the backbone feature map of the image to be detected by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units includes: The first high-frequency residual enhanced spatial depth transformation convolutional unit extracts features from the image to be detected to obtain a first downsampled feature map. The process of extracting features from the image to be detected to obtain the first downsampled feature map includes: The pixel blocks at adjacent spatial locations in the image to be detected are rearranged into the channel to obtain the rearranged feature map; The base downsampled feature map is extracted from the rearranged feature map by a 3x3 convolution; Applying the Laplacian operator to the rearranged feature map yields a high-frequency response feature map; High-frequency enhanced feature maps are extracted from high-frequency response features using 1x1 convolution; The first downsampled feature map is obtained by adding the basic downsampled feature map and the high-frequency enhanced feature map element by element. The second downsampled feature map is obtained by extracting features from the first downsampled feature map through the second high-frequency residual enhanced spatial depth transformation convolution unit; The first transition feature map is obtained by extracting features from the second downsampled feature map using the first C2f unit. The first transition feature map is extracted by the first Conv unit to obtain the third downsampled feature map; The first spatial-frequency context enhancement layer aggregation unit extracts features from the third downsampled feature map to obtain the first backbone feature map. The first spatial-frequency context enhancement layer aggregation unit includes an efficient channel attention mechanism block and several phantom convolution blocks. The process of extracting features from the third downsampled feature map to obtain the first backbone feature map includes: The first intermediate feature map is extracted from the third downsampled feature map by a 1x1 convolution; The first intermediate feature map is subjected to reparameterized convolution to obtain the second intermediate feature map; In the case of dividing the second intermediate feature map along the channel dimension to obtain the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map, the local spatial feature map is extracted from the first intermediate sub-feature map by 3x3 convolution. The context extension feature map is extracted from the second intermediate sub-feature map by a 3x3 dilated convolution with a dilation rate of 2; Frequency structure feature map is extracted from the third intermediate sub-feature map by wavelet convolution; Channel supplementary feature maps are extracted from the fourth intermediate sub-feature map using a 1x1 convolution; The local spatial feature map, the context extended feature map, the frequency structure feature map, and the channel supplementary feature map are concatenated to obtain the concatenated feature map; The efficient channel attention feature map is extracted from the stitched feature map using the efficient channel attention mechanism block; Spatial-frequency fusion feature maps are extracted from the efficient channel attention feature maps using 1x1 convolution; If the first spatial-frequency phantom convolutional feature map is extracted from the spatial-frequency fusion feature map through the first phantom convolutional block, the second spatial-frequency phantom convolutional feature map is extracted from the first spatial-frequency phantom convolutional feature map through the second phantom convolutional block, the third spatial-frequency phantom convolutional feature map is extracted from the second spatial-frequency phantom convolutional feature map through the third phantom convolutional block, and so on, until the corresponding (n-1)th spatial-frequency phantom convolutional feature map is extracted from the (n-2)th spatial-frequency phantom convolutional feature map through the (n-n-1)th phantom convolutional block, and the channel adjustment feature map is extracted from the (n-n-1)th spatial-frequency phantom convolutional feature map through a 1x1 convolution, where n-1 is the total number of phantom convolutional blocks; By concatenating the first intermediate feature map, the second intermediate feature map, the spatial-frequency fusion feature map, all the spatial-frequency phantom convolutional feature maps, and the channel adjustment feature map, an aggregated feature map is obtained. The first backbone feature map is extracted from the aggregated feature map by a 1x1 convolution; The first SCDown unit is used to downsample the feature map of the first main road to obtain a fourth downsampled feature map; The second backbone feature map is obtained by extracting features from the fourth downsampled feature map through the second spatial-frequency context enhancement layer aggregation unit; The second SCDown unit is used to downsample the second backbone feature map to obtain the fifth downsampled feature map; The second transition feature map is obtained by extracting features from the fifth downsampled feature map using the C2fCIB unit. The second transition feature map is pooled using the SPPF unit to obtain the third transition feature map; The third main feature map is obtained by performing feature interaction on the third transition feature map through the AIFI unit; The backbone feature maps are fused using a feature pyramid fusion method to obtain a fused feature map. The fused feature map includes a first fused feature map, a second fused feature map, and a third fused feature map with sequentially decreasing resolution. The plurality of spatial-frequency context enhancement layer aggregation units further includes a third spatial-frequency context enhancement layer aggregation unit and a fourth spatial-frequency context enhancement layer aggregation unit. The trained lightweight detection model further includes a second Conv unit, a third SCDown unit, a second C2f unit, and a third C2f unit. The process of fusing the backbone feature maps using the feature pyramid fusion method to obtain the fused feature map includes: The third backbone feature map is upsampled to obtain the upsampled third backbone feature map; The upsampled third backbone feature map is spliced ​​with the second backbone feature map to obtain the first spliced ​​backbone feature map; The second C2f unit is used to extract features from the first spliced ​​backbone feature map to obtain the intermediate fused feature map. The intermediate fused feature map is upsampled to obtain an upsampled intermediate fused feature map; The upsampled intermediate fused feature map is spliced ​​with the first backbone feature map to obtain the second spliced ​​backbone feature map; The first fused feature map is obtained by extracting features from the second spliced ​​backbone feature map through the third C2f unit. The first fused feature map is downsampled by the second Conv unit to obtain the downsampled first fused feature map; The first fused feature map after downsampling is spliced ​​together with the intermediate fused feature map to obtain the third spliced ​​backbone feature map; The third spliced ​​backbone feature map is extracted by the third spatial-frequency context enhancement layer aggregation unit to obtain the second fused feature map. The second fused feature map is downsampled by the third SCDown unit to obtain the downsampled second fused feature map. The second fused feature map after downsampling is spliced ​​with the third backbone feature map to obtain the fourth spliced ​​backbone feature map; The third fused feature map is obtained by extracting features from the fourth spliced ​​backbone feature map through the fourth spatial-frequency context enhancement layer aggregation unit. Based on the fused features, a neck feature map is determined through the plurality of frequency-guided semantic injection units. The neck feature map includes a high-resolution neck feature map, a medium-resolution neck feature map, and a low-resolution neck feature map. The plurality of frequency-guided semantic injection units include a first frequency-guided semantic injection unit, a second frequency-guided semantic injection unit, and a third frequency-guided semantic injection unit. The trained lightweight detection model further includes a fourth C2f unit, a fifth C2f unit, and a sixth C2f unit. The determination of the neck feature map based on the fused features through the plurality of frequency-guided semantic injection units includes: Based on the second fused feature map and the third fused feature map, an intermediate enhanced feature map is determined by the first frequency-guided semantic injection unit; Based on the first fused feature map and the second fused feature map, the top-level intermediate feature map is determined by the second frequency-guided semantic injection unit; Based on the intermediate enhanced feature map and the top intermediate feature map, the top enhanced feature map is determined by the third frequency-guided semantic injection unit; The high-resolution neck feature map is obtained by extracting features from the top-level enhanced feature map using the fourth C2f unit. The intermediate enhanced feature map is obtained by extracting features from the fifth C2f unit; The low-resolution neck feature map is obtained by extracting features from the third fused feature map using the sixth C2f unit. The detection results of the image to be detected are obtained by detecting the neck feature map using the aforementioned detection heads.

2. The lightweight small target detection method with multi-domain modeling and semantic embedding enhancement according to claim 1, characterized in that, The training process of the trained lightweight detection model includes: Obtain a training dataset, wherein the training dataset includes training images and a set of ground truth object bounding boxes; A teacher model and an initial student model are constructed. The teacher model includes several first initial high-frequency residual enhanced spatial depth transformation convolutional units, several first initial spatial-frequency context enhancement layer aggregation units, several first initial frequency-guided semantic injection units, and several first initial detection heads. The initial student model includes several second initial high-frequency residual enhanced spatial depth transformation convolutional units, several second initial spatial-frequency context enhancement layer aggregation units, several second initial frequency-guided semantic injection units, and several second initial detection heads. The training dataset is input into the teacher model to determine the first backbone feature map of the training dataset through the teacher model’s plurality of first initial high-frequency residual enhanced spatial depth transformation convolutional units and plurality of first initial spatial-frequency context enhancement layer aggregation units; The training dataset is input into the initial student model to determine the second backbone feature map of the training dataset through the plurality of second initial high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of second initial spatial-frequency context enhancement layer aggregation units of the initial student model; Based on the set of real target bounding boxes, the first backbone feature map and the second backbone feature map, the backbone network parameters of the initial student model are updated to obtain the first student model. The first student model includes several third initial high-frequency residual enhancement spatial depth transformation convolutional units, several third initial spatial-frequency context enhancement layer aggregation units, several third initial frequency guided semantic injection units and several third initial detection heads. In the case where the first backbone feature map is fused using the feature pyramid fusion method through the teacher model to obtain the first fused feature map, the first neck feature map is determined based on the first fused feature map and guided by the several first initial frequencies of semantic injection units. The training dataset is input into the first student model to determine the third backbone feature map of the training dataset through the plurality of third initial high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of third initial spatial-frequency context enhancement layer aggregation units of the first student model. In the case where the third backbone feature map is fused using the feature pyramid fusion method through the first student model to obtain the second fused feature map, the second neck feature map is determined based on the second fused feature map and guided by the several third initial frequency semantic injection units. Based on the set of real target bounding boxes, the first neck feature map, and the second neck feature map, the neck network parameters and detection head parameters of the first student model are updated to obtain the trained lightweight detection model.

3. A lightweight small target detection system with multi-domain modeling and semantic embedding enhancement, characterized in that, The lightweight small target detection system with multi-domain modeling and semantic embedding enhancement includes: The data construction module is used to acquire the image to be detected; An image detection module is used to input the image to be detected into a trained lightweight detection model to obtain the detection result of the image to be detected output by the trained lightweight detection model; wherein, the trained lightweight detection model includes several high-frequency residual enhanced spatial depth transformation convolutional units, several spatial-frequency context enhancement layer aggregation units, several frequency-guided semantic injection units, and several detection heads, and the process of the trained lightweight detection model outputting the detection result of the image to be detected includes: The backbone feature map of the image to be detected is determined by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units. The backbone feature map includes a first backbone feature map, a second backbone feature map, and a third backbone feature map. The plurality of high-frequency residual enhanced spatial depth transformation convolutional units include a first high-frequency residual enhanced spatial depth transformation convolutional unit and a second high-frequency residual enhanced spatial depth transformation convolutional unit. The plurality of spatial-frequency context enhancement layer aggregation units include a first spatial-frequency context enhancement layer aggregation unit and a second spatial-frequency context enhancement layer aggregation unit. The trained lightweight detection model further includes a first C2f unit, a first Conv unit, a first SCDown unit, a second SCDown unit, a C2fCIB unit, a SPPF unit, and an AIFI unit. The determination of the backbone feature map of the image to be detected by the plurality of high-frequency residual enhanced spatial depth transformation convolutional units and the plurality of spatial-frequency context enhancement layer aggregation units includes: The first high-frequency residual enhanced spatial depth transformation convolutional unit extracts features from the image to be detected to obtain a first downsampled feature map. The process of extracting features from the image to be detected to obtain the first downsampled feature map includes: The pixel blocks at adjacent spatial locations in the image to be detected are rearranged into the channel to obtain the rearranged feature map; The base downsampled feature map is extracted from the rearranged feature map by a 3x3 convolution; Applying the Laplacian operator to the rearranged feature map yields a high-frequency response feature map; High-frequency enhanced feature maps are extracted from high-frequency response features using 1x1 convolution; The first downsampled feature map is obtained by adding the basic downsampled feature map and the high-frequency enhanced feature map element by element. The second downsampled feature map is obtained by extracting features from the first downsampled feature map through the second high-frequency residual enhanced spatial depth transformation convolution unit; The first transition feature map is obtained by extracting features from the second downsampled feature map using the first C2f unit. The first transition feature map is extracted by the first Conv unit to obtain the third downsampled feature map; The first spatial-frequency context enhancement layer aggregation unit extracts features from the third downsampled feature map to obtain the first backbone feature map. The first spatial-frequency context enhancement layer aggregation unit includes an efficient channel attention mechanism block and several phantom convolution blocks. The process of extracting features from the third downsampled feature map to obtain the first backbone feature map includes: The first intermediate feature map is extracted from the third downsampled feature map by a 1x1 convolution; The first intermediate feature map is subjected to reparameterized convolution to obtain the second intermediate feature map; In the case of dividing the second intermediate feature map along the channel dimension to obtain the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map, the local spatial feature map is extracted from the first intermediate sub-feature map by 3x3 convolution. The context extension feature map is extracted from the second intermediate sub-feature map by a 3x3 dilated convolution with a dilation rate of 2; Frequency structure feature map is extracted from the third intermediate sub-feature map by wavelet convolution; Channel supplementary feature maps are extracted from the fourth intermediate sub-feature map using a 1x1 convolution; The local spatial feature map, the context extended feature map, the frequency structure feature map, and the channel supplementary feature map are concatenated to obtain the concatenated feature map; The efficient channel attention feature map is extracted from the stitched feature map using the efficient channel attention mechanism block; Spatial-frequency fusion feature maps are extracted from the efficient channel attention feature maps using 1x1 convolution; If the first spatial-frequency phantom convolutional feature map is extracted from the spatial-frequency fusion feature map through the first phantom convolutional block, the second spatial-frequency phantom convolutional feature map is extracted from the first spatial-frequency phantom convolutional feature map through the second phantom convolutional block, the third spatial-frequency phantom convolutional feature map is extracted from the second spatial-frequency phantom convolutional feature map through the third phantom convolutional block, and so on, until the corresponding (n-1)th spatial-frequency phantom convolutional feature map is extracted from the (n-2)th spatial-frequency phantom convolutional feature map through the (n-n-1)th phantom convolutional block, and the channel adjustment feature map is extracted from the (n-n-1)th spatial-frequency phantom convolutional feature map through a 1x1 convolution, where n-1 is the total number of phantom convolutional blocks; By concatenating the first intermediate feature map, the second intermediate feature map, the spatial-frequency fusion feature map, all the spatial-frequency phantom convolutional feature maps, and the channel adjustment feature map, an aggregated feature map is obtained. The first backbone feature map is extracted from the aggregated feature map by a 1x1 convolution; The first SCDown unit is used to downsample the feature map of the first main road to obtain a fourth downsampled feature map; The second backbone feature map is obtained by extracting features from the fourth downsampled feature map through the second spatial-frequency context enhancement layer aggregation unit; The second SCDown unit is used to downsample the second backbone feature map to obtain the fifth downsampled feature map; The second transition feature map is obtained by extracting features from the fifth downsampled feature map using the C2fCIB unit. The second transition feature map is pooled using the SPPF unit to obtain the third transition feature map; The third main feature map is obtained by performing feature interaction on the third transition feature map through the AIFI unit; The backbone feature maps are fused using a feature pyramid fusion method to obtain a fused feature map. The fused feature map includes a first fused feature map, a second fused feature map, and a third fused feature map with sequentially decreasing resolution. The plurality of spatial-frequency context enhancement layer aggregation units further includes a third spatial-frequency context enhancement layer aggregation unit and a fourth spatial-frequency context enhancement layer aggregation unit. The trained lightweight detection model further includes a second Conv unit, a third SCDown unit, a second C2f unit, and a third C2f unit. The process of fusing the backbone feature maps using the feature pyramid fusion method to obtain the fused feature map includes: The third backbone feature map is upsampled to obtain the upsampled third backbone feature map; The upsampled third backbone feature map is spliced ​​with the second backbone feature map to obtain the first spliced ​​backbone feature map; The second C2f unit is used to extract features from the first spliced ​​backbone feature map to obtain the intermediate fused feature map. The intermediate fused feature map is upsampled to obtain an upsampled intermediate fused feature map; The upsampled intermediate fused feature map is spliced ​​with the first backbone feature map to obtain the second spliced ​​backbone feature map; The first fused feature map is obtained by extracting features from the second spliced ​​backbone feature map through the third C2f unit. The first fused feature map is downsampled by the second Conv unit to obtain the downsampled first fused feature map; The first fused feature map after downsampling is spliced ​​together with the intermediate fused feature map to obtain the third spliced ​​backbone feature map; The third spliced ​​backbone feature map is extracted by the third spatial-frequency context enhancement layer aggregation unit to obtain the second fused feature map. The second fused feature map is downsampled by the third SCDown unit to obtain the downsampled second fused feature map. The second fused feature map after downsampling is spliced ​​with the third backbone feature map to obtain the fourth spliced ​​backbone feature map; The third fused feature map is obtained by extracting features from the fourth spliced ​​backbone feature map through the fourth spatial-frequency context enhancement layer aggregation unit. Based on the fused features, a neck feature map is determined through the plurality of frequency-guided semantic injection units. The neck feature map includes a high-resolution neck feature map, a medium-resolution neck feature map, and a low-resolution neck feature map. The plurality of frequency-guided semantic injection units include a first frequency-guided semantic injection unit, a second frequency-guided semantic injection unit, and a third frequency-guided semantic injection unit. The trained lightweight detection model further includes a fourth C2f unit, a fifth C2f unit, and a sixth C2f unit. The determination of the neck feature map based on the fused features through the plurality of frequency-guided semantic injection units includes: Based on the second fused feature map and the third fused feature map, an intermediate enhanced feature map is determined by the first frequency-guided semantic injection unit; Based on the first fused feature map and the second fused feature map, the top-level intermediate feature map is determined by the second frequency-guided semantic injection unit; Based on the intermediate enhanced feature map and the top intermediate feature map, the top enhanced feature map is determined by the third frequency-guided semantic injection unit; The high-resolution neck feature map is obtained by extracting features from the top-level enhanced feature map using the fourth C2f unit. The intermediate enhanced feature map is obtained by extracting features from the fifth C2f unit; The low-resolution neck feature map is obtained by extracting features from the third fused feature map using the sixth C2f unit. The detection results of the image to be detected are obtained by detecting the neck feature map using the aforementioned detection heads.

4. An electronic device, characterized in that, It includes at least one controller and a memory for communicatively connecting with the controller; the memory stores instructions executable by the at least one controller, which, when executed by the at least one controller, causes the at least one controller to perform a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement as described in any one of claims 1 to 2.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a lightweight small target detection method with multi-domain modeling and semantic embedding enhancement as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • CN120071123A

  • CN120689304A