Image detection method and device, electronic equipment, storage medium and program product
By performing feature extraction, multi-scale fusion processing and feature enhancement processing on the image to be detected, the problem of low image detection accuracy in the prior art is solved, and the accuracy and adaptability of image detection are significantly improved.
Patent Information
- Application Number
- CN202411865355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, images are detected through manual design features, and the detection accuracy is low. In addition, deep learning models often retain noise information when acquiring deep information of the image, ignoring global information, resulting in poor detection effect.
By extracting the image to be detected, a plurality of first feature maps are obtained; performing multi-scale fusion processing on the multiple first feature maps to obtain a plurality of second feature maps; performing feature enhancement processing on the multiple first feature maps to obtain a third feature map; and determining the detection result of the image to be detected based on the multiple second feature maps and the third feature map.
Through multi-scale feature fusion and feature enhancement processing, the expression ability of global and local information of the image to be detected is enhanced, the accuracy of image detection is significantly improved, and it is adapted to the detection of various complex images.
Smart Images

Figure CN119941630A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image detection method, device, electronic device, storage medium and program product. Background Art
[0002] With the rapid development of computer vision technology, image detection has been widely used in many fields, such as security monitoring, medical diagnosis, autonomous driving, etc. In the existing technology, images are detected by manually designed features, and the detection accuracy is low. Summary of the invention
[0003] To solve related technical problems, the embodiments of the present application provide an image detection method, device, electronic device, storage medium and program product, which can improve the accuracy of image detection.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides an image detection method, the method comprising:
[0006] Extracting features from the image to be detected to obtain a plurality of first feature maps;
[0007] Performing multi-scale fusion processing on the plurality of the first feature maps to obtain a plurality of second feature maps;
[0008] Performing feature enhancement processing on the plurality of the first feature maps to obtain a third feature map;
[0009] Based on the multiple second feature maps and the third feature map, a detection result of the image to be detected is determined.
[0010] The present application also provides an image detection device, the device comprising:
[0011] An image acquisition module, used for acquiring an image to be detected;
[0012] A feature extraction module, used to extract features from the image to be detected to obtain a plurality of first feature maps;
[0013] A multi-scale fusion module, used for performing multi-scale fusion processing on the plurality of the first feature maps to obtain a plurality of second feature maps;
[0014] A feature enhancement module, used for performing feature enhancement processing on the plurality of the first feature maps to obtain a third feature map;
[0015] A detection module is used to determine a detection result of the image to be detected based on the multiple second feature maps and the third feature map.
[0016] The embodiment of the present application also provides an electronic device, including: a processor and a communication interface; wherein,
[0017] The processor is used to obtain an image to be detected; perform feature extraction on the image to be detected to obtain multiple first feature maps; perform multi-scale fusion processing on the multiple first feature maps to obtain multiple second feature maps; perform feature enhancement processing on the multiple first feature maps to obtain a third feature map; and determine a detection result of the image to be detected based on the multiple second feature maps and the third feature map.
[0018] An embodiment of the present application also provides an electronic device, including a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps of any one of the above methods when running the computer program.
[0019] An embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0020] An embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0021] In the image detection method, device, electronic device, storage medium and program product provided by the embodiments of the present application, a plurality of first feature maps obtained by feature extraction of the image to be detected are subjected to multi-scale fusion processing to obtain a plurality of second feature maps, thereby enhancing the ability to express the global information of the image to be detected, and a first feature map is subjected to feature enhancement processing to obtain a third feature map, thereby improving the ability to express the local information of the image to be detected. Based on the plurality of second feature maps and the third feature map, the detection result of the image to be detected is determined. The above scheme, through multi-scale feature fusion and feature enhancement processing, enhances the ability to express the global and local information of the image to be detected, so that the image detection method is suitable for the detection of various complex images, thereby significantly improving the accuracy of image detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The process diagram of the image detection method of the present application is as follows: Figure 1 ;
[0023] Figure 2 A schematic diagram of the architecture of an image detection model according to an embodiment of the present application;
[0024] Figure 3 Schematic diagram of regional convolution of the present embodiment;
[0025] Figure 4is a schematic diagram of the structure of the characteristic reinforcement layer provided in the embodiment of the present application;
[0026] Figure 5 is a schematic diagram of the operation of the analysis convolution layer provided in an embodiment of the present application;
[0027] Figure 6 is a schematic diagram of a process of performing feature fusion on a feature map in an embodiment of the present application;
[0028] Figure 7 A schematic diagram of the structure of an image detection device according to an embodiment of the present application;
[0029] Figure 8 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0032] Sensitive image detection is an important image processing detection technology, which aims to achieve real-time detection of images by automated means. Existing sensitive image detection methods mainly include methods based on human skin color recognition, methods based on manual features, and methods based on deep learning.
[0033] Methods based on human skin color recognition have the advantages of being simple and convenient. The focus of this type of method is the proportion of human skin color in the image, that is, the higher the proportion of human skin color, the greater the probability that the image is a sensitive image. However, this type of method has obvious disadvantages: skin color images such as sand piles and earth walls can easily cause this method to fail, resulting in false alarms; some photos with a large proportion of human skin color, such as facial close-ups, bikini photos, etc., can also cause false alarms.
[0034] The method based on manual features introduces manually designed sensitive image features on the basis of the former, which has a certain degree of improvement. The detection rate of this method is significantly higher than other methods in specific situations, but the performance of this method is highly correlated with the designed manual features. In practical applications, designers need to spend a lot of energy to design features. At the same time, due to the complexity of photographic conditions and the diversity of sensitive content, the manually designed features become extremely unreliable, which also means that a lot of manpower and material resources are needed to maintain the model in the later stage.
[0035] The deep learning method is a method that uses neural networks to extract image information, which can automatically learn the features and relationships in the image. This method can effectively extract related information and has better recognition and detection effects in the complex and changing Internet environment, but this method requires a large amount of training data and computing resources.
[0036] In the process of acquiring deep information of images, existing deep learning models often retain a large amount of noise information that is irrelevant to the subject, which will affect the training of the model. At the same time, existing models pay more attention to local information and ignore the influence of global information, which is not suitable for the detection of sensitive images.
[0037] Based on this, the embodiments of the present application propose an image detection method. In each embodiment of the present application, an image to be detected is obtained; feature extraction is performed on the image to be detected to obtain multiple first feature maps; multi-scale fusion processing is performed on the multiple first feature maps to obtain multiple second feature maps; feature enhancement processing is performed on the multiple first feature maps to obtain a third feature map; based on the multiple second feature maps and the third feature map, the detection result of the image to be detected is determined. The embodiments of the present application enhance the global and local information expression capabilities of the image to be detected through multi-scale feature fusion and feature enhancement processing, so that the image detection method is suitable for the detection of various complex images, thereby significantly improving the accuracy of image detection.
[0038] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0039] An embodiment of the present application provides an image detection method, which is applied to an electronic device, where the electronic device includes a server and / or a terminal. The server can be deployed in the cloud (cloud server) or on the device side (edge server). Figure 1 The process diagram of the image detection method of the present application is as follows: Figure 1 ,like Figure 1 As shown, including:
[0040] Step 101: Acquire an image to be detected.
[0041] In the embodiment of the present application, image to be detected refers to the image data that needs to be analyzed and processed in the image detection task.Image to be detected can be image data of any format, such as three color channels (Red, Green, Blue, RGB) images, grayscale images, etc. The acquisition method of the image to be detected is not limited by the embodiment of the present application, and the image to be detected can be obtained by a variety of ways.For example, the image to be detected can be each frame image in the video stream obtained from the camera in real time, the picture uploaded by the user through various channels (such as social media, content sharing platform, etc.) or the image downloaded from the network, etc.
[0042] For example, in the sensitive image detection scenario, the real-time video stream in the live broadcast platform can be extracted frame by frame, and each frame of the image is used as the image to be detected. In the social platform or content review system, the pictures uploaded by the user can be directly used as the image to be detected.
[0043] Step 102: extract features from the image to be detected to obtain a plurality of first feature maps.
[0044] In an embodiment of the present application, the image to be detected can be input into a trained image detection model, and the pixel convolution layer in the image detection model performs feature extraction on the image to be detected to obtain multiple first feature maps.
[0045] Based on this, in one embodiment, feature extraction is performed on the image to be detected to obtain multiple first feature maps, including: feature extraction is performed on the image to be detected at at least one resolution to obtain a first feature map at each resolution.
[0046] Here, the pixel convolution layer in the image detection model can be used to extract feature maps of different levels and different spatial resolutions of the image to be detected. The resolution of the feature map of the low level is high, and the resolution of the feature map of the high level is low. The pixel convolution layer in the image detection model may include multiple convolution layers and multiple pooling layers, and the image to be detected is input into the pixel convolution layer. The multiple convolution layers and multiple pooling layers in the pixel convolution layer perform convolution and maximum pooling processing on the image to be detected, and gradually extract the high-level features of the image to be detected to obtain multiple first feature maps. Among them, each convolution layer outputs a group of first feature maps, and each pooling layer outputs a group of first feature maps. A group of first feature maps output by the same convolution layer may include one or more first feature maps, and the resolution of the one or more first feature maps is the same. A group of first feature maps output by the same pooling layer may include one or more first feature maps, and the resolution of the one or more first feature maps is the same. Since the maximum pooling operation of the pooling layer reduces the size of the image, the resolution of the image after being processed by the pooling layer is reduced. The pixel convolution layer in the image detection model can be used to extract features of the image to be detected at at least one resolution to obtain the first feature map at each resolution. It should be noted that, in the embodiment of the present application, the input itself, that is, the image to be detected, is also used as a first feature map.
[0047] Exemplarily, the pixel convolution layer is specifically a feature extraction network composed of 4 convolution layers and 3 maximum pooling layers. Among them, each layer of pooling operation is to change the input image to 1 / k of the original size, where k is the convolution kernel size. The convolution kernel size k of the maximum pooling layer is 2, which can not only greatly reduce the size of the feature map and reduce the amount of training in the later stage of the image detection model, but also control the reduction range to not be too large to cause excessive information loss. After the image to be detected is input into the pixel convolution layer, the first feature maps of 9 different layers are output. The first feature maps of 9 different layers include the image to be detected, the feature map after the first convolution layer, the feature map after the first pooling layer, the feature map after the second convolution layer, the feature map after the second pooling layer, the feature map after the third convolution layer, the feature map after the third pooling layer, and the feature map after the last convolution layer.
[0048] The embodiment of the present application generates multiple first feature maps by extracting features from the image to be detected. These first feature maps contain feature information of the image to be detected at different resolutions, providing a basis for subsequent multi-scale fusion and feature enhancement processing, thereby improving the accuracy and robustness of image detection.
[0049] Step 103: Perform multi-scale fusion processing on the multiple first feature maps to obtain multiple second feature maps.
[0050] Here, multi-scale fusion processing refers to the process of fusing feature maps of different scales (resolutions) to generate new feature maps containing multi-scale information in image processing and computer vision tasks. This method can effectively capture local details and global structures in images and improve the model's ability to detect and recognize objects of different sizes.
[0051] In an embodiment of the present application, multi-scale fusion processing refers to fusing multiple first feature maps of different resolutions to obtain multiple second feature maps. Multi-scale fusion processing can be implemented in a variety of ways. For example, it can be implemented by upsampling and downsampling, upsampling the low-resolution first feature map (such as linear interpolation), and downsampling the high-resolution first feature map (such as maximum pooling) to make the resolution of each first feature map the same, and combining multiple first feature maps with the same resolution after sampling by addition, concatenation or other fusion strategies to obtain multiple second feature maps.
[0052] Alternatively, the multi-scale fusion processing can also be implemented by cross-layer connection, that is, combining the shallow (high-resolution) first feature map with the deep (low-resolution) first feature map to form a richer and multi-layered second feature map. In the embodiment of the present application, two first feature maps with different resolutions can be selected from multiple first feature maps for cross-layer connection to obtain a second feature map, and this step is repeated to obtain multiple second feature maps.
[0053] Based on this, in one embodiment, performing multi-scale fusion processing on multiple first feature maps to obtain multiple second feature maps includes: determining each second feature map through the following steps: performing convolution processing on the nth first feature map to obtain a fourth feature map; mapping and connecting the fourth feature map and the Nnth first feature map to obtain a fifth feature map, wherein N is the number of first feature maps, and n is a positive integer that increases successively, 1≤n<N; normalizing the fifth feature map to obtain a sixth feature map; and fusing the sixth feature map with the fifth feature map to obtain a second feature map.
[0054] In an embodiment of the present application, the nth first feature map can be convolved by a 1×1 two-dimensional convolution to obtain a fourth feature map. The number N of first feature maps is the number of levels of the first feature maps of various different levels output by the pixel convolution layer of the image detection model. Taking the pixel convolution layer as an example, which is a feature extraction network composed of 4 convolution layers and 3 maximum pooling layers, the pixel convolution layer outputs 9 first feature maps of different levels, then N=9. Obtain the Nnth first feature map from multiple first feature maps, map the fourth feature map and the Nnth first feature map to obtain the fifth feature map. The mapping connection can be an operation such as concatenation or addition. For example, the fourth feature map is concatenated with the Nnth first feature map to generate a fifth feature map. The concatenation operation can be performed along the channel dimension to generate a new feature map. The fifth feature map can satisfy the following formula (1):
[0055] x i =Max(Conv2D 1×1 (t n ))⊕t N-n Formula (1);
[0056] Among them, x i is the fifth feature map in the process of generating the i-th second feature map, t N-n is the Nnth first feature map, t n is the nth first feature map, Conv2D 1×1 is a 1×1 two-dimensional convolution, Max is a maximization operation, and ⊕ is a mapping connection. The maximization operation Max refers to selecting the most important feature in the fourth feature map.
[0057] The fifth feature map is normalized to ensure that the numerical range of the feature map is within a certain range and to improve the stability of subsequent processing. Normalization can be performed using methods such as batch normalization and layer normalization. The fifth feature map is normalized in the following manner: the fifth feature map is subjected to a 1×1 two-dimensional convolution, and the fifth feature map after the 1×1 convolution is transposed to obtain a transposed fifth feature map. The product of the fifth feature map and the transposed fifth feature map is normalized using a normalization function (Softmax function) to obtain a normalized matrix. The normalized matrix is multiplied by the fifth feature map after the 1×1 convolution to obtain a sixth feature map. The sixth feature map satisfies the following formula (2).
[0058] y i =SoftMax(θ(x i ) T ·φ(x i ))g(x i ) formula (2);
[0059] Among them, y i is the sixth feature map in the process of generating the i-th second feature map, θ, φ, g are 1x1 two-dimensional convolutions, θ(x i ) T is the transposed fifth feature map.
[0060] Exemplarily, the size of the first feature map of the first layer, that is, the image to be detected, is (256×256×3), where 256×256 is width×height and 3 is the number of channels. The number of channels of the first feature map of the first layer is adjusted to 1 through a 1×1 convolution, and the most important feature in the first feature map is selected through a maximization operation to obtain a fourth feature map (256×256×1). The fourth feature map is mapped and connected to the first feature map of the eighth layer (32×32×64) to generate a fifth feature map (32×32×65).
[0061] The sixth feature map and the fifth feature map are fused to obtain the second feature map. The fusion process can use methods such as convolution operation and addition operation to generate the final second feature map. In the embodiment of the present application, after the sixth feature map is subjected to a 1×1 convolution operation, an addition operation is performed with the fifth feature to obtain the second feature map.
[0062] Starting from n=1, multiple second feature maps are determined in sequence through the above steps. Taking m=9 as an example, the first second feature map is obtained after the first feature map mapping of the first layer and the eighth layer is connected, the first second feature map is obtained after the first feature map mapping of the second layer and the seventh layer is connected, the first second feature map is obtained after the first feature map mapping of the third layer and the sixth layer is connected, and the first second feature map is obtained after the first feature map mapping of the fourth layer and the fifth layer is connected. Therefore, a total of 4 second feature maps are generated.
[0063] The embodiment of the present application performs multi-scale fusion processing on multiple first feature maps to generate multiple second feature maps. The second feature maps contain multi-scale feature information, which provides a rich feature basis for subsequent feature enhancement processing and final detection results, thereby improving the accuracy and robustness of image detection.
[0064] Step 104: Perform feature enhancement processing on the multiple first feature maps to obtain a third feature map.
[0065] Here, the first feature map with the smallest resolution is selected from multiple first feature maps for feature enhancement processing to obtain a third feature map. That is, the first feature map of the last level output by the pixel convolution layer is feature enhanced, and the size of the third feature map is the same as the size of the first feature map of the last level. The first feature map of the last level can be feature enhanced by combining regional convolution and adaptive multi-channel fusion convolution to obtain the third feature map. Regional convolution and adaptive multi-channel fusion convolution are described in detail below.
[0066] Based on this, in one embodiment, feature enhancement processing is performed on multiple first feature maps to obtain a third feature map, including: determining the first feature map with the smallest resolution among the multiple first feature maps as the main feature map; dividing the main feature map into regions to obtain multiple regional images; performing convolution processing on the multiple regional images respectively to obtain a seventh feature map corresponding to each regional image; fusing the multiple seventh feature maps to obtain an eighth feature map; and performing convolution processing on the eighth feature map to obtain a third feature map.
[0067] In the embodiment of the present application, taking the first feature map of 9 different levels output by the pixel convolution layer as an example, the first feature map with the smallest resolution output by the 9th level is determined as the main feature map. For example, assuming that the resolutions of multiple first feature maps are 512×384, 256×192, 128×96, 64×48 and 32×24, the 32×24 first feature map is selected as the main feature map. The main feature map is subjected to regional convolution and adaptive multi-channel fusion convolution to obtain the third feature map.
[0068] The regional convolution includes multiple sub-convolutions and one core convolution. Multiple sub-convolutions perform convolution on the regional graphics of their respective regions to extract features, and finally the core convolution extracts and fuses the extracted features of multiple sub-convolutions to obtain the eighth feature map. Regional convolution can extract the most concerned area in the entire Zhu feature map, thereby suppressing other non-core areas, strengthening the expression of features, and further improving the discrimination ability of the image detection model. First, the main feature map is divided into regions to obtain multiple regions, and partial regions are selected from the multiple regions, and the images corresponding to these partial regions are used as multiple regional images. For example, the images corresponding to the regions of the four corners of the main feature map can be determined as regional images. It should be noted that after splicing multiple regional images, it is not necessary to obtain a complete main feature map. Multiple sub-convolutions are used to perform convolution processing on each regional image separately, extract local features of each regional image, and generate the seventh feature map corresponding to each regional image. Fusion of multiple seventh feature maps refers to performing convolution operations on multiple seventh feature maps using core convolution to obtain the eighth feature map.
[0069] Adaptive multi-channel fusion convolution is a method for adaptively selecting the convolution kernel size based on the number of channels of the image to be convolved. The eighth feature map is subjected to adaptive multi-channel fusion convolution processing to obtain the third feature map.
[0070] The embodiment of the present application performs feature enhancement processing on multiple first feature maps to generate third feature maps. These third feature maps enhance the local information expression capability of the feature maps, improve the detection accuracy of complex backgrounds and small target objects, and provide a richer feature basis for subsequent detection results, thereby improving the accuracy and robustness of image detection.
[0071] In one embodiment, convolution processing is performed on the eighth feature map to obtain a third feature map, including: obtaining the number of channels of the eighth feature map, and determining a convolution kernel size based on the number of channels; and convolution processing is performed on the eighth feature map based on a convolution kernel having the convolution kernel size to obtain a third feature map.
[0072] In an embodiment of the present application, the eighth feature map is subjected to adaptive multi-channel fusion convolution processing to obtain the third feature map in the following process: obtain the number of channels of the eighth feature map and determine the logarithm of the number of channels. Obtain the preset first parameter and second parameter, and determine the convolution kernel size based on the logarithm, the first parameter and the second parameter. The convolution kernel size is the convolution kernel size, and a convolution kernel with the convolution kernel size is selected to perform convolution processing on the eighth feature map to obtain the third feature map. It should be noted that the convolution kernel size can only be an odd number. The process of determining the convolution kernel size satisfies the following formula (3).
[0073]
[0074] Among them, k is the convolution kernel size, c is the number of channels of the eighth feature map, γ is the first parameter, b is the second parameter, log2(c) is the logarithm of the number of channels c, odd means that the convolution kernel size can only be an odd number, the first parameter is set to 2, and the second parameter is set to 1.
[0075] Exemplarily, the number of channels of the eighth feature map is 1024, and the calculated convolution kernel size k=3 is used to perform convolution processing on the eighth feature map (16×16×1024) to obtain a third feature map (32×32×1024).
[0076] The embodiment of the present application selects a suitable convolution kernel size for feature maps with different numbers of channels, thereby avoiding the problem of too many channels and too slow a convolution kernel size causing the speed to be too slow, or too few channels and too large a convolution kernel causing poor fusion.
[0077] Step 105: Determine a detection result of the image to be detected based on the multiple second feature maps and the third feature map.
[0078] Based on this, in one embodiment, based on multiple second feature maps and third feature maps, the detection result of the image to be detected is determined, including: performing convolution processing on the multiple second feature maps and the third feature maps to obtain a ninth feature map; performing feature fusion processing on the ninth feature map to obtain a feature vector; performing mapping processing on the feature vector to obtain a prediction score of the image to be detected; when the prediction score is greater than a preset threshold, determining a detection result that characterizes the image to be detected as an abnormal image.
[0079] Here, convolution processing is performed on multiple second feature maps and third feature maps to obtain a ninth feature map, which can be achieved in the following manner: the second feature map and the third feature map are convolutionally processed by the analysis convolution layer in the image detection model, wherein the analysis convolution layer includes multiple convolution layers and a maximum pooling layer. First, the third feature map is fused with one of the multiple second feature maps to obtain an input feature map of the maximum pooling layer. The fusion process of the third feature map and the second feature map can refer to the multi-scale fusion in the above embodiment, which is not described here. After the input feature map is subjected to the pooling operation of the maximum pooling layer and the convolution operation of the convolution layer, an output feature map is obtained. The output feature map is fused with another second feature map in the multiple second feature maps as the input feature map of the next maximum pooling layer. Repeat the above process until the analysis convolution layer outputs the ninth feature map.
[0080] The feature fusion processing can adopt global pooling (such as global average pooling, global maximum pooling) and other methods to generate a fixed-length feature vector. The feature fusion processing of the ninth feature map can be implemented in the following manner: performing a maximum pooling operation on the ninth feature map to obtain a ninth feature map after maximum pooling. Performing an average pooling operation on the ninth feature map to obtain a ninth feature map after average pooling. Mapping and connecting the ninth feature map after maximum pooling and the ninth feature map after average pooling to obtain a feature vector. The feature vector can satisfy the following formula (4).
[0081] f=Max(I)·Mean(I) Formula (4);
[0082] Among them, f is the feature vector, Max is the maximum pooling operation, Mean is the average pooling operation, I is the ninth feature map, and · is the mapping connection.
[0083] After obtaining the feature vector, mapping the feature vector refers to inputting the feature vector into the fully connected layer of the image detection model, linearly mapping the feature vector through the weight matrix of the fully connected layer, multiplying the weight matrix and the feature vector, and obtaining the prediction score of the image to be detected. The prediction score refers to a numerical value obtained after the mapping process, which represents the probability that the image to be detected belongs to a certain category. In the embodiment of the present application, the category of the image to be detected includes two categories: abnormal images and normal images. The value range of the prediction score is usually between 0 and 1, and the closer to 1, the more likely the image to be detected is an abnormal image. The process of determining the prediction score satisfies the following formula (5).
[0084] y=wx formula (5);
[0085] Among them, y is the prediction score, w is the weight matrix, and x is the feature vector.
[0086] The preset threshold refers to a fixed value used to determine whether the prediction score is high enough to determine that the image to be detected is an abnormal image in the image detection task. The preset threshold can be set based on experience. For example, if the preset threshold is 0.5, when the prediction score is greater than 0.5, the image to be detected is determined to be an abnormal image, or when the prediction score is less than or equal to 0.5, the image to be detected is determined to be a normal image.
[0087] The embodiment of the present application performs multi-scale fusion processing on multiple first feature maps obtained by feature extraction of the image to be detected, obtains multiple second feature maps, enhances the ability to express the global information of the image to be detected, and performs feature enhancement processing on each first feature map to obtain a third feature map, thereby improving the ability to express the local information of the image to be detected. Based on the multiple second feature maps and the third feature map, the detection result of the image to be detected is determined. The above scheme, through multi-scale feature fusion and feature enhancement processing, enhances the ability to express global and local information of the image to be detected, so that the image detection method is suitable for the detection of various complex images, thereby significantly improving the accuracy of image detection.
[0088] In one embodiment, the image detection method provided in the embodiment of the present application is applied to a pre-trained image detection model, and the method also includes: training the image detection model through the following steps: obtaining multiple first sample data, the first sample data including a first sample image and a sample label, the sample label is used to characterize whether the first sample image is an abnormal image; for each first sample image, when the sample label of the first sample image characterizes that the first sample image is an abnormal image, determining the first sample image as a second sample image; marking abnormal areas in some of the multiple second sample images to obtain corresponding abnormal area images; determining a loss value based on the multiple first sample data and the multiple abnormal area images, and updating the parameters of the image detection model based on the loss value to obtain a trained image detection model.
[0089] In the embodiment of the present application, the specific model structure of the image detection model is not limited, for example, it can be a deep learning model. Exemplarily, 100 first sample data can be obtained, when the first sample image in the first sample data is an abnormal image, the sample label is "1"; when the first sample image in the first sample data is a normal image, the sample label is "0". From the 100 first sample images, assuming that 200 first sample images have a sample label of "1", these 200 first sample images will be determined as second sample images. Select some second sample images from multiple second sample images, for example, select 60 second sample images from 200 second sample images, annotate the abnormal areas in the 60 second sample images, and generate 60 images of abnormal areas with annotations.
[0090] The loss value is used to measure the difference between the prediction result of the image detection model and the true label, that is, to measure the difference between the prediction score output by the image detection model and the sample label. The embodiment of the present application does not limit the calculation method of the loss value, and it is also considered as a cross entropy loss. The image detection model includes a pixel convolution layer and an analysis convolution layer. The loss values of the pixel convolution layer and the analysis convolution layer can be calculated based on multiple first sample data and multiple abnormal area images, respectively. Finally, the loss value of the model is determined based on the loss values of the pixel convolution layer and the analysis convolution layer. The loss value of the pixel convolution layer is calculated based on multiple first sample data with only sample labels and multiple abnormal area images with abnormal area annotations, and the loss value of the analysis convolution layer is calculated based on multiple abnormal area images with abnormal area annotations. The loss value of the pixel convolution layer and the loss value of the analysis convolution layer are weighted based on the weight of the loss value to obtain the loss value of the model. The parameters of the image detection model are updated based on the loss value of the model using the gradient descent and back propagation algorithms to obtain the trained image detection model. The loss value of the model can satisfy the following formula (6).
[0091] L total =λ×γ×L1+(1-λ)×L2 Formula (6);
[0092] Among them, L total is the loss value of the model, L1 is the loss value of the pixel convolution layer, L2 is the loss value of the analysis convolution layer, and λ is the weight of the loss value, which balances the role of the pixel convolution layer and the analysis convolution layer in the entire image detection model. γ is an indicator of whether the abnormal area label exists. When the first sample image is an abnormal area image, γ is 1, and the parameters of the pixel convolution layer and the analysis convolution layer are updated at the same time; when the first sample image only has the first sample label, γ is 0, and the parameters of the pixel convolution layer are no longer updated.
[0093] The embodiment of the present application reduces the use of sample data and improves the training efficiency of the image detection model by mixing a first sample image with only a sample label and an abnormal area image with abnormal area annotated as sample data for training the image detection model. This reduces the manpower and material costs of constructing sample data while maintaining a high accuracy rate of the model.
[0094] The present application is described below in conjunction with application examples.
[0095] The image detection method provided in the embodiment of the present application is a detection method for sensitive images based on deep learning. In the embodiment of the present application, a neural network based on a multi-scale fusion layer and a feature enhancement layer is proposed, and the neural network is trained by a mixed data training method, and the trained neural network is an image detection model. Through further fusion processing of shallow features and deep features, the image detection model improves the recognition accuracy of large targets while taking into account smaller targets. The addition of the prediction module improves the expressive power of the feature itself. After the last few rounds of convolution operations, the prediction results are more accurate. In addition, in order to reduce the use of training data, all weakly labeled data and part of the fully labeled data are used as training samples, which reduces the manpower and material costs of constructing training data while maintaining a high accuracy rate.
[0096] Figure 2 Schematic diagram of the architecture of the image detection model of the present application embodiment. Figure 2 , the image detection model includes a pixel convolution layer 201, a multi-scale fusion layer 202, a feature enhancement layer 203, an analysis convolution layer 204, a feature fusion layer 205 and a label prediction layer 206. The pixel convolution layer 201 is used to extract the pixel-level information of the image, and the feature maps of each layer of the pixel convolution layer are processed by the multi-scale fusion layer 202 and fused with the feature maps of the specific position of the analysis convolution layer 204, so as to improve the model's understanding of global semantics in the later stage. The feature enhancement layer 203 is used between the pixel convolution layer 201 and the analysis convolution layer 204 to enhance the features, and finally the analysis convolution layer 204 judges the input image by integrating the pixel-level information and the image-level information. In addition, the embodiment of the present application also uses a joint loss based on cross entropy for model training. The image detection model is a neural network model based on a multi-scale fusion module and a prediction module trained with mixed data, which enhances the expression ability of sensitive features, improves the corresponding detection rate, ignores the complex and changeable background information, and improves the robustness of the model.
[0097] The image detection model proposed in the embodiment of the present application enhances the robustness of features, improves the image detection model's ability to extract features of different scales, and can better identify sensitive images. Compared with traditional methods, the method proposed in the embodiment of the present application has better accuracy and higher efficiency, and provides new ideas and methods for the study of sensitive images.
[0098] The image detection model provided in the embodiments of the present application is described in detail below.
[0099] Step 1, the pixel convolution layer performs the following operations: First, obtain the image to be detected, use the pixel convolution layer to process the image to be detected, and obtain feature output maps of different levels (corresponding to the first feature map in the above embodiment). The pixel convolution layer is specifically a feature extraction module composed of 4 convolution layers and 3 maximum pooling layers. Among them, the larger convolution kernel has a larger receptive field and can effectively capture complete semantic information. Each layer of pooling operation is that the input image becomes 1 / k of the original size, where k is the convolution kernel size. In this model, the size of k is 2, which can not only greatly reduce the size of the feature map and reduce the amount of training in the later stage of the model, but also control the reduction range to not be too large to cause excessive loss of information. The shallow features after the pixel convolution layer will be sent to the multi-scale fusion layer, and the deep features will be sent to the feature enhancement layer for further enhancement. The parameters of each layer of the pixel convolution layer are shown in Table 1 below.
[0100] Table 1 Parameters of pixel convolution layer
[0101]
[0102] Among them, 2×Conv2D(I1) refers to performing two 5×5 convolution operations and outputting 32 feature maps. Conv2D refers to performing a two-dimensional convolution operation, and Max-pool refers to performing a maximum pooling operation.
[0103] Step 2: The multi-scale fusion layer performs the following operations: The feature maps of 9 different layers obtained through the steps are finally output with 4 features of different sizes (corresponding to the second feature map in the above embodiment) using the multi-scale fusion layer. The specific operation is to map and connect the first layer features with the eighth layer features, and so on. The fusion formula is as follows:
[0104] x i =Max(Conv2D 1×1 (t n ))⊕t 9-n ;
[0105]
[0106] T 9-n =W·y i +x i ;
[0107] Among them, Max is the maximization operation, Conv2D 1×1 ,θ,φ,g,w are 1x1 two-dimensional convolutions. 9-n is the output feature map of the (9-n)th layer, t nis the input feature map of the nth layer, and ⊕ is the mapping connection. The use of a large number of 1x1 convolutions can change the number of channels of the feature map to 1, reducing the amount of calculation. The multi-scale fusion layer deeply fuses the features between different layers to generate a new feature map T 9-n (corresponding to the second feature map in the above embodiment) is input into the corresponding layer of the analysis convolution layer for training.
[0108] Step 3, the feature enhancement layer performs the following operations: the deep feature map outputted from step 1 (corresponding to the main feature map in the above embodiment, i.e., the output feature map of the last convolution layer, with a channel number of 1024) is used to improve the feature expression capability by using the feature enhancement layer, and the output result size is consistent with the input. The feature enhancement layer consists of regional convolution and an adaptive multi-channel fusion convolution, wherein the regional convolution consists of four sub-convolutions and one core convolution, and the four sub-convolution regions will each be convolved to extract features, and finally the core convolution extracts the result of the fusion of the four sub-convolution regions. Regional convolution can extract the most concerned area in the entire feature map, thereby suppressing other non-core areas, enhancing the expression of features, and further improving the discrimination ability of the model. The pixel area required for the core convolution M region is k=2k′+2d+1, wherein the size of each sub-convolution (the side length of the sub-convolution kernel) is k′, and the expansion rate (the spacing between each element of the convolution kernel in the convolution operation, used to expand the receptive field of the convolution kernel without increasing the number of parameters) is d. Figure 3 Schematic diagram of regional convolution according to an embodiment of the present invention.
[0109] Adaptive multi-channel fusion convolution is an improvement on one-dimensional convolution. It uses the characteristics of two-dimensional convolution to highlight the channels that help improve the expression of core features. In order to adapt the model to different numbers of channels, an adaptive formula that can adaptively select the convolution kernel size is designed. That is, the convolution kernel size is: Where k represents the size of the convolution kernel, c represents the number of channels, || odd k can only be an odd number, and γ and b are set to 2 and 1 in the model to change the ratio between the number of channels and the size of the convolution kernel. The adaptive formula can select the appropriate convolution kernel size for feature maps with different numbers of channels, thereby avoiding the problem of too many channels and too slow convolution kernel size, or too few channels and too large convolution kernel, which leads to poor fusion. The result after convolution uses an activation function to obtain the eigenvalues of each channel, so as to make the expression of important features clearer.
[0110] By using adaptive convolution kernels, the image detection model can be modified as needed in the later stage without worrying about whether the feature enhancement layer can achieve the desired effect after the image detection model is modified.
[0111] After the feature map is input into the feature enhancement layer, the four regional convolutions will respectively extract the most discriminative features. The central region M will integrate the results of the four regional convolutions and set the central region feature value as the most discriminative feature among the four regional convolutions, thereby achieving feature enhancement of the feature map.
[0112] Figure 4 It is a structural diagram of the feature enhancement layer provided in an embodiment of the present application. The feature enhancement layer includes regional convolution 301 and adaptive multi-channel fusion convolution 302. The feature enhancement layer cuts in from two angles to improve the feature expression capability. Although both the individual regional convolution or the adaptive multi-channel fusion convolution have been improved, the former lacks the interaction of the channel direction and cannot reduce the influence of other irrelevant channel maps on the results; the latter lacks the feature extraction of a single feature map and only focuses on the channel direction, which greatly reduces the effect. The fusion of the two can ignore some channel maps that focus on background information, making the effect of the entire feature enhancement module more obvious.
[0113] Step 4: The analysis convolution layer performs the following operations: The analysis convolution layer integrates the four feature maps output by the multi-scale fusion layer into the main feature map (a feature map with 1024 channels) respectively, and finally outputs a feature map with 18 channels (corresponding to the ninth feature map in the above embodiment). Figure 5 : This is an operation diagram of the analysis convolution layer provided in the embodiment of the present application. The analysis convolution layer includes multiple convolution layers and pooling layers, each of which performs convolution and pooling operations on the input feature map. The input features of each layer need to be fused with one of the four feature maps output by the multi-scale fusion layer.
[0114] Similar to the pixel convolution layer, the purpose of the pooling operation is still to extract features and reduce the size of the feature map. The feature map output by the final analysis convolution layer is 1 / 64 of the input map. The specific parameter settings of the convolution operation and pooling operation are shown in Table 2.
[0115] Table 2 Analysis of the parameter settings of the convolutional layer
[0116] Convolutional Layer Convolution Kernel Number of features enter 1025 Max-pool 2×2 1025 Conv2D 5×5 9 Max-pool 2×2 8 Conv2D 5×5 9 Max-pool 2×2 16 Conv2D 5×5 17 Max-pool 2×2 17 Conv2D 5×5 18
[0117] Step 5: The feature fusion layer performs the following operations: The analysis convolution layer of step 4 finally outputs a feature map with 18 channels. The feature fusion layer operates on this feature map. The specific operation content is to send the feature map to the maximum pooling and average pooling respectively, and output the final result (corresponding to the feature vector in the above embodiment).
[0118] f = Max(I)·Mean(I);
[0119] Among them, f is the final result, Max represents the maximum pooling operation, Mean represents the average pooling, I is the input feature map, and · is the mapping connection. Maximum pooling can save the most comprehensive semantic information, while average pooling saves more background information. Figure 6 It is a schematic diagram of the process of performing feature fusion on the feature map according to an embodiment of the present application.
[0120] Step 6: The label prediction layer performs the following operations: After obtaining the final result of step 5, the fully connected layer is used to obtain the final result (corresponding to the prediction score in the above embodiment). y = wx, where y represents the one-dimensional variable in the label prediction layer, w represents the weight of the feature, and x is the input feature vector. The final output result y is between 0 and 1. The larger the value, the greater the probability that the image to be detected is a sensitive image.
[0121] Step 7: Train the image detection model. All training data are weakly labeled, that is, only whether the image is a pornographic image is labeled; some images are fully labeled, that is, not only whether it is a pornographic image is labeled, but also sensitive areas need to be labeled. The best effect is achieved when the fully labeled images account for 30% of all images. In addition, the increase in the proportion of fully labeled images will still improve the accuracy to a certain extent, but the degree of improvement will gradually decrease as the number of fully labeled images increases.
[0122] Using the joint cross entropy loss as the loss function, we calculate the cross entropy loss of the pixel convolution layer and the analysis convolution layer respectively, and finally calculate the joint loss. total =λ×γ×L1+(1-λ)×L2, where L1 is the pixel convolution layer, L2 is the analysis convolution layer, λ is used as a balance factor to balance the role of the two layers in the entire model; γ is used as an indicator of whether pixel-level annotation exists. When the image is a fully annotated image, γ is 1, and the parameters of the pixel convolution layer and the analysis convolution layer are updated at the same time; when the image is a weakly annotated image, γ is 0, and the parameters of the pixel convolution layer are no longer updated.
[0123] During actual testing, the image detection method provided by the embodiment of the present application has an average accuracy rate of over 80% and an FPS of 10.2, which meets the actual application standard.
[0124] Compared with the existing methods, the embodiment of the present application utilizes the feature enhancement layer to improve the expression ability of sensitive image features, thereby enhancing the discrimination ability of the image detection model. The multi-scale fusion is used to achieve the recognition ability of objects of different sizes on the basis of low computational complexity. The combined cross entropy loss can realize low-complexity training of the image detection model.
[0125] In order to implement the image detection method of the embodiment of the present application, the embodiment of the present application also provides an image detection device, Figure 7Schematic diagram of the composition structure of the image detection device according to the embodiment of the present application. Figure 7 As shown, the image detection device comprises:
[0126] The image acquisition module 601 is used to acquire the image to be detected; the feature extraction module 602 is used to extract features of the image to be detected to obtain multiple first feature maps; the multi-scale fusion module 603 is used to perform multi-scale fusion processing on the multiple first feature maps to obtain multiple second feature maps; the feature enhancement module 604 is used to perform feature enhancement processing on the multiple first feature maps to obtain a third feature map; the detection module 605 is used to determine the detection result of the image to be detected based on the multiple second feature maps and the third feature map.
[0127] In one embodiment, the feature extraction module 602 is further used to extract features from the image to be detected at at least one resolution to obtain a first feature map at each resolution.
[0128] In one embodiment, the multi-scale fusion module 603 is also used to determine each second feature map through the following steps: convolution processing is performed on the nth first feature map to obtain a fourth feature map; the fourth feature map and the Nnth first feature map are mapped and connected to obtain a fifth feature map, wherein N is the number of first feature maps, n is a successively increasing positive integer, 1≤n<N; normalization processing is performed on the fifth feature map to obtain a sixth feature map; and the sixth feature map and the fifth feature map are fused to obtain a second feature map.
[0129] In one embodiment, the feature enhancement module 604 is further used to determine the first feature map with the smallest resolution among multiple first feature maps as the main feature map; divide the main feature map into regions to obtain multiple regional images; perform convolution processing on the multiple regional images respectively to obtain the seventh feature map corresponding to each regional image; fuse the multiple seventh feature maps to obtain the eighth feature map; and perform convolution processing on the eighth feature map to obtain the third feature map.
[0130] In one embodiment, the feature enhancement module 604 is further used to obtain the number of channels of the eighth feature map, and determine the convolution kernel size based on the number of channels; and perform convolution processing on the eighth feature map based on the convolution kernel with the convolution kernel size to obtain the third feature map.
[0131] In one embodiment, the detection module 605 is further used to perform convolution processing on multiple second feature maps and third feature maps to obtain a ninth feature map; perform feature fusion processing on the ninth feature map to obtain a feature vector; perform mapping processing on the feature vector to obtain a prediction score of the image to be detected; when the prediction score is greater than a preset threshold, determine a detection result that characterizes the image to be detected as an abnormal image.
[0132] In one embodiment, the image detection device also includes a model training module, which is used to train the image detection model through the following steps: obtaining multiple first sample data, the first sample data including a first sample image and a sample label, the sample label is used to characterize whether the first sample image is an abnormal image; for each first sample image, when the sample label of the first sample image characterizes that the first sample image is an abnormal image, the first sample image is determined as a second sample image; abnormal areas in some of the multiple second sample images are marked to obtain corresponding abnormal area images; based on the multiple first sample data and the multiple abnormal area images, a loss value is determined, and based on the loss value, the parameters of the image detection model are updated to obtain a trained image detection model.
[0133] In actual application, the image acquisition module 601, the feature extraction module 602, the multi-scale fusion module 603, the feature enhancement module 604, the detection module 605 and the model updating module can be implemented by a processor in the image detection device.
[0134] It should be noted that: when the image detection device provided in the above embodiment performs image detection, only the division of the above program modules is used as an example. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the image detection device provided in the above embodiment belongs to the same concept as the image detection method embodiment. The specific implementation process is detailed in the image detection method embodiment, which will not be repeated here.
[0135] Based on the hardware implementation of the above program modules, and in order to implement an image detection method provided in an embodiment of the present application, an embodiment of the present application further provides an electronic device, such as Figure 8 As shown, the electronic device 700 includes:
[0136] Communication interface 701, capable of exchanging information with other network nodes;
[0137] The processor 702 is connected to the communication interface 701 to implement information exchange with other network nodes, and is used to execute the method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 703.
[0138] Specifically, the processor 702 is used to obtain an image to be detected; perform feature extraction on the image to be detected to obtain multiple first feature maps; perform multi-scale fusion processing on the multiple first feature maps to obtain multiple second feature maps; perform feature enhancement processing on the multiple first feature maps to obtain a third feature map; and determine the detection result of the image to be detected based on the multiple second feature maps and the third feature map.
[0139] In one embodiment, the processor 702 is further used to train a second image reconstruction model based on the at least one third image and a fourth image corresponding to the at least one third image until a first convergence condition is reached; and to train the second image reconstruction model based on at least one fifth image corresponding to each third image in the at least one third image until a second convergence condition is reached, thereby obtaining the first image reconstruction model.
[0140] In one embodiment, the processor 702 is further configured to perform feature extraction on the image to be detected at at least one resolution to obtain a first feature map at each resolution.
[0141] In one embodiment, the processor 702 is further used to determine each second feature map through the following steps: performing convolution processing on the nth first feature map to obtain a fourth feature map; mapping and connecting the fourth feature map and the Nnth first feature map to obtain a fifth feature map, wherein N is the number of first feature maps, n is a successively increasing positive integer, 1≤n<N; normalizing the fifth feature map to obtain a sixth feature map; and fusing the sixth feature map and the fifth feature map to obtain a second feature map.
[0142] In one embodiment, the processor 702 is further used to determine a first feature map with the smallest resolution among multiple first feature maps as a main feature map; perform region division on the main feature map to obtain multiple regional images; perform convolution processing on the multiple regional images respectively to obtain a seventh feature map corresponding to each regional image; fuse the multiple seventh feature maps to obtain an eighth feature map; and perform convolution processing on the eighth feature map to obtain a third feature map.
[0143] In one embodiment, the processor 702 is further used to obtain the number of channels of the eighth feature map, and determine the convolution kernel size based on the number of channels; and perform convolution processing on the eighth feature map based on the convolution kernel having the convolution kernel size to obtain the third feature map.
[0144] In one embodiment, the processor 702 is further used to perform convolution processing on multiple second feature maps and third feature maps to obtain a ninth feature map; perform feature fusion processing on the ninth feature map to obtain a feature vector; perform mapping processing on the feature vector to obtain a prediction score of the image to be detected; when the prediction score is greater than a preset threshold, determine a detection result that characterizes the image to be detected as an abnormal image.
[0145] In one embodiment, the processor 702 is further used to train the image detection model through the following steps: obtaining multiple first sample data, the first sample data including a first sample image and a sample label, the sample label being used to characterize whether the first sample image is an abnormal image; for each first sample image, when the sample label of the first sample image characterizes that the first sample image is an abnormal image, determining the first sample image as a second sample image; marking abnormal areas in some of the multiple second sample images to obtain corresponding abnormal area images; determining a loss value based on the multiple first sample data and the multiple abnormal area images, and updating the parameters of the image detection model based on the loss value to obtain a trained image detection model.
[0146] It should be noted that the specific processing process of the processor 702 can be understood by referring to the above method.
[0147] Of course, in actual application, the various components in the electronic device 700 are coupled together through the bus system 704. It is understandable that the bus system 704 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 8 Various buses are labeled as bus system 704 .
[0148] The memory 703 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device 700. Examples of such data include: any computer program used to operate on the electronic device 700.
[0149] The method disclosed in the above embodiment of the present application can be applied to the processor 702, or implemented by the processor 702. The processor 702 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 702 or an instruction in the form of software. The above-mentioned processor 702 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 702 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the memory 703, and the processor 702 reads the information in the memory 703 and completes the steps of the above method in combination with its hardware.
[0150] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.
[0151] It can be understood that the memory 703 of the embodiment of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 703 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.
[0152] In an exemplary embodiment, the embodiment of the present application further provides an electronic device, including a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps of any of the above methods when running the computer program.
[0153] The embodiment of the present application further provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 703 storing a computer program, and the computer program can be executed by a processor 702 of an electronic device 700 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0154] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0155] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. The term "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "one or more" herein represents any combination of at least two of any one or more of a plurality of items. For example, including one or more of A, B, and C can represent including any one or at least two or more elements selected from the set consisting of A, B, and C.
[0156] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0157] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. An image detection method, characterized in that: The method comprises: Acquire the image to be detected; Extracting features from the image to be detected to obtain a plurality of first feature maps; Performing multi-scale fusion processing on the plurality of first feature maps to obtain a plurality of second feature maps; Performing feature enhancement processing on the plurality of the first feature maps to obtain a third feature map; Based on the multiple second feature maps and the third feature map, a detection result of the image to be detected is determined.
2. The method according to claim 1, characterized in that The step of extracting features from the image to be detected to obtain a plurality of first feature maps includes: Feature extraction is performed on the image to be detected at at least one resolution to obtain a first feature map at each resolution.
3. The method according to claim 1, characterized in that The performing multi-scale fusion processing on the plurality of the first feature maps to obtain a plurality of second feature maps includes: Each of the second feature maps is determined by the following steps: Perform convolution processing on the nth first feature map to obtain a fourth feature map; Mapping and connecting the fourth feature map and the Nnth first feature map to obtain a fifth feature map, wherein N is the number of the first feature maps, n is a positive integer that increases successively, and 1≤n<N; Normalizing the fifth feature map to obtain a sixth feature map; The sixth feature map and the fifth feature map are fused to obtain the second feature map.
4. The method according to claim 1, characterized in that The performing feature enhancement processing on the plurality of the first feature maps to obtain a third feature map includes: Determine a first feature map having the smallest resolution among the plurality of first feature maps as a main feature map; Performing region division on the main feature map to obtain a plurality of region images; Performing convolution processing on the plurality of regional images respectively to obtain a seventh feature map corresponding to each of the regional images; Fusing a plurality of the seventh feature maps to obtain an eighth feature map; Perform convolution processing on the eighth feature map to obtain the third feature map.
5. The method according to claim 4, characterized in that The performing convolution processing on the eighth feature map to obtain the third feature map includes: Obtaining the number of channels of the eighth feature map, and determining a convolution kernel size based on the number of channels; The eighth feature map is convolved based on a convolution kernel having the convolution kernel size to obtain the third feature map.
6. The method according to claim 1, characterized in that The step of determining the detection result of the image to be detected based on the plurality of second feature maps and the third feature map comprises: Performing convolution processing on the multiple second feature maps and the third feature map to obtain a ninth feature map; Performing feature fusion processing on the ninth feature map to obtain a feature vector; Performing mapping processing on the feature vector to obtain a prediction score of the image to be detected; When the prediction score is greater than a preset threshold, a detection result is determined that indicates that the image to be detected is an abnormal image.
7. The method according to any one of claims 1 to 6, characterized in that: Applied to a pre-trained image detection model, the method further comprises: training the image detection model by the following steps: Acquire a plurality of first sample data, wherein the first sample data includes a first sample image and a sample label, wherein the sample label is used to indicate whether the first sample image is an abnormal image; For each first sample image, when the sample label of the first sample image indicates that the first sample image is an abnormal image, determining the first sample image as a second sample image; marking abnormal regions in some of the second sample images among the plurality of the second sample images to obtain corresponding abnormal region images; Based on the plurality of first sample data and the plurality of abnormal region images, determining a loss value, Based on the loss value, the parameters of the image detection model are updated to obtain a trained image detection model.
8. An image detection device, characterized in that: The device comprises: An image acquisition module, used for acquiring an image to be detected; A feature extraction module, used to extract features from the image to be detected to obtain a plurality of first feature maps; A multi-scale fusion module, used for performing multi-scale fusion processing on the plurality of the first feature maps to obtain a plurality of second feature maps; A feature enhancement module, used for performing feature enhancement processing on the plurality of the first feature maps to obtain a third feature map; A detection module is used to determine a detection result of the image to be detected based on the multiple second feature maps and the third feature map.
9. An electronic device, characterized in that: comprising a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that The computer program implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.