Object Detection Method, Device, Storage Medium and Electronic Device

Through the object detection method of multi-layer residual convolutional layer, the problems of real-time face detection and low computing resource utilization efficiency in the prior art are solved, efficient detection of non-positive faces is achieved, and the real-time and efficiency of passenger flow statistics are improved.

CN114155567BActive Publication Date: 2025-05-30BOE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010933809.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-08
Publication Date
2025-05-30
Estimated Expiration
2040-11-30

AI Technical Summary

Technical Problem

When the prior art applies deep learning algorithms in the field of face detection, it is difficult to effectively detect non-frontal faces, resulting in low real-time customer flow statistics and low computing resource utilization efficiency.

Method used

Using the object detection method of a multi-layer residual convolution layer, feature extraction is performed through the first and second convolutional layers, and then the target feature image is further extracted using the third and fourth convolutional layers, and a plurality of reference preselected images are generated for object detection.

Benefits of technology

This method can improve the real-time and computing efficiency of object detection without adding convolutional layers, enhance the detection ability of non-positive faces, and reduce the waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155567B_ABST
    Figure CN114155567B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image processing technology, and specifically relates to an object detection method and device, a computer-readable storage medium, and an electronic device. The method includes: obtaining an input image, and performing feature extraction on the input image by using a first convolutional layer and a second convolutional layer to obtain a reference feature image; performing feature extraction on the reference feature image by using a third convolutional layer to obtain a first object feature map, and performing feature extraction on the first object feature image by using a fourth convolutional layer to obtain a second object feature image; obtaining a reference preselected image for each point on the first object feature image and the second object feature image; determining a target preselected image from multiple reference preselected images, and completing object detection by using an object detection algorithm. The technical solution of the embodiment of the present disclosure overcomes the deficiencies of the existing object detection methods in wasting computing resources and having poor real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Deep learning algorithms have been quite mature in the field of face detection. However, in the application of passenger flow statistics, not every customer's face is directly facing the camera. To correctly detect the number of people, the head features should be extracted to detect the number of heads in the video, so as to achieve the purpose of passenger flow statistics.

[0003] However, the head features are more complex than the face features. Therefore, simply applying the face detection network does not achieve good results, and it is necessary to increase the number of convolutional kernels in the detection network. However, increasing the number of convolutional kernels will increase the computational burden significantly with limited computing resources at the edge, and the real-time performance is poor.

[0004] Therefore, it is necessary to design a new object detection method.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide an object detection method and device, a computer-readable storage medium, and an electronic device, so as to at least overcome the deficiencies of the existing object detection methods in wasting computing resources and poor real-time performance to a certain extent.

[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned partially through the practice of the present disclosure.

[0008] According to the first aspect of the present disclosure, an object detection method is provided, including:

[0009] Obtain an input image, and perform feature extraction on the input image using a first convolutional layer and a second convolutional layer to obtain a reference feature image;

[0010] Perform feature extraction on the reference feature image using a third convolutional layer to obtain a first target feature map, and perform feature extraction on the first target feature image using a fourth convolutional layer to obtain a second target feature image;

[0011] Obtain a reference preselected image for each point on the first target feature image and the second target feature image;

[0012] Determine a target preselected image from multiple reference preselected images, and complete object detection using an object detection algorithm;

[0013] Wherein the first convolutional layer and the second convolutional layer are both single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer are double-module residual convolutional layers.

[0014] In an exemplary embodiment of the present disclosure, the input image is subjected to feature extraction using a first convolutional layer and a second convolutional layer to obtain a reference feature image, including:

[0015] The input image is subjected to feature extraction using the first convolutional layer to obtain an initial feature image;

[0016] The initial feature image is subjected to feature extraction using the second convolutional layer to obtain a reference feature image.

[0017] In an exemplary embodiment of the present disclosure, the single-module residual convolutional layer includes:

[0018] A first convolutional unit, including a depthwise separable convolutional kernel and a first convolutional kernel in a serial design;

[0019] A residual network unit, including a max pooling subunit and a padding subunit in a serial design;

[0020] Wherein, the first convolutional kernel is a standard convolutional kernel.

[0021] In an exemplary embodiment of the present disclosure, the input image is subjected to feature extraction using the first convolutional layer to obtain an initial feature map;

[0022] The input image is input into the first convolutional unit of the first convolutional layer for feature extraction to obtain a first feature image;

[0023] The input image is input into the residual network unit of the first convolutional layer to obtain a second feature image having the same format as the first feature image;

[0024] The first feature image and the second feature image are added together to obtain the initial feature image.

[0025] In an exemplary embodiment of the present disclosure, the initial feature image is subjected to feature extraction using the second convolutional layer to obtain a reference feature image, including:

[0026] The initial feature image is input into the first convolutional unit of the second convolutional layer for feature extraction to obtain a third feature image;

[0027] The initial feature map is input into the residual network unit of the second convolutional layer to obtain a fourth feature image having the same format as the third feature image;

[0028] The third feature image and the fourth feature image are added together to obtain the reference feature image.

[0029] In an exemplary embodiment of the present disclosure, the dual-module residual convolutional layer includes:

[0030] A second convolution unit, comprising a depthwise separable convolution kernel and a second convolution kernel designed in series;

[0031] A third convolution unit, connected in series with the second convolution sub-unit, wherein the third convolution unit includes a depth-separable convolution kernel and a third convolution kernel designed in series;

[0032] Residual network unit, including a serially designed max pooling subunit and a padding subunit.

[0033] In an exemplary embodiment of the present disclosure, a third convolutional layer is used to extract features from the reference feature image to obtain a first target feature map, including:

[0034] The fifth feature image is obtained by extracting features from the reference feature image using the second convolution unit and the third convolution unit of the third convolution layer.

[0035] Inputting the reference feature image into the residual network unit of the third convolutional layer to obtain a sixth feature image having the same format as the fifth feature image;

[0036] The fifth feature image and the sixth feature image are added together to obtain the first target feature image.

[0037] In an exemplary embodiment of the present disclosure, a fourth convolutional layer is used to extract features from a first target feature image to obtain a second target feature image, including:

[0038] The second convolution unit and the third convolution unit of the fourth convolution layer are used to extract features of the first target feature image to obtain a seventh feature image.

[0039] Inputting the first target feature image into the residual network unit of the fourth convolutional layer to obtain an eighth feature image having the same format as the seventh feature image;

[0040] The seventh feature image and the eighth feature image are added to obtain the second target feature image.

[0041] In an exemplary embodiment of the present disclosure, obtaining a reference pre-selected image of each point on the first target feature image and the second target feature image includes:

[0042] Taking each point of the first target feature image and the second target feature image as a center, at least one reference pre-selected image is generated based on each point.

[0043] In an exemplary embodiment of the present disclosure, determining a target pre-selected image from a plurality of reference pre-selected images, and performing target detection using a target detection algorithm, includes:

[0044] Perform class regression on the features of all reference preselected images respectively, and calculate the class scores of each class;

[0045] The target preselected image can be selected according to the class score;

[0046] Perform non-maximum suppression operation on the target preselected figure to complete target detection.

[0047] According to one aspect of the present disclosure, there is provided an object detection device, including:

[0048] An image acquisition module, configured to acquire an input image, and perform feature extraction on the input image by using a first convolutional layer and a second convolutional layer to obtain a reference feature image;

[0049] A feature extraction module, configured to perform feature extraction on the reference feature image by using a third convolutional layer to obtain a first target feature image, and perform feature extraction on the first feature map by using a fourth convolutional layer to obtain a second target feature image;

[0050] An image selection module, configured to acquire reference preselected images for each point on the first target feature image and the second target feature image;

[0051] An image determination module, configured to determine a target preselected image from multiple reference preselected images, and complete target detection by using an object detection algorithm;

[0052] Wherein the first convolutional layer and the second convolutional layer are both single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer are double-module residual convolutional layers.

[0053] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the object detection method described in any one of the above is implemented.

[0054] According to one aspect of the present disclosure, there is provided an electronic device, including:

[0055] A processor; and

[0056] A memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the object detection method described in any one of the above.

[0057] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects:

[0058] In a target detection method provided by an embodiment of the present disclosure, a reference image is obtained by performing feature extraction on an input image using a first convolutional layer and a second convolutional layer, a first target feature map is obtained by performing feature extraction on the reference image using a third convolutional layer, a second target image is obtained by performing feature extraction on the first feature map using a fourth convolutional layer, a reference preselected image of each point on the first target image and the second target image is acquired, and the target preselected image is determined among multiple budget images to complete target detection, where the first convolutional layer and the second convolutional layer are both single-module residual convolutional kernels, and the third convolutional layer and the fourth convolutional layer are dual-module residual convolutional kernels; compared with the prior art, the first convolutional layer and the second convolutional layer are both single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer are dual-module residual convolutional kernels. By using the above four convolutional layers to obtain the target preselected image, the receptive field can be increased, target detection can be completed without adding convolutional layers, the computational burden is reduced, the computational efficiency is accelerated, and the real-time performance of target detection is enhanced.

[0059] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:

[0061] Figure 1 Schematically showing a flowchart of a target detection method in an exemplary embodiment of the present disclosure;

[0062] Figure 2 Schematically showing a schematic diagram of a single-module residual convolutional layer in an exemplary embodiment of the present disclosure;

[0063] Figure 3 Schematically showing a schematic diagram of a dual-module residual convolutional layer in an exemplary embodiment of the present disclosure;

[0064] Figure 4 Schematically showing a framework diagram of the data flow of a target detection method in an exemplary embodiment of the present disclosure;

[0065] Figure 5 Schematically showing a schematic diagram of the composition of a target detection device in an exemplary embodiment of the present disclosure;

[0066] Figure 6A schematic structural diagram of a computer system of an electronic device suitable for implementing the exemplary embodiments of the present disclosure is shown;

[0067] Figure 7 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is shown. Detailed implementation manners

[0068] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0069] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0070] In the present exemplary embodiment, first, a target detection method is provided. Referring to Figure 1 as shown in, the above-mentioned target detection method may include the following steps:

[0071] S110, obtain an input image, and perform feature extraction on the input image using a first convolutional layer and a second convolutional layer to obtain a reference feature image;

[0072] S120, perform feature extraction on the reference feature image using a third convolutional layer to obtain a first target feature map, and perform feature extraction on the first target feature image using a fourth convolutional layer to obtain a second target feature image;

[0073] S130, obtain a reference preselected image for each point on the first target feature image and the second target feature image;

[0074] S140, determine the video to be played in the event occurrence area according to the weight, where the first convolutional layer and the second convolutional layer are both single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer are double-module residual convolutional layers.

[0075] In the object detection method provided in this exemplary embodiment, compared with the prior art, both the first convolutional layer and the second convolutional layer are single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer become dual-module residual convolutional kernels. By using the above four convolutional layers to obtain a target preselection image, the receptive field can be increased, object detection can be completed without adding convolutional layers, the computational burden is reduced, the computational efficiency is accelerated, and the real-time performance of object detection is enhanced.

[0076] Next, each step of the object detection method in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.

[0077] In step S110, an input image is obtained, and the first convolutional layer and the second convolutional layer are used to perform feature extraction on the input image to obtain a reference feature image.

[0078] In an exemplary embodiment of the present disclosure, the input image can be a square image or a rectangular image, and then the first convolutional layer and the second convolutional layer can be used to perform feature extraction on the input image to obtain a reference feature image.

[0079] In this exemplary embodiment, reference can be made to Figure 2 and Figure 4 As shown, both the first convolutional layer 410 and the second convolutional layer 420 are single-module residual convolutional layers. Among them, the single-module residual convolutional layer includes a first convolutional unit and a residual network unit. The first convolutional unit may include a depthwise separable convolutional kernel 210 and a first convolutional kernel 220 designed in series. The depthwise separable convolutional kernel 210 can be a 5*5 depthwise separable convolutional kernel 210. The size of the depth convolutional kernel can be customized according to user requirements and is not specifically limited in this exemplary embodiment. For example, a 6*6 depthwise separable convolutional kernel 210. Since the specific operation mode of the depthwise separable convolutional kernel 210 is already relatively mature in the related art, it will not be elaborated here.

[0080] In this exemplary embodiment, the first convolutional kernel 220 can be a standard convolutional kernel, and the first convolutional kernel 220 can be a 1*1 standard convolutional kernel in this embodiment.

[0081] In this exemplary embodiment, the residual network unit may include a max-pooling sub-unit 230 and a padding sub-unit 240 designed in series. Among them, the max-pooling sub-unit 230 divides the input image into several rectangular regions and outputs the maximum value for each sub-region, that is, it is used to perform a max-pooling operation on the above input image. The max-pooling operation is already relatively mature in the related art, so it will not be elaborated here. The padding sub-unit 240 is used to pad the feature map obtained after max-pooling, that is, to fill with zeros, to adjust the size of the feature map after max-pooling.

[0082] In the present exemplary embodiment, the input image can be first input into the first convolutional unit of the first convolutional layer 410 for feature extraction to obtain a first feature image. Then, the input image is input into the residual network unit of the first convolutional layer 410. After the above-mentioned max-pooling operation and padding operation, a second feature image with the same format as the first feature image can be obtained. The padding operation can make the second feature image have the same matrix format as the first feature image. Therefore, the first feature image and the second feature image can be directly added as matrices to obtain an initial feature image.

[0083] After obtaining the above-mentioned first feature image, the initial feature image can be input into the first convolutional unit of the second convolutional layer 420 for feature extraction to obtain a third feature image. Then, the initial feature map is input into the residual network unit of the second convolutional layer 420. After the above-mentioned max-pooling operation and padding operation, a fourth feature image with the same format as the third feature image can be obtained; the padding operation can make the third feature image have the same matrix format as the fourth feature image. Therefore, the first feature image and the second feature image can be directly added as matrices to obtain an initial feature image.

[0084] In step S120, the third convolutional layer 430 is used to perform feature extraction on the reference feature image to obtain a first target feature map, and the fourth convolutional layer 440 is used to perform feature extraction on the first target feature image to obtain a second target feature image.

[0085] In an exemplary embodiment of the present disclosure, referring to Figure 3 and Figure 4 as shown, the third convolutional layer and the fourth convolutional layer can be dual-module residual convolutional layers, where the dual-module residual convolutional layer can include a second convolutional unit, a third convolutional unit, and a residual network unit.

[0086] In the present exemplary embodiment, the second convolutional unit can include a depthwise separable convolution kernel 330 and a second convolution kernel 310 in a serial design. The second convolution kernel 310 can include a standard 1*1 convolution kernel, a normalization layer (BN layer), and an activation layer (relu layer). The depthwise separable convolution kernel 330 can be a depthwise separable convolution kernel 330 with a size of 5*5. The size of the depth convolution kernel can be customized according to user requirements and is not specifically limited in the present exemplary embodiment. For example, a depthwise separable convolution kernel 330 with a size of 6*6. Since the specific operation mode of the depthwise separable convolution kernel 330 is already relatively mature in the related art, it will not be elaborated here.

[0087] The third convolutional unit may also include a depthwise separable convolutional kernel 340 and a third convolutional kernel 320 designed in series. The third convolutional kernel 320 may include a 1×1 standard convolutional kernel and a 3×3 standard convolutional kernel. In the present exemplary embodiment, only the 1×1 standard convolutional kernel in the third convolutional kernel 320 is adopted. The depth convolution is the same as the structure and function of the depth convolutional layers in the above-mentioned first convolutional layer 410 and second convolutional layer 420, and has been described in detail above, so it will not be elaborated here.

[0088] In the present exemplary embodiment, the residual network unit may include a max-pooling subunit 230 and a padding subunit 240 designed in series. The max-pooling subunit 230 divides the input image into several rectangular regions and outputs the maximum value for each sub-region, that is, it is used to perform a max-pooling operation on the above input image. The max-pooling operation is already relatively mature in the related art, so it will not be elaborated here. The padding subunit 240 is used to pad the feature map obtained after max-pooling, that is, to fill with zeros, for adjusting the size of the feature map after max-pooling.

[0089] In the present exemplary embodiment, the second convolutional unit and the third convolutional unit of the third convolutional layer 430 may be used to extract features from the reference feature image to obtain a fifth feature image, and the reference image is input into the residual network unit of the third convolutional layer 430 to obtain a sixth feature image with the same format as the fifth feature image; after the above-mentioned max-pooling operation and padding operation, a sixth feature image with the same format as the fifth feature image can be obtained. The padding operation can make the sixth feature image have the same matrix format as the fifth feature image, so the fifth feature image and the sixth feature image can be directly added as matrices to obtain a first target feature map.

[0090] After obtaining the above first target feature image, the second convolutional unit and the third convolutional unit of the fourth convolutional layer 440 may be used to extract features from the first target feature image to obtain a seventh feature image, and then the first target feature image may be input into the residual network unit of the fourth convolutional layer 440 to obtain an eighth feature image with the same format as the seventh feature image; after the above-mentioned max-pooling operation and padding operation, an eighth feature image with the same format as the seventh feature image can be obtained. The padding operation can make the eighth feature image have the same matrix format as the seventh feature image, so the seventh feature image and the eighth feature image can be directly added as matrices to obtain a second target feature map.

[0091] In this exemplary embodiment, the number of the first target feature images is multiple, that is, the above-mentioned third convolutional layer performs dimensionality reduction processing on the above-mentioned reference image to obtain a relatively large number of first target feature images. Similarly, the fourth convolutional layer also performs dimensionality reduction processing on the multiple first target feature images to obtain a relatively large number of second target feature images, thereby obtaining a relatively large number of reference preselected images and improving the accuracy of target detection.

[0092] In step S130, obtain the reference preselected images for each point on the first target feature image and the second target feature image.

[0093] In an exemplary embodiment of the present disclosure, it is possible to obtain the reference preselected images for each point in the first target feature image and the second target feature image. Specifically, with each point in the first target image and the second target feature image as the center, at least one reference preselected image is generated based on each point.

[0094] In this exemplary embodiment, the number of the first target images and the second target images can both be multiple, and the sizes of the first target feature images and the second target feature images are different. It is possible to obtain corresponding reference preselected images on feature images of multiple sizes, which can improve the accuracy of target detection.

[0095] In this exemplary embodiment, the number of reference preselected images for each point can be multiple, and the above-mentioned reference preselected images can be used to form the reference preselected image combination for this point.

[0096] In this exemplary embodiment, referring to Figure 4 As shown, both the above-mentioned first target feature image and the second target feature image can be input into the preselected image acquisition layer 450. Since for each reference preselected image in each of the k reference preselected image combinations at each position of each feature image in the first target image and the second target feature image, c classes and the scores of each class need to be calculated, and 4 offset values (offsets) of the reference preselected image relative to its default reference preselected image also need to be calculated. Therefore, on each feature image in the first target image and the second target feature image, (c + 4) × k filters are required. Thus, if the size of one of the feature images is m × n, then (c + 4) × k × m × n reference preselected images are generated based on this feature image. Among them, the size of the above-mentioned offset value can be 0.5, or it can be customized according to requirements, and no specific limitation is made in this exemplary embodiment.

[0097] In step S140, determine the target preselected image among the multiple reference preselected images, and complete the target detection using the target detection algorithm.

[0098] In an exemplary embodiment of the present disclosure, category regression can be performed separately on the features of all reference preselected images, and the category scores of each category can be calculated.

[0099] In this exemplary embodiment, based on object detection, the detection targets of the object detection method of this solution can be human heads, human faces, etc. Hereinafter, taking the detection target as a human head as an example for illustration, the categories of the reference preselected images of this solution can include background and human heads. Therefore, in this exemplary embodiment, the scores of the above-mentioned human head category and background category can be calculated, and then the reference preselected images with the scores of the human head category less than a preset value can be excluded, and the reference preselected images with the scores of the human head category greater than or equal to the above-mentioned preset value can be selected as the target preselected images. The above-mentioned preset value can be set by the user, for example, 0.01, 0.02, etc., and no specific limitation is made in this exemplary embodiment.

[0100] In this exemplary embodiment, the obtained multiple target preselected images can be input into the non-maximum suppression operation layer 460, and NMS (Non Maximum Suppression) operation is performed according to the scores of the human head category in the above-mentioned target preselected images and the overlapping parts of each target preselected image to obtain the final preselected images. The detection of human heads can be completed by using the final preselected images. The maximum threshold in the NMS operation can be set according to requirements. For example, 0.45, 0.5, etc., and no specific limitation is made in this exemplary embodiment.

[0101] In summary, the present application can complete the detection of multiple human heads through the above four convolutional layers without adding convolutional layers, which can reduce the computational burden, increase the computational efficiency, and improve the real-time performance of object detection.

[0102] The following introduces the device embodiments of the present disclosure, which can be used to execute the above object detection method of the present disclosure. In addition, in an exemplary embodiment of the present disclosure, an object detection device is also provided. Referring to Figure 5 As shown, the object detection device 500 includes: an image acquisition module 510, a feature extraction module 520, an image selection module 530, and an image determination module 540.

[0103] Among them, the obtaining module 510 can be used to obtain an input image, and perform feature extraction on the input image by using the first convolutional layer 410 and the second convolutional layer 420 to obtain a reference feature image; the matching module 520 can be used to perform feature extraction on the reference feature image by using the third convolutional layer 430 to obtain a first target feature map, and perform feature extraction on the first feature map by using the fourth convolutional layer 440 to obtain a second target feature image; the calculation module 530 can be used to obtain a reference preselected image for each point on the first target feature image and the second target feature image; the determination module 540 can be used to determine a target preselected image from multiple reference preselected images, and complete target detection by using a target detection algorithm.

[0104] Since each functional module of the target detection device in the exemplary embodiment of the present disclosure corresponds to the steps in the exemplary embodiment of the above target detection method, for details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the above target detection method of the present disclosure.

[0105] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0106] In addition, in the exemplary embodiment of the present disclosure, an electronic device capable of implementing the above target detection is also provided.

[0107] Those skilled in the art to which the present disclosure pertains can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0108] Next, refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present disclosure. Figure 6 The shown electronic device 600 is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0109] As Figure 6As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one of the above-mentioned processing units 610, at least one of the above-mentioned storage units 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), and a display unit 640.

[0110] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification above. For example, the processing unit 610 may execute steps S110 shown in Figure 1 as follows: Obtain an input image, and perform feature extraction on the input image using the first convolutional layer 410 and the second convolutional layer 420 to obtain a reference feature image; S120: Perform feature extraction on the reference feature image using the third convolutional layer 430 to obtain a first target feature map, and perform feature extraction on the first target feature image using the fourth convolutional layer 440 to obtain a second target feature image; S130: Obtain a reference preselected image for each point on the first target feature image and the second target feature image; S140: Determine a target preselected image among multiple reference preselected images, and complete target detection using a target detection algorithm; where the first convolutional layer 410 and the second convolutional layer 420 are both single-module residual convolutional layers, and the third convolutional layer 430 and the fourth convolutional layer 440 are double-module residual convolutional layers.

[0111] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 621 and / or a cache storage unit 622, and may further include a read-only storage unit (ROM) 623.

[0112] The storage unit 620 may further include a program / utilities 624 having a set (at least one) of program modules 625. Such program modules 625 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0113] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0114] The electronic device 600 can also communicate with one or more external devices 670 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 660. As shown in the figure, the network adapter 660 communicates with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0115] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0116] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0117] Referring to Figure 7 , a program product 700 for implementing the above method according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.

[0118] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0120] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0121] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0122] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.

[0123] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0124] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A target detection method, It is characterized in that include: Acquire an input image, and use a first convolutional layer and a second convolutional layer to perform feature extraction on the input image to obtain a reference feature image; Using a third convolutional layer to extract features from the reference feature image to obtain a first target feature image, and using a fourth convolutional layer to extract features from the first target feature image to obtain a second target feature image; Acquire a reference pre-selected image of each point on the first target feature image and the second target feature image; Determining a target pre-selected image from the plurality of reference pre-selected images, and performing target detection using a target detection algorithm; The first convolution layer and the second convolution layer are both single-module residual convolution layers, and the third convolution layer and the fourth convolution layer are dual-module residual convolution layers; the single-module residual convolution layer includes a first convolution unit and a residual network unit designed in parallel; the dual-module residual convolution layer includes a second convolution unit and a third convolution unit designed in serial, and a residual network unit designed in parallel with the second convolution unit and the third convolution unit.

2. The method according to claim 1, It is characterized in that Using a first convolutional layer and a second convolutional layer to extract features from the input image to obtain a reference feature image, comprising: Using the first convolutional layer to extract features from the input image to obtain an initial feature image; The second convolutional layer is used to extract features from the initial feature image to obtain a reference feature image.

3. The method according to claim 2, It is characterized in that The first convolution unit includes a depth-separable convolution kernel and a first convolution kernel designed in series; Residual network units, including serially designed max-pooling subunits and padding subunits; Among them, the first convolution kernel is a standard convolution kernel.

4. The method according to claim 3, It is characterized in that Using the first convolutional layer to extract features from the input image to obtain an initial feature map; Inputting the input image into the first convolution unit of the first convolution layer to extract features to obtain a first feature image; Inputting the input image into the residual network unit of the first convolutional layer to obtain a second feature image having the same format as the first feature image; The first characteristic image and the second characteristic image are added to obtain the initial characteristic image.

5. The method according to claim 4, It is characterized in that Using the second convolutional layer to extract features from the initial feature image to obtain a reference feature image includes: Inputting the initial feature image into the first convolution unit of the second convolution layer to perform feature extraction to obtain a third feature image; Inputting the initial feature map into the residual network unit of the second convolutional layer to obtain a fourth feature image having the same format as the third feature image; The third characteristic image and the fourth characteristic image are added to obtain the reference characteristic image.

6. The method according to claim 1, It is characterized in that The second convolution unit includes a depth-separable convolution kernel and a second convolution kernel designed in series; A third convolution unit, connected in series with the second convolution sub-unit, wherein the third convolution unit includes a depth-separable convolution kernel and a third convolution kernel designed in series; Residual network unit, including a serially designed max pooling subunit and a padding subunit.

7. The method according to claim 6, It is characterized in that Using a third convolutional layer to extract features from the reference feature image to obtain a first target feature map, including: The fifth feature image is obtained by extracting features from the reference feature image using the second convolution unit and the third convolution unit of the third convolution layer. Inputting the reference feature image into the residual network unit of the third convolutional layer to obtain a sixth feature image having the same format as the fifth feature image; The fifth feature image and the sixth feature image are added together to obtain the first target feature image.

8. The method according to claim 7, It is characterized in that The fourth convolutional layer is used to extract features from the first target feature image to obtain a second target feature image, including: The second convolution unit and the third convolution unit of the fourth convolution layer are used to extract features of the first target feature image to obtain a seventh feature image. Inputting the first target feature image into the residual network unit of the fourth convolutional layer to obtain an eighth feature image having the same format as the seventh feature image; The seventh feature image and the eighth feature image are added to obtain the second target feature image.

9. The method according to claim 1, It is characterized in that Acquiring a reference pre-selected image of each point on the first target feature image and the second target feature image, comprising: Taking each point of the first target feature image and the second target feature image as a center, at least one reference pre-selected image is generated based on each point.

10. The method according to claim 1, It is characterized in that Determining a target pre-selected image from the plurality of reference pre-selected images, and performing target detection using a target detection algorithm, comprises: Perform category regression on the features of all reference pre-selected images and calculate the category scores of each category; The target pre-selected images can be selected based on the category scores; Perform non-maximum suppression operation on the target pre-selected graphics to complete target detection.

11. A target detection device, It is characterized in that include: An image acquisition module, used to acquire an input image, and use a first convolutional layer and a second convolutional layer to perform feature extraction on the input image to obtain a reference feature image; A feature extraction module, configured to extract features from the reference feature image using a third convolutional layer to obtain a first target feature image, and to extract features from the first feature image using a fourth convolutional layer to obtain a second target feature image; An image selection module, used to obtain a reference pre-selected image of each point on the first target feature image and the second target feature image; An image determination module, used to determine a target pre-selected image from a plurality of the reference pre-selected images, and perform target detection using a target detection algorithm; Among them, the first convolutional layer and the second convolutional layer are both single-module residual convolutional layers, and the third convolutional layer and the fourth convolutional layer are double-module residual convolutional layers; the single-module residual convolutional layer includes a first convolutional unit and a residual network unit designed in parallel; the double-module residual convolutional layer includes a second convolutional unit and a third convolutional unit designed in series, and a residual network unit designed in parallel with the second convolutional unit and the third convolutional unit.

12. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the object detection method according to any one of claims 1 to 10.

13. An electronic device, characterized in that, comprising: a processor; and a memory for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the object detection method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Target detection method for aerial image analysis of unmanned aerial vehicle

    CN114998757A

  • Target detection method and apparatus, and storage medium and electronic device

    WO2022052785A1