Target detection method and device, equipment and medium

By improving the YOLO11 model, the wavelet convolution submodule and feature fusion technology are introduced, and the accuracy problem of traditional object detection methods in complex industrial scenarios is solved, achieving higher detection accuracy and robustness.

CN120279494AActive Publication Date: 2025-07-08GOERTEK INC

Patent Information

Application Number
CN202510756674.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In industrial production, traditional target detection methods are low in complex and changeable product packing scenarios, especially when facing small targets and occlusion problems, and are difficult to meet the needs.

Method used

Using the improved YOLO11 model, a wavelet convolution submodule is introduced, feature extraction is performed through parallel convolution and variable convolution kernel, and feature decomposition maps are combined with wavelet transformation to enhance the contour feature weights, and feature fusion is performed through the neck network and detection head to generate accurate object detection results.

Benefits of technology

It improves the detection accuracy in product packing scenarios, reduces the false detection and missed detection rates, and improves the target detection capability under occlusion and dense arrangement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279494A_ABST
    Figure CN120279494A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and device, equipment and a medium, and relates to the technical field of industrial production. The target detection method comprises the steps that a to-be-detected target image is acquired, and the target image is an image which is acquired on the basis of a shooting device and contains a to-be-boxed product and / or a box body in a product boxing scene; the target image is input into a pre-generated target detection model, a target detection result output by the target detection model is obtained, and the target detection result comprises position information of the product and / or the box body; wherein the target detection model comprises a backbone network, a neck network and a detection head which are connected in sequence, the backbone network at least comprises N first convolution modules for performing feature extraction based on parallel convolution and a variable convolution kernel, M first convolution modules in the N first convolution modules are convolution modules containing wavelet convolution sub-modules, M is smaller than or equal to N, and N is a positive integer greater than or equal to 1; and the wavelet convolution sub-module is used for improving the weight of the contour feature of the product in the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial production technologies, and more specifically, to an object detection method, device, equipment, and medium. Background Art

[0002] In industrial production, the product assembly and packaging links are crucial, and the detection of the product boxing process is a key process in product assembly and packaging. At present, with the rapid development of computer vision technology, object detection technology has gradually been applied to industrial production to achieve the detection of the product boxing process.

[0003] However, in the industrial production scenario, the actual application scenario of object detection technology has the problem of complex and variable environments. For example, complex scenarios such as too small products resulting in too small objects and occlusion between products. This leads to the problem of low detection accuracy in traditional object detection methods. Summary of the Invention

[0004] An object of this application is to provide a new technical solution for object detection.

[0005] According to a first aspect of this application, an object detection method is provided. The method includes: Obtain a target image to be detected, where the target image is an image including the product to be boxed and / or the box in the product boxing scenario collected by a photographing device; Input the target image into a pre-generated object detection model to obtain an object detection result output by the object detection model, where the object detection result includes the position information of the product and / or the box; Wherein, the object detection model includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to perform feature extraction on the target image to obtain multiple first feature maps of different scales; the neck network is used to perform feature fusion processing on the multiple first feature maps of different scales to obtain multiple second feature maps of different scales; the detection head is used to generate the object detection result corresponding to each of the second feature maps; the backbone network includes at least N first convolution modules for feature extraction based on parallel convolution and variable convolution kernels, and M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, where M is less than or equal to N, and the wavelet convolution sub-module is used to increase the weight of the contour features of the product in the target image.

[0006] Optionally, the backbone network further includes a plurality of second convolution modules for performing convolution operations, normalization operations, and activation operations. At least one of the second convolution modules is configured to perform convolution operations, normalization operations, and activation operations on the target image to obtain a third feature map, and input the third feature map into the wavelet convolution sub-module in the first convolution module to generate the first feature map.

[0007] Optionally, the wavelet convolution sub-module includes at least a first wavelet transform unit and a second wavelet transform unit. The first wavelet transform unit is configured to decompose the third feature map into a low-frequency sub-band feature and a high-frequency sub-band feature, and input the low-frequency sub-band feature into the second wavelet transform unit; the low-frequency sub-band feature is used to characterize the contour feature of the product in the third feature map, and the high-frequency sub-band feature is used to characterize the local detail feature of the product in the third feature map; the second wavelet transform unit is configured to decompose the low-frequency sub-band feature into a first low-frequency sub-band feature and a second low-frequency sub-band feature to improve the feature weight of the contour feature in the generated first feature map.

[0008] Optionally, M is less than N, and among the N first convolution modules, the other N - M first convolution modules except the M first convolution modules do not include the wavelet convolution sub-module.

[0009] Optionally, the neck network includes a feature fusion extraction module and a spatial enhancement attention module. The feature fusion extraction module is configured to perform feature fusion extraction based on the first feature maps of multiple different scales to obtain fourth feature maps of multiple different scales. The spatial enhancement attention module is configured to perform spatial attention weighting and multi-scale feature fusion processing on each of the fourth feature maps to generate the second feature map, and input the second feature map into the detection head.

[0010] Optionally, the pre-generated target detection model is pre-generated based on the following method: Obtain a sample image and a training label corresponding to the sample image. The sample image is an image including the product to be packed and / or the box in the product packing scenario collected by a photographing device. Train a preset first detection model according to the sample image and the training label to generate a trained second detection model. Perform pruning processing on the first detection model to generate the target detection model.

[0011] Optionally, the target image is a plurality of consecutive target images in time; the method further includes: Obtain the target detection result corresponding to each frame of image. Assign product identifiers to the products in the target detection results based on the BOT-SORT algorithm; Determine the packing status of the products according to the product identifiers of the products and the change of the area of the products in the multi-frame target images; Determine the number of products in the package according to the packing status of the products.

[0012] According to a second aspect of the present application, there is provided an object detection device, the device comprising: An acquisition module, configured to acquire a target image to be detected, where the target image is an image including products to be packed and / or boxes in a product packing scenario collected by a photographing device; A detection module, configured to input the target image into a pre-generated target detection model, and obtain target detection results output by the target detection model, where the target detection results include position information of the products and / or the boxes; Wherein, the target detection model includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is configured to perform feature extraction on the target image to obtain a plurality of first feature maps with different scales; the neck network is configured to perform feature fusion processing on the plurality of first feature maps with different scales to obtain a plurality of second feature maps with different scales; the detection head is configured to generate the target detection results corresponding to each of the second feature maps; the backbone network includes at least N first convolution modules for feature extraction based on parallel convolution and variable convolution kernels, and M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, where M is less than or equal to N, and the wavelet convolution sub-module is configured to increase the weight of the contour features of the products in the target image.

[0013] According to a third aspect of the present application, there is provided an electronic device, the electronic device comprising a memory and a processor, the memory is configured to store computer instructions, and the processor is configured to call the computer instructions from the memory to execute the method according to any one of the first aspect.

[0014] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the method according to any one of the first aspect.

[0015] The present application provides an object detection method, which includes: obtaining a target image to be detected, where the target image is an image containing the product to be packed and / or the box in the product packing scenario collected by a photographing device; inputting the target image into a pre-generated object detection model to obtain an object detection result output by the object detection model, where the object detection result includes the position information of the product and / or the box; wherein, the object detection model includes a backbone network, a neck network, and a detection head connected in sequence, the backbone network is used to extract features from the target image to obtain a plurality of first feature maps with different scales; the neck network is used to perform feature fusion processing on the plurality of first feature maps with different scales to obtain a plurality of second feature maps with different scales; the detection head is used to generate an object detection result corresponding to each second feature map; the backbone network includes at least N first convolution modules that perform feature extraction based on parallel convolution and variable convolution kernels, and M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, M is less than or equal to N, and the wavelet convolution sub-module is used to increase the weight of the contour features of the product in the target image. In this way, by introducing the wavelet convolution sub-module into the object detection model, the weight of the contour features of the product in the target image can be increased, and the position of the product can still be accurately detected when there are occlusions and dense arrangements of products in the product packing scenario, improving the detection accuracy.

[0016] Other features and advantages of the present application will become clear from the following detailed description of the exemplary embodiments of the present application with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present application and, together with the description, are used to explain the principles of the present application.

[0018] Figure 1 is a block diagram of the hardware configuration of an electronic device for implementing the object detection method according to an embodiment of the present application Figure 1 ; Figure 2 is a schematic flow chart of an object detection method according to an embodiment of the present application; Figure 3 is a schematic structural diagram of an object detection model according to an embodiment of the present application; Figure 4 is a schematic structural diagram of a wavelet convolution sub-module according to an embodiment of the present application; Figure 5 is a schematic structural diagram of a first convolution module according to an embodiment of the present application; Figure 6 is a schematic structural diagram of a spatial attention module according to an embodiment of the present application; Figure 7It is a schematic flowchart of a method for generating a target detection model according to an embodiment of the present application; Figure 8 It is a schematic structural diagram of an apparatus for implementing target detection according to an embodiment of the present application; Figure 9 It is a block diagram of the hardware configuration of an electronic device for implementing a target detection method according to an embodiment of the present application Figure 2 。 Detailed Embodiments

[0019] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present application.

[0020] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present application or its application or use.

[0021] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the above technologies, methods, and devices should be regarded as part of the specification.

[0022] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0023] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0024] Figure 1 It is a block diagram of the hardware configuration of an electronic device for implementing a target detection method according to an embodiment of the present application Figure 1 。

[0025] The electronic device 1000 can be a terminal or a server. Further, the terminal can be a hardware device (such as an AR device, an MR device, and a VR device), a portable computer, a tablet computer, a handheld computer, etc. The server can be a cloud server, etc.

[0026] The electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and so on. Among them, the processor 1100 may be a central processing unit CPU, a microprocessor MCU, etc. The memory 1200 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-transitory memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 can perform wired or wireless communication, for example. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The input device 1600 may include, for example, a touch screen, a keyboard, etc. The user can input / output voice information through the speaker 1700 and the microphone 1800.

[0027] Although a plurality of devices are shown for the electronic device 1000 in Figure 1 , embodiments of the present application may only relate to some of the devices. For example, the electronic device 1000 only includes the memory 1200 and the processor 1100.

[0028] Applied to the embodiments of the present application, the memory 1200 of the electronic device 1000 can be used to store instructions for controlling the processor 1100 to execute the object detection method provided by the embodiments of the present application.

[0029] In the above description, those skilled in the art can design instructions according to the solutions disclosed in the present application. How the instructions control the processor to operate can refer to the descriptions in the related art and will not be elaborated here.

[0030] The present application provides an object detection method applied to an electronic device as shown in Figure 1 . As shown in Figure 2 , the object detection method provided by the present application includes the following steps S2100 to step S2200.

[0031] Step S2100, obtain a target image to be detected.

[0032] Among them, the target image is an image containing the product to be packed and / or the box in the product packing scenario collected by the shooting device.

[0033] Exemplarily, the box can be used to load the product, and the shooting device can be located above the box to shoot the product packing scenario, such as a packing video of the product being loaded into the box. The target image can be one or more frames of the packing video.

[0034] It should be noted that in the scenario of product packing, there are complex and variable scenarios such as occlusion, dense arrangement of products, and relatively small detection targets. The detection accuracy in related technologies often fails to meet the requirements.

[0035] Step S2200: Input the target image into a pre-generated target detection model to obtain the target detection result output by the target detection model.

[0036] Among them, the target detection result includes the position information of the product and / or the box. Optionally, the target detection result may further include parameters such as the type and area of the product and / or the box.

[0037] The target detection model includes a backbone network, a neck network, and a detection head connected in sequence. Among them, the backbone network is used to extract features from the target image to obtain a plurality of first feature maps of different scales, and input the first feature maps into the neck network; the neck network is used to perform feature fusion processing on the plurality of first feature maps of different scales to obtain a plurality of second feature maps of different scales, and input the second feature maps into the detection head; the detection head is used to generate the target detection result corresponding to each second feature map.

[0038] In an embodiment of the present application, the above backbone network may at least include N first convolution modules that perform feature extraction based on parallel convolution and variable convolution kernels. M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, M is less than or equal to N, and both N and M are positive integers. The wavelet convolution sub-module is used to increase the feature weight of the contour features of the product in the target image. The feature weight can be used to characterize the proportion of each feature in the first feature map. The greater the feature weight of the contour features, the more accurate the detected contour of the product. In such a complex scenario as product packing, the accuracy of the target detection model for product detection can be improved.

[0039] Figure 3 It is a schematic structural diagram of a target detection model provided according to an embodiment of the present application. As Figure 3 shown, the target detection model may be an improved YOLO11 model.

[0040] It should be noted that although the YOLO11 model in related technologies can provide fast detection, when facing the above complex scenarios in the product packing process, there are still problems of insufficient accuracy, and there are cases of missed detection or false detection. In this embodiment, the YOLO11 model is improved by introducing a wavelet transform convolution sub-module, which can improve the accuracy of the target detection model for product detection and effectively reduce false detection and missed detection in the product packing process.

[0041] As Figure 3As shown, the object detection model (improved YOLO11 model) may include a backbone network, a neck network, and a detection head. The following is an introduction to each of them: The backbone network can be used for feature extraction. For example, feature extraction is performed on the target image to obtain multiple first feature maps of different scales. The backbone network may include multiple first convolutional modules and multiple second convolutional modules. The first convolutional module may be the C3K2 module in YOLO11, and the second convolutional module may be the CBS (Convolution-BatchNorm-Silu) module in YOLO11, that is, the second convolutional module is used to sequentially perform a convolution operation, a normalization operation, and an activation operation. As Figure 3 shown, the backbone network in this embodiment may include multiple layers connected in series. For example, it includes an input layer, a CBS module, a CBS module, a C3K2 module, a CBS module, a C3K2 module, a CBS module, a C3K2 module, a CBS module, a C3K2 module, an SPPF (Spatial Pyramid Pooling-Fast) module, and a C2PSA (Global Guided Pathway Attention) module, totaling 12 layers.

[0042] In an embodiment of the present application, the above M may be less than N, and among the N first convolutional modules, the other N - M first convolutional modules except the M first convolutional modules do not include a wavelet convolutional sub-module.

[0043] In some examples, the backbone network includes four C3K2 modules. A wavelet convolutional sub-module can be added to the C3K2 modules (i.e., the first convolutional modules) of the eighth and tenth layers, that is, it is transformed into C3K2_WTConv. Optionally, the first convolutional module including the wavelet convolutional sub-module and the first convolutional module without the wavelet convolutional sub-module can be distinguished in name. For example, the first convolutional module including the wavelet convolutional sub-module can be called the third convolutional module (such as Figure 3 the C3K2_WTConv in Figure 3 ), and the first convolutional module without the wavelet convolutional sub-module can be called the fourth convolutional module (such as

[0044] It should be noted that in the backbone network of the YOLO11 model, the scale feature maps processed by the eighth and tenth layers are mainly responsible for extracting high-level semantic information. The wavelet convolution sub-module uses wavelet transform to decompose the input into multi-scale components, and applies small convolutional kernels to the low-frequency sub-bands, significantly expanding the receptive field while the number of parameters only grows logarithmically, and at the same time enhancing the ability to capture low-frequency features (such as shapes and structures). This design improves the extraction efficiency of complex features. In the product packing tracking and counting system, the large receptive field of WTConv and its sensitivity to low-frequency features enable the model to extract structural information from occluded targets, infer their positions, thereby improving recognition robustness; at the same time, by expanding the receptive field, the model can detect small targets in a wider context and use surrounding information to improve counting accuracy. This improvement significantly enhances the performance of the system in complex scenarios and ensures the accuracy of tracking and counting.

[0045] The shallow layers of the backbone network (such as the C3K2 modules in the fourth and sixth layers) mainly extract low-level features, such as edges and textures, which are relatively simple and can be efficiently processed by traditional convolutions. Therefore, the wavelet convolution sub-module can be omitted to improve the operation efficiency of the model.

[0046] At least one second convolution module (i.e., the CBS module) is used to perform convolution operation, normalization operation and activation operation on the target image to obtain a third feature map, and input the third feature map into the wavelet convolution sub-module in the first convolution module to generate a first feature map. The at least one second convolution module can be a module that inputs the feature map of the C3K2 module containing the wavelet convolution sub-module, such as the CBS modules in the seventh and ninth layers.

[0047] Figure 4 It is a schematic structural diagram of a wavelet convolution sub-module provided according to an embodiment of the present application. As Figure 4 shown, the wavelet convolution sub-module C3K2_WTConv of this embodiment can at least include a first wavelet transform unit and a second wavelet transform unit. The first wavelet transform unit is used to decompose the input third feature map into low-frequency sub-band features and high-frequency sub-band features, and input the low-frequency sub-band features into the second wavelet transform unit; the low-frequency sub-band features are used to characterize the contour features of the products in the third feature map, and the high-frequency sub-band features are used to characterize the local detail features of the products in the third feature map; the second wavelet transform unit is used to further decompose the low-frequency sub-band features into a first low-frequency sub-band feature and a second low-frequency sub-band feature to improve the feature weight of the contour features in the generated first feature map.

[0048] That is to say, the wavelet convolution sub-module decomposes the input feature map into low-frequency sub-band features and high-frequency sub-band features through two-dimensional discrete wavelet transform. The low-frequency sub-band features are used to capture global structural features (such as the contour features of the product), and the high-frequency sub-band features are used to retain local detail features. After applying ordinary convolution to the low-frequency sub-band features, the features are recombined using the inverse wavelet transform and added to the result of the ordinary convolution of the input to generate the output. To enhance the multi-scale ability, multiple layers of wavelet transform can be nested to expand the receptive field. Compared with traditional convolution (such as increasing the convolution kernel or stacking layers), the parameter growth of wavelet convolution is logarithmic rather than square. Wavelet convolution can significantly reduce the number of parameters and the amount of computation, laying the foundation for model lightweighting. At the same time, it significantly expands the receptive field and improves the object detection accuracy.

[0049] In this embodiment, the second wavelet transform unit decomposes the low-frequency sub-band features again to form a hierarchical abstraction structure: by cascaded decomposition, the perception range of the underlying convolution kernel is doubled step by step (for example, the secondary decomposition enables a 3×3 kernel to cover a 12×12 area in the original input). Under the constraint of only linear increase in parameters, the network is forced to focus on the basic geometric structure of the target, thereby achieving more robust global feature association in occlusion scenarios. That is to say, by repeatedly decomposing and strengthening the low-frequency sub-band features (such as the contour features of the product), the contour capture ability for occluded targets can be significantly improved, and the computational efficiency is close to that of standard convolution.

[0050] Figure 5 It is a schematic structural diagram of a first convolution module provided according to an embodiment of the present application. As Figure 5 shown, the first convolution module includes a CBS module, a splitting module, M first sub-modules (the first sub-module can be referred to as a C3K module), a dimension fusion module, and a CBS module connected in series in sequence. Among them: The CBS module can include three parts: convolution, batch normalization, and SiLU activation function, which are used to perform convolution operations on the input data, and then perform batch normalization and activation function processing, so as to output a feature map. The splitting module is used to split the feature map generated by the CBS into multiple parts for parallel processing in subsequent convolution operations, which helps the model learn richer feature representations.

[0051] M (for example, 2 or 3) first sub-modules can be respectively used to extract features and enhance the expression ability of the features. The dimension fusion module is used to fuse the multiple feature maps output by the splitting module to restore to the original dimension. It can be achieved through operations such as feature map splicing, addition, or other fusion techniques. The purpose is to integrate the feature information learned by different paths. The CBS module after dimension fusion is used to perform further convolution, batch normalization, and activation operations on the fused feature map to enhance the expression ability of the model.

[0052] Among them, the first sub-module includes a CBS module, N bottleneck units, a dimension fusion module, and a CBS module connected in sequence. Among them, the bottleneck unit includes a CBS module and a wavelet convolution sub-module connected in sequence. That is to say, Figure 5 The first convolution module shown includes a wavelet convolution sub-module. The bottleneck unit may also include an optional residual connection. Based on the residual connection, the input of the bottleneck unit can be directly superimposed on the output of the wavelet convolution sub-module to obtain the output of the bottleneck unit. The bottleneck unit in the first convolution module including the wavelet convolution sub-module can set this residual connection, that is, it is allowed to directly superimpose the input of the bottleneck unit on the output of the wavelet convolution sub-module. For the bottleneck unit in the first convolution module that does not include the wavelet convolution sub-module, this residual connection may not be set, that is, there is no need to superimpose the input of the bottleneck unit again. It should be noted that it can also be determined by the user whether to set this residual connection for the bottleneck units in each first convolution module according to different requirements, and this embodiment does not limit this.

[0053] In some examples, the backbone network can perform feature extraction on the target image to obtain at least three first feature maps of different scales and input the three first feature maps of different scales into the feature fusion extraction module of the neck network. Exemplarily, the at least three first feature maps can be output by the C3K2 module of the sixth layer, the C3K2_WTConv module of the eighth layer, and the C2PSA module of the twelfth layer respectively.

[0054] The neck network can be used to perform feature fusion processing on multiple first feature maps of different scales to obtain multiple second feature maps of different scales, and input the second feature maps into the detection head.

[0055] In an embodiment of the present application, the neck network includes a feature fusion extraction module and a Spatially Enhanced Attention Module (SEAM). Among them: The feature fusion extraction module is used to perform feature fusion extraction based on multiple first feature maps of different scales to obtain multiple fourth feature maps of different scales. The feature fusion extraction module can adopt the structure of the neck network of YOLO11 in related technologies.

[0056] The space enhanced attention module is used to perform spatial attention weighting and multi-scale feature fusion processing on each fourth feature map respectively, generate a second feature map, and input the second feature map into the detection head. The space enhanced attention module can be a module added to the neck network of YOLO11 in related technologies.

[0057] Such as Figure 3As shown in the figure, the neck network may include an upsampling module, a dimension fusion module, a C3K2 module, an upsampling module, a dimension fusion module, a C3K2 module, a SEAM module, a CBS module, a dimension fusion module, a C3K2 module, a SEA module, a CBS module, a dimension fusion module, a C3K2 module, and a SEA module that are connected in sequence. Among them, the three SEAM modules are spatial enhancement attention modules respectively used to output a second feature map to the detection head. The other modules constitute a feature fusion extraction module, which is used to perform feature fusion extraction on multiple first feature maps of different scales to obtain multiple fourth feature maps of different scales, and input the fourth feature maps into the SEAM module respectively, so that the SEAM module generates a second feature map based on the fourth feature map.

[0058] It should be noted that, as Figure 3 shown, the spatial attention module SEAM can be applied before the neck network outputs to the detection head, helping the model adjust the feature map weights at different scales (such as the target detection layers of large scale, medium scale, and small scale), and optimizing the occlusion processing ability. For example, in the product boxing scenario, SEAM can effectively reduce the missed detection rate caused by partial occlusion and improve the detection robustness.

[0059] Figure 6 is a schematic structural diagram of a spatial attention module SEAM provided according to an embodiment of the present application. As Figure 6 shown, the spatial attention module may include an input layer, a CSMM module (Channel and Spatial Mixing Module), an average pooling layer (AveragePooling), a linear layer (Linear), a ReLU activation function (ReLU), a linear layer (Linear), a Sigmoid activation function (Sigmoid), and a channel expansion (Channel exp) that are connected in sequence, where: The CSMM module (Channel and Spatial Mixing Module) is used to process the input data with different patch sizes (such as Patch = 6, 7, 8) to extract multi-scale features. The CSMM module processes the input data through patches of different sizes to capture feature information at different scales. Among them, a patch can be a small area extracted from the target image. These small areas can be square or rectangular, and they can be used to capture local features in the image. When processing an image, splitting the large image into multiple small patches (patches) allows the model to pay more attention to the detailed parts of the image, thereby improving the efficiency and accuracy of feature extraction. In this embodiment, using patches of different sizes means that the model will extract local regions of different sizes from the input data for feature extraction. The advantage of doing this is that it can capture features at different scales, enabling the model to better understand various details and structures in the image. For example, smaller patches can capture finer details of the product, while larger patches can capture more extensive context information, such as the contour features of the product. The average pooling layer (Average Pooling) is used to perform average pooling on the features processed by the CSMM module to reduce the dimension of the feature map while retaining key information. The linear layer (Linear) is used to perform a linear transformation on the pooled features to adjust the dimension of the features. The ReLU activation function (ReLU) is used to introduce non-linearity and enhance the expressive power of the model. The linear layer (Linear) is used to perform a linear transformation on the features again to further adjust the dimension of the features. The Sigmoid activation function (Sigmoid) is used to map the feature values to the interval (0, 1) for subsequent channel expansion operations. Channel expansion (Channel exp) is used to enhance the expressive power of the features by expanding the number of channels.

[0060] In one embodiment of the present application, the above CSMM module may include at least three CSMMs, such as three CSMMs with Patches of 6, 7, and 8 respectively. Each CSMM includes modules such as patch embedding (Patch Embedding), GELU activation function (GELU), batch normalization (Batch Normalization), depthwise convolution (Depthwise Convolution), GELU activation function (GELU), batch normalization (Batch Normalization), pointwise convolution (Pointwise Convolution), GELU activation function (GELU), and batch normalization (Batch Normalization) connected in sequence. Among them: Patch Embedding is used to split the input data into patches (regions) of different sizes and perform embedding processing to capture local features. The GELU activation function (GELU) is used to introduce the GELU activation function after Patch Embedding to enhance the non-linear characteristics. Batch Normalization is used to perform batch normalization on the features to accelerate the convergence speed of the model and improve stability. Depthwise Convolution is used to perform convolution operations on the features channel by channel to extract spatial features and learn the correlation between the spatial dimension and channels. Pointwise Convolution is used to fuse channel information through 1x1 convolution to learn the correlation between channels.

[0061] In this way, the SEAM module uses its CSMM module to extract multi-scale features from different patches, and uses depthwise separable convolution to extract spatial features channel by channel and learn the correlation between the spatial dimension and channels, thereby reducing the computational complexity. At the same time, SEAM combines residual connections to enhance feature transmission to avoid information loss, 1x1 convolution to fuse channel information to learn the correlation, and a fully connected layer to highlight the feature expression of the unoccluded region through a two-layer network, thereby adaptively enhancing the feature response of the unoccluded region and compensating for the feature loss caused by occlusion.

[0062] In an embodiment of the present application, the detection head in the object detection model can be used to generate object detection results corresponding to each second feature map. For example, each second feature map can be detected to obtain the object detection results. It should be noted that the specific implementation manners of some modules in the backbone network, neck network, and detection head of the object detection model can also adopt the manners in related technologies, and this embodiment does not limit this.

[0063] In an embodiment of the present application, a method for generating an object detection model is also provided, and this method may include the following steps S11 to step S12.

[0064] Step S11, obtain a sample image and the training label corresponding to the sample image.

[0065] Among them, the sample image is an image containing the product to be packed and / or the box in the product packing scenario collected by the shooting device.

[0066] Step S12, train a preset first detection model according to the sample image and the training label to generate a trained second detection model.

[0067] Step S13, perform pruning processing on the first detection model to generate an object detection model.

[0068] Exemplarily, the first detection model can be pruned based on the Layer-wise Adaptive Magnitude Pruning (LAMP) method to generate a target detection model. LAMP can adaptively remove redundant parameters according to the importance of the weights of each layer, successfully reducing the computational complexity while maintaining the high accuracy of the model. This optimization makes the model smaller and faster, very suitable for deployment on edge devices, and capable of meeting the requirements of real-time applications.

[0069] In this way, by using the above method to pre-generate a target detection model and optimizing the model based on pruning technology, redundant parameters are reduced and the inference speed is improved. The target detection model can be applied to Figure 2 the target detection method shown in

[0070] The embodiment of the present application also provides a target detection method, which can be applied to an electronic device as shown in Figure 1 As shown in Figure 7 The target detection method may further include the following steps S7100 to S7500.

[0071] Step S7100, obtain multiple frames of target images.

[0072] Exemplarily, multiple frames of target images including products to be packed and / or boxes in the product packing scenario collected by a shooting device can be obtained.

[0073] The shooting device can be located above the box to shoot the product packing scenario, such as the packing video of the process of the product being loaded into the box, so as to obtain multiple frames of target images in the packing video.

[0074] Step S7200, obtain the target detection result corresponding to each frame of target image.

[0075] The above target images include multiple frames of target images that are continuous in time. The target detection result corresponding to each frame of target image can be obtained based on Figure 2 the method of the embodiment shown in

[0076] Step S7300, assign product identifiers to the products in the target detection results based on a preset tracking algorithm.

[0077] Among them, the preset tracking algorithm can be the BOT-SORT (Bag of Tricks for SORT) algorithm.

[0078] Step S7400: Determine the packing status of the product based on the product identifier of the product and the change in the area of the product in multiple frames of target images.

[0079] Step S7500: Determine the number of products in the package according to the packing status of the product.

[0080] In some examples, the packing status of the product may include one or more of packing start, in the process of packing, or packing completed. Exemplarily, when the product corresponding to the product identifier is first detected, the packing status of the product may be packing start, and then it is in the state of being packed. According to the change in the area of the product in multiple frames of target images, it can be determined whether the packing status of the product changes to packing completed. After a product is packed, the data of the products in the package is incremented by one.

[0081] In an embodiment of the present application, the data of the products in the package can also be output and displayed to the user. For example, the image of the product being packed, the product identifier, and the current number of products in the package can be displayed through a display device.

[0082] By adopting the above method, on the basis of improving the detection accuracy and efficiency based on the target detection model, the BoT-SORT algorithm is combined for target tracking, further improving the stability and accuracy in a multi-target environment. Through this multi-level optimization scheme, false detections and missed detections can be effectively reduced, and the overall performance of product packing tracking and counting can be improved.

[0083] Figure 8 It is a schematic diagram of a target detection device provided by an embodiment of the present application. As Figure 8 shown, the target detection device 800 may include: An acquisition module 810, configured to acquire a target image to be detected, where the target image is an image including the product to be packed and / or the box in the product packing scenario collected by a shooting device; A detection module 820, configured to input the target image into a pre-generated target detection model to obtain a target detection result output by the target detection model, where the target detection result includes the position information of the product and / or the box; Among them, the target detection model includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to extract features from the target image to obtain a plurality of first feature maps with different scales; the neck network is used to perform feature fusion processing on the plurality of first feature maps with different scales to obtain a plurality of second feature maps with different scales; the detection head is used to generate the target detection results corresponding to each of the second feature maps; the backbone network includes at least N first convolution modules that perform feature extraction based on parallel convolution and variable convolution kernels. M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, where M is less than or equal to N, and the wavelet convolution sub-module is used to increase the weight of the contour features of the product in the target image.

[0084] In one embodiment of the present application, the backbone network further includes a plurality of second convolution modules for performing convolution operations, normalization operations, and activation operations. At least one of the second convolution modules is used to perform convolution operations, normalization operations, and activation operations on the target image to obtain a third feature map, and input the third feature map into the wavelet convolution sub-module in the first convolution module to generate the first feature map.

[0085] In one embodiment of the present application, the wavelet convolution sub-module includes at least a first wavelet transform unit and a second wavelet transform unit. The first wavelet transform unit is used to decompose the third feature map into a low-frequency sub-band feature and a high-frequency sub-band feature, and input the low-frequency sub-band feature into the second wavelet transform unit; the low-frequency sub-band feature is used to characterize the contour features of the product in the third feature map, and the high-frequency sub-band feature is used to characterize the local detail features of the product in the third feature map; the second wavelet transform unit is used to decompose the low-frequency sub-band feature into a first low-frequency sub-band feature and a second low-frequency sub-band feature to increase the feature weight of the contour features in the generated first feature map.

[0086] In one embodiment of the present application, M is less than N, and the other N - M first convolution modules among the N first convolution modules do not include the wavelet convolution sub-module.

[0087] In one embodiment of the present application, the neck network includes a feature fusion extraction module and a spatial enhancement attention module. The feature fusion extraction module is used to perform feature fusion extraction based on the plurality of first feature maps with different scales to obtain a plurality of fourth feature maps with different scales. The spatial enhancement attention module is used to perform spatial attention weighting and multi-scale feature fusion processing on each of the fourth feature maps to generate the second feature map, and input the second feature map into the detection head.

[0088] In one embodiment of the present application, the device further includes a generation module, and the generation module is used to pre-generate the target detection model based on the following method: obtain a sample image and a training label corresponding to the sample image, where the sample image is an image including the product to be packed and / or the box in the product packing scenario collected by the shooting device; train a preset first detection model according to the sample image and the training label to generate a trained second detection model; perform pruning processing on the first detection model to generate the target detection model.

[0089] In one embodiment of the present application, the target image is multiple frames of target images that are continuous in time; the device further includes: A processing module, configured to obtain the target detection result corresponding to each frame of image; assign product identifiers to the products in the target detection result based on the BOT-SORT algorithm; determine the packing status of the products according to the product identifiers of the products and the change in the area of the products in the multiple frames of target images; determine the number of products in the package according to the packing status of the products.

[0090] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0091] The present application also provides an electronic device, and the electronic device includes any one of the target detection devices 800 provided in the above device embodiments.

[0092] Or, as Figure 9 shown, the electronic device 1000 includes a memory 1200 and a processor 1100. The memory 1200 is used to store computer instructions, and the processor 1100 is used to call the computer instructions from the memory 1200 to execute any one of the methods provided in the above method embodiments.

[0093] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements any one of the methods in the foregoing embodiments of the present application. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto, and it may also be a temporary storage medium.

[0094] The embodiments of the present application also provide a computer program product, and the computer program product may include a computer program, and when the computer program is executed by a processor, it may implement any one of the methods in the foregoing embodiments of the present application.

[0095] This application may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement any of the methods in the foregoing embodiments of this application.

[0096] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0097] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0098] The computer program instructions for performing the operations of the present application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, which may include object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present application.

[0099] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0100] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0101] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0102] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It should be noted that implementing through hardware, through software, and through a combination of software and hardware are all equivalent.

[0103] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the technical field to understand the embodiments disclosed herein. The scope of the present application is defined by the appended claims.

Claims

1. A target detection method, characterized in that, The method includes: Obtaining a target image to be detected, where the target image is an image containing the product to be packed and / or the box in the product packing scenario collected by a photographing device; Inputting the target image into a pre-generated target detection model to obtain a target detection result output by the target detection model, where the target detection result includes the position information of the product and / or the box; Wherein, the target detection model includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to perform feature extraction on the target image to obtain a plurality of first feature maps with different scales; the neck network is used to perform feature fusion processing on the plurality of first feature maps with different scales to obtain a plurality of second feature maps with different scales; the detection head is used to generate the target detection result corresponding to each of the second feature maps; the backbone network includes at least N first convolution modules for feature extraction based on parallel convolution and variable convolution kernels. M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, where M is less than or equal to N, and the wavelet convolution sub-module is used to increase the weight of the contour feature of the product in the target image.

2. The method according to claim 1, wherein The backbone network further includes a plurality of second convolution modules for performing convolution operations, normalization operations, and activation operations. At least one of the second convolution modules is used to perform convolution operations, normalization operations, and activation operations on the target image to obtain a third feature map, and input the third feature map into the wavelet convolution sub-module in the first convolution module to generate the first feature map.

3. The method according to claim 2, wherein The wavelet convolution sub-module includes at least a first wavelet transform unit and a second wavelet transform unit. The first wavelet transform unit is used to decompose the third feature map into a low-frequency sub-band feature and a high-frequency sub-band feature, and input the low-frequency sub-band feature into the second wavelet transform unit; the low-frequency sub-band feature is used to represent the contour feature of the product in the third feature map, and the high-frequency sub-band feature is used to represent the local detail feature of the product in the third feature map; the second wavelet transform unit is used to decompose the low-frequency sub-band feature into a first low-frequency sub-band feature and a second low-frequency sub-band feature to increase the feature weight of the contour feature in the generated first feature map.

4. The method according to claim 1, characterized in that, M is less than N, and the other N - M first convolution modules among the N first convolution modules do not include the wavelet convolution sub-module.

5. The method according to claim 1, wherein The neck network includes a feature fusion extraction module and a spatial enhancement attention module, The feature fusion extraction module is used to perform feature fusion extraction based on the plurality of first feature maps with different scales to obtain a plurality of fourth feature maps with different scales; The spatial enhancement attention module is used to perform spatial attention weighting and multi-scale feature fusion processing on each of the fourth feature maps to generate the second feature map, and input the second feature map into the detection head.

6. The method according to claim 1, characterized in that, The pre-generated target detection model is pre-generated based on the following method: Obtain a sample image and training labels corresponding to the sample image, where the sample image is an image containing products to be packed and / or boxes in a product packing scenario collected by a photographing device; Train a preset first detection model according to the sample image and the training labels to generate a trained second detection model; Perform pruning processing on the first detection model to generate the target detection model.

7. The method according to any one of claims 1 to 6, characterized in that The target images are multiple consecutive frames of target images in time; the method further includes: Obtain the target detection results corresponding to each frame of target image; Assign product identifiers to the products in the target detection results based on the BOT-SORT algorithm; Determine the packing status of the products according to the product identifiers of the products and the change in the area of the products in the multiple frames of target images; Determine the number of products packed according to the packing status of the products.

8. A target detection device, characterized in that, The device includes: An acquisition module, configured to acquire a target image to be detected, where the target image is an image containing products to be packed and / or boxes in a product packing scenario collected by a photographing device; A detection module, configured to input the target image into a pre-generated target detection model to obtain target detection results output by the target detection model, where the target detection results include position information of the products and / or the boxes; Wherein, the target detection model includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is configured to perform feature extraction on the target image to obtain multiple first feature maps of different scales; the neck network is configured to perform feature fusion processing on the multiple first feature maps of different scales to obtain multiple second feature maps of different scales; the detection head is configured to generate the target detection results corresponding to each of the second feature maps; the backbone network includes at least N first convolution modules that perform feature extraction based on parallel convolution and variable convolution kernels. M of the N first convolution modules are convolution modules including wavelet convolution sub-modules, where M is less than or equal to N, and the wavelet convolution sub-module is configured to increase the weight of the contour features of the products in the target image.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is configured to store computer instructions, and the processor is configured to call the computer instructions from the memory to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Infrared weak and small target detection method and device

    CN118397458A

  • Pedestrian detection method and device, equipment, storage medium and product

    CN119206853A

  • Method for detecting spinning box, detection device, electronic device, storage medium, and program

    JP7669600B1

Cited By

  • Tooth brushing action detection method, device, equipment and medium

    CN121545213A