Object counting method, object target detection model training method and device
Through the item object detection model, the problem of inaccurate counting of obscured items in image recognition is solved, and the accuracy of inventory inventory is improved.
Patent Information
- Application Number
- CN202310715903.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-06-15
AI Technical Summary
In image recognition, items are blocked due to stacking of columns, and the prior art cannot accurately count, resulting in inaccurate inventory counting.
Visible items and obscured items are identified through the item target detection model, and the feature extraction network and classification regression network are used, combined with the obscured item detection subnet, identify and count occluded items, and use feature fusion and attention mechanisms to improve detection accuracy.
Accurate counting of obscured items is achieved and the accuracy of inventory inventory is improved.
Smart Images

Figure CN116883722B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, image processing, deep learning and other technical fields, and can be applied to scenarios such as smart cities. Background Art
[0002] In scenarios where inventory or shelf counting of items (e.g., goods, stacks, cartons, merchandise, and any other items) is performed through image recognition, object detection technology is typically used to detect the number of items contained in the image in order to count and count the items. However, items are typically stacked in multiple columns, and items in the front column may obscure items in the back column. Therefore, the obscured items cannot be counted, resulting in inaccurate counts. Summary of the Invention
[0003] The present disclosure provides an object counting method, an object target detection model training method, and an apparatus for solving at least one of the above-mentioned technical problems.
[0004] According to one aspect of the present disclosure, a method for counting items is provided, the method comprising:
[0005] Inputting the target object image to be processed into an object detection model to perform object detection, thereby obtaining prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects;
[0006] Obtaining a count of visible items under each item category according to the predicted boxes of the multiple visible items and item category information of the multiple visible items;
[0007] Determining an occluded object count for each object category based on the predicted boxes of the multiple visible objects, the occlusion category information of the multiple visible objects, and the object category information of the multiple visible objects;
[0008] Obtaining an actual item count for each item category based on the visible item count for each item category and the obscured item count for each item category;
[0009] The object target detection model is pre-trained based on at least object images, visible object labels, and occluded object labels.
[0010] According to another aspect of the present disclosure, a method for training an object detection model is provided, the method comprising:
[0011] Obtain object images, visible object labels, and occluded object labels;
[0012] Training the object detection model based on at least the object image, the visible object label, and the occluded object label;
[0013] The object target detection model is used to perform object target detection based on the input object image, and obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
[0014] According to one aspect of the present disclosure, there is provided an object counting device, the device comprising:
[0015] An object detection module is configured to perform object detection by inputting an image of a target object to be processed into an object detection model, thereby obtaining prediction boxes of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects;
[0016] a visible object counting module, configured to obtain a visible object count for each object category based on the prediction boxes of the multiple visible objects and the object category information of the multiple visible objects;
[0017] an occluded object counting module, configured to determine a count of occluded objects under each object category based on the predicted frames of the multiple visible objects, the occlusion category information of the multiple visible objects, and the object category information of the multiple visible objects;
[0018] an actual item counting module, configured to obtain an actual item count for each item category based on the visible item count for each item category and the obscured item count for each item category;
[0019] The object target detection model is pre-trained based on at least object images, visible object labels, and occluded object labels.
[0020] According to one aspect of the present disclosure, a training device for an object detection model is provided, the device comprising:
[0021] An acquisition module, used to acquire object images, visible object labels, and occluded object labels;
[0022] a training module, configured to train the object detection model based on at least the object image, the visible object label, and the occluded object label;
[0023] The object target detection model is used to perform object target detection based on the input object image, and obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
[0024] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0025] at least one processor; and
[0026] a memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0028] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the above method.
[0029] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.
[0030] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0032] Figure 1 This is a flow chart of a method for counting items provided in the first embodiment of the present disclosure;
[0033] Figure 2 This is a schematic diagram of the structure of an exemplary object detection model;
[0034] Figure 3 is a flow chart of a method for counting items provided in a second embodiment of the present disclosure;
[0035] Figure 4 is an exemplary schematic diagram of item detection;
[0036] Figure 5 3 is a flowchart of a method for training an object detection model provided by the third embodiment of the present disclosure;
[0037] Figure 6 A schematic diagram of the structure of an exemplary object detection model training network;
[0038] Figure 7 is a structural diagram of an object counting device provided in a fourth embodiment of the present disclosure;
[0039] Figure 8 2 is a schematic structural diagram of a training device for an object detection model provided in a fifth embodiment of the present disclosure;
[0040] Figure 9 is a block diagram of an electronic device for implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0041] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0042] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0043] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0044] The terminology used herein is for describing particular embodiments only and is not intended to limit the present disclosure.As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0045] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.
[0046] The method for copying a folder of vehicle security assets according to the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in a memory. Alternatively, the method for counting items provided by the present disclosure can be executed by a server.
[0047] It should be noted that the items described in this disclosure can be any type of items depending on the application scenario. For example, in a smart vending scenario, the items are goods, cargo boxes, etc.; in a warehouse inventory application scenario, the items are stacks, containers, etc., which are not limited here.
[0048] In the first embodiment of the disclosure, see Figure 1 , Figure 1 A flowchart of a method for counting items provided by the first embodiment of the present disclosure is shown. The method includes:
[0049] S101: Inputting the target object image to be processed into an object detection model to perform object detection, thereby obtaining prediction boxes of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
[0050] The occlusion classification information is used to indicate whether a visible item is an occlusion representation item, specifically, to classify each visible item as a regular visible item or an occlusion representation item. Therefore, the occlusion classification information includes: occlusion representation items and regular visible items; wherein, an occlusion representation item indicates: a visible item among multiple visible items that represents an occluded item behind itself. For example, an occlusion representation item is a top item (e.g. Figure 4 A11 in the image), and the top item features are: a visible item placed at the top of the stack, with other visible items detectably above it in the item image. If a visible item with these features is detected in the item image, the occlusion classification information for the visible item is determined to be an occlusion-representing item, which effectively means that other occluded items are placed behind the visible item.
[0051] The object detection model can detect visible objects contained in an image and their classification information, and identify regular visible objects and occluded objects from the visible objects. The object detection model can be any object detection model that can identify visible objects from an image, such as: Regional Convolutional Neural Network (R-CNN), Fast Regional Convolutional Neural Network (FastR-CNN), High-Speed Regional Convolutional Neural Network (FasterR-CNN), Single Deep Neural Network (SDD), Single Stage Detection Neural Network (YOLO), and various object detection networks based on these neural networks (such as PPYOLOE).
[0052] S102: Obtain a count of visible items in each item category based on the predicted boxes of the multiple visible items and the item category information of the multiple visible items.
[0053] After the object detection model detects the prediction boxes and item classification information of the visible items in the object image, it can perform classification counting based on the prediction boxes and item classification information of the visible items to determine the count of visible items under each item category.
[0054] In some examples, specifically, the horizontal and vertical axes can be determined based on the vertex positions of the prediction boxes of each visible item and the overall stacked positions of each item. Then, based on the horizontal and vertical axes, the item classification information of each item can be determined from bottom to top according to the vertex positions of the prediction boxes of each item, starting from the item closest to the origin, and classification counting can be performed in this way.
[0055] S103: Determine the occluded object count for each object category based on the predicted boxes of the multiple visible objects, the occlusion category information of the multiple visible objects, and the object category information of the multiple visible objects.
[0056] Through the prediction box of the visible object and the occlusion classification information, the occlusion representation object and its position can be detected. Therefore, the number of occluded objects can be determined based on the position of the occlusion representation object. For example, if the occlusion representation object is determined to be the third object from the bottom up based on the position, it can be known that there are three objects in the back row of the occlusion representation object that are occluded, and the object count is +3. Usually in warehouses or on shelves, multiple items of the same category will be stacked together. Therefore, the item category of the occluded item can be determined, and the number of occluded items can be added to the corresponding item category to determine the occluded item count under each item category.
[0057] S104 : Obtain an actual count of items in each item category according to the count of visible items in each item category and the count of obscured items in each item category.
[0058] The actual item count for each item class is determined by adding the visible item count and the occluded item count for each item class.
[0059] The object detection model is pre-trained based on at least object images, visible object labels, and occluded object labels. Object images and visible object labels (including visible object classification labels and visible object position labels) can be used to train the object detection model to detect visible objects; object images and occluded object labels can be used to train the object detection model to detect occluded objects.
[0060] Using the above method, the object target detection model identifies the visible objects and object categories included in the object image, so that the number of visible objects can be classified and counted to obtain the count of visible objects under each object category; and a new branch for target detection of occluded representation objects is added, so that the object target detection model can identify visible objects that are occluded representation objects among the visible objects, thereby determining the count of occluded objects under each object category based on the occluded representation objects; and then the count of visible objects and the count of occluded objects under each object category are added together to obtain the actual count of objects under each object category. Since the count of occluded objects can be calculated, the accuracy of object counting is effectively improved.
[0061] In disclosing the second embodiment, see Figure 2 , Figure 2 The following is a schematic diagram showing the structure of an exemplary object detection model of the present disclosure. The object detection model includes a connected feature extraction network 01 and a classification regression network 02. The specific structure of the object network model can be implemented in many ways. Figure 2 The basic model structure of the item target network model is PPYOLOE model as an example for explanation. Figure 2 PPYOLOE is a single-stage target detection model, including a backbone network, a neck network, and a head network. The feature extraction network 01 includes a feature extraction subnetwork 011 and a feature fusion subnetwork 012; the classification and regression network 02 includes an occluded object detection subnetwork 021, a regression subnetwork 022, and an object classification subnetwork 023; wherein, the feature extraction subnetwork 011 is the backbone network, the feature fusion subnetwork 012 is the neck network, and the classification and regression network 02 is the head network. Figure 2 The “+” in represents element-wise feature fusion (element-wise add). For ease of explanation, the second embodiment of the present disclosure is described based on the above model.
[0062] See also Figure 3 , Figure 3 A flowchart of a method for counting items provided by the second embodiment of the present disclosure is shown. The method includes:
[0063] S201 : Obtaining the object image features of the target object image to be processed by inputting the image into a pre-trained feature extraction network.
[0064] In some examples, the feature extraction network includes a feature extraction subnetwork and a feature fusion subnetwork. S201 includes the following steps:
[0065] S2011. Input the target object image to be processed into a pre-trained feature extraction subnetwork to extract image features and obtain initial object image features.
[0066] Specifically, the object image (target object image) is input into the feature extraction subnetwork for feature extraction to obtain initial object image features at multiple scales.
[0067] Among them, the feature extraction subnetwork (also known as the backbone network) can be a variety of neural networks that can realize feature extraction, such as: residual network (ResNet), inception network (Inception), visual geometry network (VGG), cross-stage local residual network (CSPRepResNet), etc., which are not limited here.
[0068] In some examples, based on Figure 2 In the present disclosure, the feature extraction subnetwork of the object detection model can be CSPRepResNet, a feature extraction network that adds a cross-stage partial network (CSPNet) to the residual network ResNet50. CSPRepResNet uses the idea of structural reparameterization to extract image features using a fusion structure of multiple feature maps and multiple receptive fields with multiple branches and jump connections during model training. This can improve the detection accuracy of model training. Moreover, during the detection process, CSPRepResNet can degenerate into a single-channel structure with the same effect, which can save video memory and significantly improve detection efficiency.
[0069] S2012. Input the initial object image features into a pre-trained feature fusion subnetwork for feature fusion to obtain object image features.
[0070] Specifically, the initial item image features at multiple scales can be input into the feature fusion sub-network for feature fusion to obtain the item image features at multiple scales output by the feature fusion sub-network. Figure 2 In the
[15] , the three initial object image features (C3, C4, C5) of different scales output by the feature extraction sub-network 011 are input into the feature fusion sub-network 012 for feature fusion to obtain three object image features (P3, P4, P5) of different scales.
[0071] In some examples, the feature fusion sub-network can be a path aggregation network (PAN), a feature pyramid network (FPN), a cross-stage partial path aggregation network (CrossStagePartialPathAggregatedNetwork, CSPPAN), etc., which is not limited here.
[0072] For some examples, see Figure 2 PAN is the feature fusion sub-network 012 of the neck network. Since the PAN network is a bottom-up upsampling network, it can transmit strong positioning information, which is beneficial for detection frame positioning. Therefore, this network structure can better fuse the initial object image features output by the feature extraction sub-network 011, thereby improving network performance.
[0073] S202: Inputting the object image features into a pre-trained classification regression network to obtain prediction boxes of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
[0074] Among them, the classification regression network is used to predict the location and classification of items based on the image features of items.
[0075] In some examples, the classification and regression network includes an object classification subnetwork, a regression subnetwork, and an occluded object detection subnetwork. S202 includes the following steps:
[0076] S2021. By inputting the object image features into the pre-trained regression subnetwork, the prediction boxes of multiple visible objects are obtained.
[0077] The regression subnetwork is used to perform target regression prediction on multiple object image features to obtain the position information of the prediction box to which the feature points of multiple visible objects in the object image features belong (including the coordinate information of the four vertices of the prediction box).
[0078] For some examples, see Figure 2 , Figure 2 The regression subnetwork 022 includes an attention subnetwork for paying attention to multiple channels of object image features (assigning weights to multiple channels) to highlight important channels; and then convolving and integrating the processed object image features to obtain the position information of the prediction boxes to which the feature points of multiple visible objects belong.
[0079] Specifically, the attention mechanism sub-network may be an effective squeezing and concentration sub-network (Effective Squeeze and Extraction Block, ESEBlock), which includes a fully connected layer.
[0080] S2022. Obtain item classification information of multiple visible items by inputting item image features into a pre-trained item classification subnetwork.
[0081] Multiple object image features are input into the object classification subnetwork to perform object classification prediction to obtain a first probability distribution of feature points of the above-mentioned multiple visible objects, wherein the first probability distribution is used to indicate the probability that the corresponding feature points belong to multiple object categories.
[0082] In some examples, the item classification subnetwork includes an attention mechanism subnetwork (which can also be ESEBlock) for paying attention to multiple channels of item image features (assigning weights to multiple channels) to highlight important channels; and then convolving the processed item image features to obtain a first probability distribution of feature points of multiple visible items.
[0083] In the present disclosure, for each prediction frame, the first probability distribution of the prediction frame can be determined based on the first probability distribution of each feature point in the prediction frame, and then the item category to which the target (item) in the prediction frame belongs can be determined based on the first probability distribution of the prediction frame. For example, the predicted item category corresponding to the maximum probability in the first probability distribution can be used as the item category to which the item in the prediction frame belongs.
[0084] In this way, it is possible to simultaneously determine the position information and the first probability distribution of each feature point in the prediction box based on the image features of objects at multiple different scales, which can improve the accuracy and reliability of the determination results.
[0085] S2023. Obtain occlusion classification information of multiple visible objects by inputting object image features into a pre-trained occluded object detection subnetwork.
[0086] The occluded object detection subnetwork is used to classify multiple visible objects into: occluded representation objects and regular visible objects based on occlusion classification information, wherein the occlusion classification information includes whether the visible object is an occluded representation object or a regular visible object.
[0087] In some examples, the occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork. S2023 includes the following substeps:
[0088] Sub-step 1: Input the object image features into the pre-trained first attention sub-network for attention mechanism processing to obtain the processed object image features.
[0089] Sub-step 2: Input the processed object image features into the pre-trained binary classification sub-network for binary classification to obtain occlusion classification information of multiple visible objects.
[0090] Specifically, the first attention subnetwork can be an ESEBlock, which is used to apply attention to multiple channels of object image features (by assigning weights to the multiple channels) to highlight important channels. The processed object image features are then connected through residuals to obtain classification features. This is then passed through a binary classification subnetwork to perform a binary classification task. The two types of classification (binary classification) include regular visible object classification and occlusion representation object classification. Based on the binary classification task, a second probability distribution of the feature points of the visible objects can be obtained. The second probability distribution indicates the probability that the corresponding feature point belongs to the two categories.
[0091] For each prediction box, the second probability distribution of the prediction box can be determined based on the second probability distribution of each feature point in the prediction box, and then the binary classification structure to which the target (item) in the prediction box belongs can be determined based on the second probability distribution of the prediction box. For example, the maximum probability of the second probability distribution of the prediction box corresponds to the occluded representation item classification, which means that the visible item corresponding to the prediction box is not occluded by the representation item classification, and its occlusion classification information is the occluded representation item.
[0092] It should be noted that there are many ways to implement the structure of the obstructed object detection sub-network, which are not limited here.
[0093] It should be noted that in the above example, the feature extraction network and the classification regression network constitute the object target detection model. S201 and S202 are one implementation method of S101. Depending on the structure of the object target detection model, S101 also has other implementation methods, which are not limited here.
[0094] S203 , traversing the prediction frames of the multiple visible objects, and constructing a horizontal sub-axis and a vertical sub-axis perpendicular to the direction of the horizontal sub-axis according to the vertex coordinates of the prediction frames of the multiple visible objects.
[0095] The horizontal sub-axis may be an X-axis, and the vertical sub-axis may be a Y-axis.
[0096] See also Figure 4 , Figure 4An exemplary object detection diagram is shown. After identifying multiple (6+5+3=14) prediction frames of visible items, the four vertex coordinates of the prediction frames can be obtained. Therefore, by traversing the prediction frames of the visible items, the vertex coordinates of the same position of each visible item (for example, the vertex coordinates of the lower left vertex of the prediction frame of each visible item) can be compared to identify visible items with similar Y coordinates (i.e., located in the same row). The X-axis is then constructed based on the vertex coordinates of the bottom edge of the prediction frame of the row of visible items with the smallest Y coordinate (i.e., located in the bottom row). Similarly, the vertex coordinates of the same position of each visible item (for example, the vertex coordinates of the lower left vertex of the prediction frame of each visible item) can be compared to identify visible items with similar X coordinates (i.e., located in the same column). The Y-axis is then constructed based on the vertex coordinates to the left of the prediction frame of the column of visible items with the smallest X coordinate (i.e., located in the leftmost column). Furthermore, multiple columns and rows of visible items can be determined based on the vertex coordinates.
[0097] S204 : Determine visible objects in the same column among the multiple visible objects based on the horizontal sub-axis, the vertical sub-axis, and the vertex coordinates of the prediction boxes of the multiple visible objects, and divide the multiple visible objects into multiple same-column visible object counting groups.
[0098] In some examples, considering that items are usually placed in the same column in the same classification, it is possible to traverse the items and count them by column counting. First, based on the above, according to the determined X-axis, Y-axis and vertex coordinates of the prediction boxes of each visible item (vertex coordinates of the same relative position, for example: the lower left vertex), the visible items can be divided into multiple columns, each column is called a visible item counting group in the same column, and the vertices of the prediction boxes of the visible items in the visible item counting group in the same column have the same X coordinates and different Y coordinates. Of course, in other examples, multiple visible commodities can also be divided into multiple rows, which is not limited here.
[0099] S205. For a same-column visible item count group, traverse the visible items in the same-column visible item count group along the column direction, and perform category counting on the visible items in the same-column visible item count group based on the category information of the visible items, to obtain a single-column visible item count for each category in the same-column visible item count group.
[0100] S206 : Obtain a visible item count for each item category of the multiple visible items according to the visible item counts for each item category in the multiple visible item count groups in the same column.
[0101] After establishing the spatial relationship, traverse the visible items in columns and count the visible items in each column in order from bottom to top (i.e., in the column direction). Figure 4Taking the first visible item count group from left to right in the same column as an example, the visible items in a single column are traversed from bottom to top. Based on the item category information of each visible item, a count is performed under the item category information to which it belongs. The category count information of multiple columns is then aggregated to obtain the visible item count for each item category. For ease of explanation, in this embodiment, the visible items in the same column belong to the same item category. Therefore, the visible item counts of the three visible item count groups in the same column from left to right are: Category A - 6 items, Category B - 5 items, and Category C - 3 items.
[0102] It should be noted that, in the above example, S203 to S206 is one implementation of S102. Depending on the counting method, S102 may have other implementations, which are not limited here.
[0103] S207 , for a visible object counting group in the same column, traverse the visible objects in the visible object counting group in the same column along the column direction, and determine the visible objects in the visible object counting group in the same column that are occlusion representation objects according to the occlusion classification information of each visible object.
[0104] Specifically, Figure 4 Taking the first visible item counting group in the same column from left to right as an example, traverse the visible items from bottom to top and determine whether each visible item is a visible item that occludes the representation item based on the occlusion classification information of the visible item. For example, the visible item in the third row of the first column in the figure is the occlusion representation item A11, and the visible item in the first row of the third column is the occlusion representation item A31.
[0105] S208 : Determine stacking position information of the visible objects that are occluding the representative objects according to the vertex coordinates of the predicted boxes of the visible objects that are occluding the representative objects.
[0106] According to the vertex coordinates of the predicted box of the visible object, the stacking position information of the occluded representation object can be obtained. The stacking position information is used to indicate the number of visible object rows between the occluded representation object and the X-axis. The stacking position information can specifically be the stacking height obtained according to the Y coordinate.
[0107] S209 : Determine, based on the stacking position information and item classification information of the visible items that are occluding representative items, a single-column occluding item count of the item classification corresponding to the visible items that are occluding representative items in the same-column visible item count group.
[0108] According to the stacking position information, the number of visible rows of items between the occluded representation item and the X-axis can be known. For example, in the first visible item counting group in the same column, there are 2 rows of visible items between the occluded representation item A11 and the X-axis, which means that there are 2+1 obscured items. Since the same products are usually placed together, the item classification information of the occluded representation item A11 can be used to determine that the single-column occluded item count of the first visible item counting group in the same column is 2 - A-classified items; similarly, in the third visible item counting group in the same column, there are 0 rows of visible items between the occluded representation item A31 and the X-axis, which means that there is 1 obscured item. Since the same products are usually placed together, the item classification information of the occluded representation item A31 can be used to determine that the single-column occluded item count of the third visible item counting group in the same column is 1 - C-classified item.
[0109] S210 , obtaining an obstructing object count in each item category of the multiple visible objects based on the single-column obstructing object counts of the item category corresponding to the visible objects that obstruct the representative objects in the multiple same-column visible object count groups.
[0110] Sum up the occluded item counts for each column of items under each item category to obtain the occluded item counts for each item category among multiple visible items.
[0111] It should be noted that, in the above example, S207 to S210 are one implementation of S103. Depending on the counting method, S103 may have other implementations, which are not limited here.
[0112] S211 : Obtain an actual item count for each item category according to the visible item count for each item category and the obstructed item count for each item category.
[0113] By adding the count of visible items in each item category to the count of occluded items in each item category, we can get the actual count of items in each item category. Figure 4 , and finally the actual item counts of the three columns of items are: Category A - 6 + 3 = 9 items, Category B - 5 items, Category C - 3 + 1 = 4 items.
[0114] Through the above method, the newly added occluded object detection subnetwork including ESEBlock and binary classification subnetwork performs target detection on occluded representation objects that have occluded objects behind them. It can accurately locate the occluded representation objects using a simple model structure and infer the count of occluded objects based on their placement, ultimately obtaining an accurate object count.
[0115] See also Figure 5 , Figure 5 A flowchart of a method for training an object detection model provided by the third embodiment of the present disclosure is shown. The method includes:
[0116] S301: Acquire an object image, a visible object label, and an occluded object label.
[0117] Visible item labels include visible item category labels and visible item position labels. Classification training for item classification and classification training for occluded representation items / regularly visible items require different datasets. Item images and visible item labels are used for classification training, while item images and occluded representation item labels are used for classification training for occluded representation items / regularly visible items.
[0118] S302: Train an object detection model based on at least the object image, the visible object label, and the occlusion representation object label.
[0119] Among them, the object target detection model is used to perform object target detection based on the input object image, and obtain prediction boxes of multiple visible objects, object classification information of multiple visible objects, and occlusion classification information of multiple visible objects, wherein the occlusion classification information is used to indicate whether the visible object is an occlusion representation object.
[0120] The object detection model can detect visible objects contained in the image and the object classification information of the visible objects, and identify regular visible objects and occluded objects from the visible objects. The object detection model can be trained based on various basic object detection models, such as: Regional Convolutional Neural Network (R-CNN), Fast Regional Convolutional Neural Network (FastR-CNN), High-Speed Regional Convolutional Neural Network (FasterR-CNN), Single Deep Neural Network (SDD), Single Stage Detection Neural Network (YOLO), and various object detection networks based on the above neural networks (such as PPYOLOE).
[0121] See also Figure 6 , Figure 6 The following figure shows a schematic diagram of the training network structure of an exemplary object detection model. The basic model of this object detection model is the PPYOLOE model. The object detection model includes a connected feature extraction network 01 and a classification and regression network 02. In addition, a new branch is added during the model training phase: the metric network 03. The specific structure of the object network model can be implemented in many ways. Figure 6 The basic model structure of the item target network model is PPYOLOE model as an example for explanation. Figure 6PPYOLOE is a single-stage target detection model, including a backbone network, a neck network, and a head network. The feature extraction network 01 includes a feature extraction subnetwork 011 and a feature fusion subnetwork 012; the classification and regression network 02 includes an occluded object detection subnetwork 021, a regression subnetwork 022, and an object classification subnetwork 023; wherein, the feature extraction subnetwork 011 is the backbone network, the feature fusion subnetwork 012 is the neck network, and the classification and regression network 02 is the head network; the measurement network 03 includes a candidate region subnetwork 031 and a feature recognition loss subnetwork 032. Figure 6 The “+” in represents element-wise feature fusion (element-wise add). For ease of explanation, the third embodiment of the present disclosure is described based on the above model.
[0122] In some examples, the object detection model includes a feature extraction network and a classification regression network. The visible object labels include visible object classification labels and visible object position labels.
[0123] Based on this, S202 includes:
[0124] Step 1: Input the object image into the feature extraction network to obtain the object image features of the object image.
[0125] In some examples, the feature extraction network includes a feature extraction subnetwork and a feature fusion subnetwork.
[0126] Specifically, S202 includes: inputting the object image for training into the feature extraction subnetwork to extract image features to obtain initial object image features; and inputting the initial object image features into the feature fusion subnetwork to perform feature fusion to obtain object image features.
[0127] The feature extraction subnetwork (also known as the backbone network) can be a variety of neural networks capable of feature extraction, such as ResNet, Inception, VGG, CSPRepResNet, etc. The feature fusion subnetwork can be PAN, FPN, CSPPAN, etc.
[0128] Step 2: Input the item image features of the item image into a classification regression network to obtain prediction boxes of multiple visible items, item classification information of the multiple visible items, and occlusion classification information of the multiple visible items.
[0129] Step 3: Determine the classification regression loss value based on the prediction boxes of multiple visible objects, the visible object location labels, the item classification information of multiple visible objects, the visible item classification labels, the occlusion classification information of multiple visible objects, and the occlusion representation item labels.
[0130] The classification regression network is used to predict the location and classification of items based on the image features of items.
[0131] After obtaining the prediction boxes of multiple visible objects from step 2, compare them with the visible object position labels to obtain the regression loss value; after obtaining the item classification information of multiple visible objects, compare them with the visible item classification labels to obtain the item classification loss value; after obtaining the occlusion classification information of multiple visible objects, compare them with the occlusion representation item labels to obtain the occlusion classification loss value; then, based on the three loss values, obtain the total classification regression loss value.
[0132] Step 4: Train the object detection model based on the classification regression loss value.
[0133] The classification regression loss value is back-propagated to the input of the object detection model, and the object detection model is optimized and trained to obtain a trained object detection model.
[0134] In some examples, the classification and regression network includes: a regression subnetwork, an object classification subnetwork, and an occluded object detection subnetwork.
[0135] Based on the above, step 2 includes:
[0136] Sub-step 2A: Input the object image features into the regression sub-network to identify the prediction boxes of visible objects, and obtain the prediction boxes of multiple visible objects.
[0137] The regression subnetwork is used to perform target regression prediction on multiple object image features to obtain the position information of the prediction box to which the feature points of multiple visible objects in the object image features belong (including the coordinate information of the four vertices of the prediction box).
[0138] For some examples, see Figure 6 , Figure 6 The regression subnetwork 022 includes an attention subnetwork for paying attention to multiple channels of object image features (assigning weights to multiple channels) to highlight important channels; and then convolving and integrating the processed object image features to obtain the position information of the prediction boxes to which the feature points of multiple visible objects belong.
[0139] Specifically, the attention mechanism sub-network may be an effective squeezing and concentration sub-network (Effective Squeeze and Extraction Block, ESEBlock), which includes a fully connected layer.
[0140] Sub-step 2B: Input the item image features into the item classification sub-network to classify and identify visible items, and obtain item classification information of multiple visible items.
[0141] Multiple object image features are input into the object classification subnetwork to perform object classification prediction to obtain a first probability distribution of feature points of the above-mentioned multiple visible objects, wherein the first probability distribution is used to indicate the probability that the corresponding feature points belong to multiple object categories.
[0142] In some examples, the item classification subnetwork includes an attention mechanism subnetwork (which can also be ESEBlock) for paying attention to multiple channels of item image features (assigning weights to multiple channels) to highlight important channels; and then convolving the processed item image features to obtain a first probability distribution of feature points of multiple visible items.
[0143] In the present disclosure, for each prediction frame, the first probability distribution of the prediction frame can be determined based on the first probability distribution of each feature point in the prediction frame, and then the item category to which the target (item) in the prediction frame belongs can be determined based on the first probability distribution of the prediction frame. For example, the predicted item category corresponding to the maximum probability in the first probability distribution can be used as the item category to which the item in the prediction frame belongs.
[0144] In this way, it is possible to simultaneously determine the position information and the first probability distribution of each feature point in the prediction box based on the image features of objects at multiple different scales, which can improve the accuracy and reliability of the determination results.
[0145] Sub-step 2C: Input the object image features into the occluded object detection sub-network to identify the occluded objects and obtain occlusion classification information of multiple visible objects.
[0146] The occluded object detection subnetwork is used to classify multiple visible objects into: occluded representation objects and regular visible objects based on occlusion classification information, wherein the occlusion classification information includes whether the visible object is an occluded representation object or a regular visible object.
[0147] In some examples, the occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork. Sub-step 2C includes the following sub-steps:
[0148] Sub-step 2 C1: Input the object image features into the first attention sub-network for attention mechanism processing to obtain the processed object image features.
[0149] Sub-step 2 C2: Input the processed object image features into the binary classification sub-network for binary classification to obtain occlusion classification information of multiple visible objects.
[0150] Specifically, the first attention subnetwork can be an ESEBlock, which is used to apply attention to multiple channels of object image features (by assigning weights to the multiple channels) to highlight important channels. The processed object image features are then connected through residuals to obtain classification features. This is then passed through a binary classification subnetwork to perform a binary classification task. The two types of classification (binary classification) include regular visible object classification and occlusion representation object classification. Based on the binary classification task, a second probability distribution of the feature points of the visible objects can be obtained. The second probability distribution indicates the probability that the corresponding feature point belongs to the two categories.
[0151] For each prediction box, the second probability distribution of the prediction box can be determined based on the second probability distribution of each feature point in the prediction box, and then the binary classification structure to which the target (item) in the prediction box belongs can be determined based on the second probability distribution of the prediction box. For example, the maximum probability of the second probability distribution of the prediction box corresponds to the occluded representation item classification, which means that the visible item corresponding to the prediction box is not occluded by the representation item classification, and its occlusion classification information is the occluded representation item.
[0152] It should be noted that there are many ways to implement the structure of the obstructed object detection sub-network, which are not limited here.
[0153] It should be noted that sub-steps 2 AC can be executed in parallel without any restriction on the order of execution.
[0154] Based on the above, sub-step three includes:
[0155] Sub-step 3A: For the regression sub-network, determine the regression loss value based on the predicted boxes and visible object location labels of multiple visible objects.
[0156] After obtaining the predicted boxes of multiple visible objects from sub-step 2A, they are compared with the visible object position labels to obtain the regression loss value.
[0157] In some examples, the regression subnetwork uses distribution focal loss (DFL) to supervise the learning of the integral module in the regression subnetwork, and combines it with the intersection-over-union loss (GIoULoss) to jointly supervise the learning of the regression task. The regression loss value can be determined by DFL and GIoU.
[0158] Sub-step 3B: For the item classification sub-network, determine an item classification loss value based on the item classification information of multiple visible items and the visible item classification labels.
[0159] After obtaining the item classification information of multiple visible items from sub-step 2B, the item classification loss value is obtained by comparing it with the visible item classification label.
[0160] In some examples, the item classification loss value can adopt the variable focus loss (VFL), which uses the intersection over union-aware classification score (IACS) as the training target, so that the model can learn the joint distribution of classification confidence (classification score) and intersection over union confidence (IoU score), eliminating the gap between different prediction branches.
[0161] Sub-step 3C: For the occluded object detection sub-network, determine the occlusion classification loss value based on the occlusion classification information of multiple visible objects and the occlusion representation object labels.
[0162] After obtaining the occlusion classification information of multiple visible objects from sub-step 2C, compare it with the occlusion representation object label to obtain the occlusion classification loss value.
[0163] In some examples, similar to the object classification loss value, the occlusion classification loss value can adopt VFL. VFL uses IACS as the training target, so that the model can learn the joint distribution of classification confidence and intersection-over-union confidence, eliminating the gap between different prediction branches.
[0164] It should be noted that sub-steps 3 AC can be executed in parallel without any restriction on the order of execution.
[0165] Sub-step 3D: Determine the classification regression loss value based on the regression loss value, the item classification loss value, and the occlusion classification loss value.
[0166] After obtaining the three loss values described above, the final classification and regression loss value is then calculated based on these three loss values. There are many ways to calculate the classification and regression loss value, which are not limited here. The classification and regression loss value is backpropagated to the input of the object detection model to optimize it, resulting in a trained object detection model.
[0167] The above classification and regression network is the head network of the object target detection model. In some embodiments, the head network module uses the Task Alignment Learning (TAL) algorithm to perform model evaluation during the model training process.
[0168] Specifically, after outputting multiple detection prediction boxes during model training, the model uses metrics such as classification confidence and intersection-over-union (IoU) to select the detection prediction box for model evaluation. The model is then evaluated using the selected detection prediction box and the preset detection annotation box. Classification confidence describes the probability that the model believes a detection prediction box contains an object of a certain category. IoU is used to characterize the degree of overlap between two bounding boxes.
[0169] In some embodiments, the intersection-over-union parameter may be the ratio between the intersection and the union of two bounding boxes. For example, the intersection-over-union parameter may be an IoU (Intersection over Union) parameter. For example, the classification confidence and the intersection-over-union parameter are sorted in descending order, and the detection prediction box with the highest number of targets is selected as the detection prediction box for model evaluation, and it is ensured that the points covered by the detection prediction box of the target number are within the detection annotation box. In this way, the problem of misalignment between classification and regression tasks can be solved.
[0170] Continue to see Figure 6 In some examples, the object target detection model further includes: a measurement network 03 connected to the feature extraction network 01.
[0171] The metric network is used to obtain the candidate region features corresponding to the visible item labels based on the input object image features (output by the feature extraction network) and the visible item labels, and perform metric learning based on the candidate region features corresponding to the visible item labels and the preset high-precision candidate region features to obtain the feature recognition loss value.
[0172] In this case, before step 4, the method further includes:
[0173] Update the regression loss value based on the feature recognition loss value.
[0174] Among them, the role of the metric network is to perform metric learning based on the high-precision candidate area features and the candidate area features corresponding to the visible object labels obtained based on the object image features of this application, and obtain the feature recognition loss value based on the comparison results of the two. The feature recognition loss value is then combined with the regression loss value to update the regression loss value, and the regression loss value updated based on the feature recognition loss value is back-propagated to the object target detection model for training. Since the high-precision candidate area features have high recognition accuracy, this method can effectively improve the feature recognition accuracy of the target detection model.
[0175] Among them, the high-precision candidate region features are high-precision candidate region features extracted in advance using an image classification model. Specifically, as long as the recognition accuracy of the high-precision candidate region features is higher than the candidate region features corresponding to the visible object labels output by the candidate region sub-network, it will be sufficient.
[0176] In some examples, the image classification model is, for example, VGG, ResNet, dense convolutional network (DenseNet), etc., which are not limited here.
[0177] Specifically, in some examples, see Figure 6, the metric network 03 includes: a candidate region subnetwork 031 and a feature recognition loss subnetwork 032.
[0178] The candidate region subnetwork is used to obtain candidate region features corresponding to visible item labels based on the input item image features and visible item labels. It should be noted that the visible item labels here specifically refer to visible item location labels. This step essentially inputs the item image features and visible item location labels into the candidate region subnetwork, determines the candidate regions within the item image features based on the visible item location labels, and outputs candidate region features corresponding to the visible item labels.
[0179] In some examples, the candidate region subnetwork is, for example, a candidate region alignment (ROIalign) network. Based on the position information of the predicted box, a candidate region alignment operation (i.e., a ROIAlign operation, which is a pooling operation) can be performed on the object image features to obtain candidate region features corresponding to the visible object labels.
[0180] The feature recognition loss subnetwork is used to perform metric learning based on the candidate region features corresponding to the input visible object labels and the high-precision candidate region features to obtain the feature recognition loss value.
[0181] Metric learning is performed based on the candidate regions corresponding to the visible object labels and the high-precision candidate region features, and both are input into the feature recognition loss sub-network to determine the feature recognition loss value based on the learning results.
[0182] It should be noted that feature recognition loss can be calculated in a variety of ways. In some examples, it is triplet loss or contrastive loss. Contrastive loss is directly derived from the similarity score between the candidate region corresponding to the visible object label and the high-precision candidate region features. Triplet loss is derived by introducing a negative sample and calculating the first similarity score between the candidate region corresponding to the visible object label and the negative sample, and the second similarity score between the high-precision candidate region features and the negative sample. This approach, combined with the introduction of a metric network, effectively improves model accuracy.
[0183] In the fourth embodiment of the disclosure, based on Figure 1 The same principle, Figure 7 An article counting device 70 provided in a fourth embodiment of the present disclosure is shown, and the device includes:
[0184] The object detection module 701 is configured to perform object detection by inputting the target object image to be processed into an object detection model, thereby obtaining prediction boxes of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects, wherein the occlusion classification information indicates whether the visible object is an occlusion representation object;
[0185] A visible object counting module 702 is configured to obtain a visible object count for each object category based on the prediction boxes of the multiple visible objects and the object category information of the multiple visible objects;
[0186] an occluded object counting module 703 for determining the occluded object count for each object category based on the predicted frames of the multiple visible objects, the occlusion category information of the multiple visible objects, and the object category information of the multiple visible objects;
[0187] An actual item counting module 704 is configured to obtain an actual item count for each item category based on a visible item count for each item category and a blocked item count for each item category;
[0188] The object target detection model is pre-trained based on at least object images, visible object labels, and occluded object labels.
[0189] In some examples, the object detection model includes: a feature extraction network and a classification regression network;
[0190] The target detection module 701 specifically includes:
[0191] A feature extraction submodule is used to obtain the object image features of the target object image to be processed by inputting the image of the target object to be processed into a pre-trained feature extraction network;
[0192] The classification and regression submodule is used to obtain prediction boxes of multiple visible objects, object classification information of multiple visible objects, and occlusion classification information of multiple visible objects by inputting object image features into a pre-trained classification and regression network.
[0193] In some examples, the feature extraction network includes a feature extraction subnetwork and a feature fusion subnetwork;
[0194] The feature extraction submodule is specifically used for:
[0195] Input the target object image to be processed into the pre-trained feature extraction subnetwork to extract image features and obtain the initial object image features;
[0196] The initial object image features are input into the pre-trained feature fusion subnetwork for feature fusion to obtain the object image features.
[0197] In some examples, the classification and regression network includes an object classification subnetwork, a regression subnetwork, and an occluded object detection subnetwork;
[0198] The classification and regression submodule is specifically used for:
[0199] By inputting the object image features into the pre-trained regression sub-network, the predicted boxes of multiple visible objects are obtained;
[0200] By inputting the item image features into the pre-trained item classification subnetwork, the item classification information of multiple visible items is obtained;
[0201] By inputting the object image features into the pre-trained occluded object detection subnetwork, the occlusion classification information of multiple visible objects is obtained.
[0202] In some examples, the occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork;
[0203] The classification and regression submodule is specifically used for:
[0204] Input the item image features into the pre-trained first attention sub-network for attention mechanism processing to obtain the processed item image features;
[0205] The processed object image features are input into the pre-trained binary classification sub-network for binary classification to obtain the occlusion classification information of multiple visible objects;
[0206] The occlusion classification information includes: whether the visible object is an occlusion representation object or whether the visible object is a regular visible object.
[0207] In some examples, the visible item counting module 702 is specifically configured to:
[0208] Traverse the prediction boxes of multiple visible objects, and construct a horizontal axis and a vertical axis perpendicular to the horizontal axis according to the vertex coordinates of the prediction boxes of the multiple visible objects;
[0209] Determine visible objects located in the same column among the multiple visible objects based on the horizontal sub-axis, the vertical sub-axis, and the vertex coordinates of the prediction boxes of the multiple visible objects, and divide the multiple visible objects into multiple counting groups of visible objects in the same column;
[0210] For a visible item count group in the same column, traverse the visible items in the visible item count group in the same column direction, and perform classification counting based on the item classification information of the visible items in the visible item count group in the same column to obtain a single-column visible item count for each item classification in the visible item count group in the same column;
[0211] The visible item counts under each item category of the multiple visible items are obtained according to the visible item counts in a single column under each item category in the multiple visible item count groups in the same column.
[0212] In some examples, the obstructed object counting module 703 is specifically configured to:
[0213] For a visible object count group in the same column, traverse the visible objects in the visible object count group in the same column along the column direction, and determine the visible objects in the visible object count group in the same column that are occlusion representation objects according to the occlusion classification information of each visible object;
[0214] Determining stacking position information of the visible object that is the occluded representation object according to vertex coordinates of the predicted box of the visible object that is the occluded representation object;
[0215] Determining, based on the stacking position information and item classification information of the visible items that are the occluded representative items, a single-column occluded item count of the item classification corresponding to the visible items that are the occluded representative items in the same-column visible item count group;
[0216] Obstructing object counts in each item category of the plurality of visible objects are obtained according to the obstructing object counts in a single column of the item category corresponding to the visible objects that obstruct the representative items in the plurality of visible object count groups in the same column.
[0217] In the disclosed fifth embodiment, based on Figure 8 The same principle, Figure 8 A training device 80 for an object detection model provided in a fifth embodiment of the present disclosure is shown, and the device includes:
[0218] An acquisition module 801 is configured to acquire an object image, a visible object label, and an occluded object label;
[0219] A training module 802 is configured to train an object detection model based on at least object images, visible object labels, and occluded object labels;
[0220] Among them, the object target detection model is used to perform object target detection based on the input object image, and obtain prediction boxes of multiple visible objects, object classification information of multiple visible objects, and occlusion classification information of multiple visible objects, wherein the occlusion classification information is used to indicate whether the visible object is an occlusion representation object.
[0221] In some examples, the object target detection model includes: a feature extraction network and a classification regression network; the visible object label includes: a visible object classification label and a visible object position label;
[0222] The training module 802 includes:
[0223] A feature extraction submodule is used to input the object image into the feature extraction network to obtain the object image features of the object image;
[0224] A classification and regression submodule is used to input the item image features of the item image into a classification and regression network to obtain prediction boxes of multiple visible items, item classification information of the multiple visible items, and occlusion classification information of the multiple visible items;
[0225] A first loss submodule is configured to determine a classification regression loss value based on the prediction boxes of the multiple visible objects, the visible object position labels, the item classification information of the multiple visible objects, the visible item classification labels, the occlusion classification information of the multiple visible objects, and the occlusion representation item labels;
[0226] The loss inversion submodule is used to train the object detection model based on the classification regression loss value.
[0227] In some examples, the classification and regression network includes: a regression subnetwork, an object classification subnetwork, and an occluded object detection subnetwork;
[0228] The classification and regression submodule is specifically used for:
[0229] Input the item image features into the regression sub-network to identify the prediction boxes of visible items and obtain the prediction boxes of multiple visible items;
[0230] Input the item image features into the item classification subnetwork to classify and identify visible items, and obtain item classification information of multiple visible items;
[0231] The object image features are input into the occluded object detection subnetwork to identify the occluded objects and obtain the occlusion classification information of multiple visible objects.
[0232] In some examples, the first loss submodule is specifically configured to:
[0233] For the regression subnetwork, the regression loss value is determined based on the predicted boxes and position labels of multiple visible objects;
[0234] For the item classification subnetwork, determine the item classification loss value based on the item classification information and the visible item classification labels of multiple visible items;
[0235] For the occluded object detection subnetwork, the occlusion classification loss value is determined based on the occlusion classification information of multiple visible objects and the occlusion representation object labels;
[0236] The classification regression loss value is determined based on the regression loss value, the item classification loss value, and the occlusion classification loss value.
[0237] In some examples, the occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork;
[0238] The first loss submodule is specifically used for:
[0239] Input the item image features into the first attention sub-network for attention mechanism processing to obtain the processed item image features;
[0240] The processed object image features are input into the binary classification sub-network for binary classification to obtain the occlusion classification information of multiple visible objects;
[0241] The occlusion classification information includes: whether the visible object is an occlusion representation object or whether the visible object is a regular visible object.
[0242] In some examples, the object detection model further includes: a metric network connected to the feature extraction network;
[0243] The metric network is used to obtain candidate region features corresponding to visible item labels based on the input object image features and visible item labels. It also performs metric learning based on the candidate region features corresponding to the visible item labels and the preset high-precision candidate region features to obtain the feature recognition loss value.
[0244] In this case, the device also includes:
[0245] The loss update module is used to update the regression loss value based on the feature recognition loss value.
[0246] In some examples, the metric network includes: a candidate region subnetwork, a feature recognition loss subnetwork;
[0247] The candidate region subnetwork is used to obtain candidate region features corresponding to visible object labels based on the input object image features and visible object labels;
[0248] The feature recognition loss subnetwork is used to perform metric learning based on the candidate region features corresponding to the input visible object labels and the high-precision candidate region features to obtain the feature recognition loss value.
[0249] In some examples, the feature recognition loss value is a triplet loss value or a contrastive loss value.
[0250] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0251] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0252] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0253] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 903 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 902 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0254] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0255] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the object counting method or the object detection model training method. For example, in some embodiments, the object counting method or the object detection model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the object counting method or the object detection model training method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute an object counting method or an object target detection model training method in any other appropriate manner (eg, by means of firmware).
[0256] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0257] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0258] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0259] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0260] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0261] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0262] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0263] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for counting items, wherein: The method comprises: Inputting the target object image to be processed into an object detection model to perform object detection, thereby obtaining prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects; Obtaining a count of visible items under each item category according to the predicted boxes of the multiple visible items and item category information of the multiple visible items; Determining a count of occluded items under each item category based on the predicted frames of the multiple visible items, the occlusion classification information of the multiple visible items, and the item classification information of the multiple visible items; wherein, based on the predicted frames of the multiple visible items and the occlusion classification information of the multiple visible items, occluding representative items and positions of the occluded representative items are detected; based on the positions of the occluded representative items, the number of occluded occluding items is determined; and based on the item classification information of the multiple visible items, the number of occluding items under the corresponding item category is added to determine the count of occluded items under each item category; Obtaining an actual item count for each item category based on the visible item count for each item category and the obscured item count for each item category; The object target detection model is pre-trained based on at least object images, visible object labels, and occluded object labels.
2. The method according to claim 1, wherein The object target detection model includes: a feature extraction network and a classification regression network; The step of inputting the target object image to be processed into an object detection model to perform object detection to obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects includes: Obtaining object image features of the target object image by inputting the image of the target object to be processed into the pre-trained feature extraction network; By inputting the object image features into the pre-trained classification regression network, prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects are obtained.
3. The method according to claim 2, wherein: The feature extraction network includes a feature extraction subnetwork and a feature fusion subnetwork; The step of inputting the target object image to be processed into a pre-trained feature extraction network to obtain the object image features of the target object image to be processed includes: Inputting the target object image to be processed into the pre-trained feature extraction subnetwork to extract image features to obtain initial object image features; The initial object image features are input into the pre-trained feature fusion subnetwork for feature fusion to obtain the object image features.
4. The method according to claim 2 or 3, wherein: The classification and regression network includes an object classification subnetwork, a regression subnetwork, and an occluded object detection subnetwork; The step of inputting the object image features into a pre-trained classification regression network to obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects includes: Obtaining prediction boxes of the multiple visible objects by inputting the object image features into the pre-trained regression subnetwork; Obtaining item classification information of the plurality of visible items by inputting the item image features into the pre-trained item classification subnetwork; By inputting the object image features into the pre-trained occluded object detection subnetwork, occlusion classification information of the multiple visible objects is obtained.
5. The method according to claim 4, wherein The occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork; The step of inputting the object image features into the pre-trained occluded object detection subnetwork to obtain occlusion classification information of the multiple visible objects includes: Inputting the item image features into the pre-trained first attention sub-network for attention mechanism processing to obtain processed item image features; Inputting the processed object image features into a pre-trained binary classification subnetwork for binary classification to obtain occlusion classification information of the multiple visible objects; The occlusion classification information is: the visible object is an occlusion representation object or the visible object is a regular visible object.
6. The method according to any one of claims 1 to 3, wherein: Obtaining a count of visible items under each item category according to the predicted boxes of the multiple visible items and the item category information of the multiple visible items includes: Traversing the prediction frames of the multiple visible items, and constructing a horizontal sub-axis and a vertical sub-axis perpendicular to a direction of the horizontal sub-axis according to vertex coordinates of the prediction frames of the multiple visible items; determining, based on the horizontal sub-axis, the vertical sub-axis, and vertex coordinates of the prediction boxes of the multiple visible items, the visible items located in the same column among the multiple visible items, and dividing the multiple visible items into a plurality of same-column visible item counting groups; For a visible item count group in the same column, traverse the visible items in the visible item count group in the same column along a column direction, and perform category counting on the visible items in the visible item count group based on the item category information, to obtain a single-column visible item count for each item category in the visible item count group in the same column; The visible object count of each item category in the plurality of visible objects is obtained according to the visible object counts in a single column under each item category in the plurality of visible object count groups in the same column.
7. The method according to claim 6, wherein: The determining, based on the predicted frames of the multiple visible objects, the occlusion classification information of the multiple visible objects, and the object classification information of the multiple visible objects, a count of occluded objects under each object classification includes: For one of the visible object counting groups in the same column, traverse the visible objects in the visible object counting group in the same column in a column direction, and determine the visible objects in the visible object counting group in the same column that are the occlusion-characterized objects according to the occlusion classification information of each of the visible objects; Determining stacking position information of the visible object that is the occluded representation object according to vertex coordinates of the predicted box of the visible object that is the occluded representation object; Determining, according to the stacking position information of the visible objects that are the obstructing representative objects and the object classification information, a count of obstructing objects in a single column of the object classification corresponding to the visible objects that are the obstructing representative objects in the same-column visible object count group; Obstructing object counts in each item category of the multiple visible objects are obtained according to the single-column obstructing object counts of the item category corresponding to the visible objects that obstruct the representative items in the multiple same-column visible object count groups.
8. A method for training an object detection model, wherein: The method is used for training the object detection model of claim 1, and the method comprises: Obtain object images, visible object labels, and occluded object labels; Training the object detection model based on at least the object image, the visible object label, and the occluded object label; The object target detection model is used to perform object target detection based on the input object image, and obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
9. The method according to claim 8, wherein The object target detection model includes: a feature extraction network and a classification regression network; the visible object label includes: a visible object classification label and a visible object position label; The training of the object detection model based on at least the object image, the visible object label, and the occlusion representation object label includes: Inputting the object image into the feature extraction network to obtain object image features of the object image; Inputting the item image features of the item image into the classification regression network to obtain prediction boxes of multiple visible items, item classification information of the multiple visible items, and occlusion classification information of the multiple visible items; Determining a classification regression loss value based on the predicted boxes of the multiple visible items, the visible item position labels, the item classification information of the multiple visible items, the visible item classification labels, the occlusion classification information of the multiple visible items, and the occlusion representation item labels; The object detection model is trained based on the classification regression loss value.
10. The method according to claim 9, wherein: The classification and regression network includes: a regression subnetwork, an object classification subnetwork, and an occluded object detection subnetwork; The inputting the item image features of the item image into the classification regression network to obtain prediction boxes of multiple visible items, item classification information of the multiple visible items, and occlusion classification information of the multiple visible items includes: Inputting the object image features into the regression subnetwork to perform prediction box recognition of visible objects, thereby obtaining prediction boxes of the multiple visible objects; Inputting the item image features into the item classification subnetwork to classify and identify visible items, thereby obtaining item classification information of the multiple visible items; The object image features are input into the occluded object detection subnetwork to identify the occluded objects, thereby obtaining occlusion classification information of the multiple visible objects.
11. The method according to claim 10, wherein: The determining of the classification regression loss value based on the predicted boxes of the multiple visible items, the visible item position labels, the item classification information of the multiple visible items, the visible item classification labels, the occlusion classification information of the multiple visible items, and the occlusion representation item labels includes: For the regression subnetwork, determining a regression loss value based on the predicted boxes of the multiple visible objects and the visible object position labels; determining, for the item classification subnetwork, an item classification loss value based on the item classification information of the plurality of visible items and the visible item classification labels; determining, for the occluded object detection subnetwork, an occlusion classification loss value based on the occlusion classification information of the plurality of visible objects and the occlusion representation object labels; The classification regression loss value is determined according to the regression loss value, the item classification loss value, and the occlusion classification loss value.
12. The method according to claim 10 or 11, wherein: The occluded object detection subnetwork includes a first attention subnetwork and a binary classification subnetwork; Inputting the object image features into the occluded object detection subnetwork to identify occluded objects and obtain occlusion classification information of the multiple visible objects includes: Inputting the item image features into the first attention sub-network for attention mechanism processing to obtain processed item image features; Inputting the processed object image features into the binary classification subnetwork for binary classification to obtain occlusion classification information of the plurality of visible objects; The occlusion classification information is: the visible object is an occlusion representation object or the visible object is a regular visible object.
13. The method according to any one of claims 9 to 11, wherein: The object detection model further includes: a measurement network connected to the feature extraction network; The metric network is used to obtain candidate region features corresponding to the visible object labels based on the input object image features and the visible object labels, and perform metric learning based on the candidate region features corresponding to the visible object labels and preset high-precision candidate region features to obtain a feature recognition loss value; Before training the object detection model based on the classification regression loss value, the method further includes: The regression loss value is updated according to the feature recognition loss value.
14. The method according to claim 13, wherein The measurement network includes: a candidate region subnetwork and a feature recognition loss subnetwork; The candidate region subnetwork is used to obtain candidate region features corresponding to the visible object labels based on the input object image features and the visible object labels; The feature recognition loss subnetwork is used to perform metric learning based on the candidate region features corresponding to the input visible object labels and the high-precision candidate region features to obtain the feature recognition loss value.
15. The method according to claim 14, wherein The feature recognition loss value is a triplet loss value or a contrast loss value.
16. An article counting device, comprising: An object detection module is configured to perform object detection by inputting an image of a target object to be processed into an object detection model, thereby obtaining prediction boxes of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects; a visible object counting module, configured to obtain a visible object count for each object category based on the prediction boxes of the multiple visible objects and the object category information of the multiple visible objects; an occluded object counting module, configured to determine a count of occluded objects for each object category based on the predicted frames of the multiple visible objects, the occlusion classification information of the multiple visible objects, and the item classification information of the multiple visible objects; wherein, based on the predicted frames of the multiple visible objects and the occlusion classification information of the multiple visible objects, occluded representative objects and their positions are detected; based on the positions of the occluded representative objects, the number of occluded occluded objects is determined; and based on the item classification information of the multiple visible objects, the number of occluded objects is added to the number of occluded objects for the corresponding item category, thereby determining the count of occluded objects for each item category; an actual item counting module, configured to obtain an actual item count for each item category based on the visible item count for each item category and the obscured item count for each item category; The object target detection model is pre-trained based on at least object images, visible object labels, and occluded object labels.
17. A training device for an object detection model, wherein: The device is used for training the object detection model of claim 1, and the device comprises: An acquisition module, used to acquire object images, visible object labels, and occluded object labels; a training module, configured to train the object detection model based on at least the object image, the visible object label, and the occluded object label; The object target detection model is used to perform object target detection based on the input object image, and obtain prediction frames of multiple visible objects, object classification information of the multiple visible objects, and occlusion classification information of the multiple visible objects.
18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7, or the method of any one of claims 8 to 15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7, or the method according to any one of claims 8-15.
20. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 7 or the method according to any one of claims 8 to 15.
Citation Information
Patent Citations
Container system and goods detection device and method
CN111898417A
Method and Apparatus for counting the number of person
KR1020160103844A