Cargo quantity identification method, device, electronic device and storage medium
By using a multi-layer convolutional neural network model to extract cargo image features and determine object candidate boxes, the problem of low cargo quantity recognition accuracy in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202111534478.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The accuracy of the cargo quantity identification method in the prior art is not high and depends heavily on threshold setting and contour detection accuracy.
The pre-trained neural network model with multi-layer convolutional layers is used to extract the image feature of the cargo image, determine the object candidate box set, and determine the final cargo quantity through screening and deduplication.
Improves the accuracy of cargo quantity recognition, can identify objects of different locations and sizes, and enhances the robustness and recall of the model.
Smart Images

Figure CN114220011B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition and deep learning technology, and in particular to a method, device, electronic device, storage medium and program product for identifying the quantity of goods. Background Art
[0002] To prevent damage to goods during transportation that could make it difficult to determine responsibility, cargo must be counted during loading and unloading. For regular-shaped goods, such as steel pipes, image recognition technology can be used for this. A photo is taken during loading and another taken during unloading, and both uploaded to the logistics management system. The system then uses image recognition technology to count and compare the number of steel pipes. If the numbers in the two photos don't match, it proves that the goods were lost during transportation.
[0003] In the existing technology, the method for identifying the quantity of goods is usually as follows: first, the steel pipe image is converted into a grayscale image with pixel values between 0 and 255, and then a threshold is set so that values greater than the threshold are 1 and values less than the threshold are 0; then the outer contour of the steel pipe is obtained through an edge detection algorithm; the outer contour value is de-noised through expansion and corrosion, the existing contour display is enhanced, and the circular shape of the steel pipe cross-section is detected based on the contour line; finally, the number of times the circle appears is counted to obtain the number of steel pipes in the image.
[0004] However, the cargo quantity recognition method in the above-mentioned prior art is heavily dependent on factors such as whether the threshold setting is reasonable and the contour detection accuracy, resulting in low recognition accuracy. Summary of the Invention
[0005] The present application provides a method, device, electronic device, storage medium and program product for identifying the quantity of goods, so as to solve the problem of low accuracy in the method for identifying the quantity of goods in the prior art.
[0006] In a first aspect, the present application provides a method for identifying the quantity of goods, the method comprising:
[0007] Extracting image features from the cargo images using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers;
[0008] Determining a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of the object candidate frames;
[0009] screening and removing duplicate object candidate frames in the object candidate frame set according to the object category and the positioning coordinates of the object candidate frames in the image features;
[0010] The number of target object candidate frames obtained after the screening and deduplication is used as the number of goods in the goods image.
[0011] In a second aspect, the present application further provides a device for identifying the quantity of goods, the device comprising:
[0012] An image feature extraction module, configured to extract image features from cargo images using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers;
[0013] an object candidate frame set determination module, configured to determine a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of object candidate frames;
[0014] a screening and deduplication module, configured to screen and dedupe the object candidate frames in the object candidate frame set according to the object categories and the positioning coordinates of the object candidate frames in the image features;
[0015] The cargo quantity determination module is configured to use the number of target object candidate frames obtained after the screening and deduplication as the cargo quantity in the cargo image.
[0016] In a third aspect, the present application further provides an electronic device, comprising:
[0017] one or more processors;
[0018] a storage device for storing one or more programs,
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the cargo quantity identification method as described above.
[0020] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the cargo quantity identification method as described above.
[0021] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the cargo quantity identification method as described above when executed by a processor.
[0022] The technical solution of this application uses a neural network model with multiple layers of convolution to extract image features. Image features output by at least two convolutional layers in the model are selected to determine a set of candidate object frames in the cargo image. Then, through screening and deduplication, the final identified target object candidate frames are determined from the set of candidate object frames, and the number of target object candidate frames is used as the number of cargo in the cargo image. Because image features output by at least two convolutional layers in the model are selected, the model can identify objects of different locations and sizes in the cargo image, thereby discovering multiple targets in the image and accurately identifying the cargo in the image, thereby improving the model's accuracy in cargo quantity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a flow chart of the method for identifying the quantity of goods in Example 1 of the present application;
[0024] Figure 2 This is a flow chart of the method for identifying the quantity of goods in Example 2 of the present application;
[0025] Figure 3 This is a flow chart of the method for identifying the quantity of goods in Example 3 of the present application;
[0026] Figure 4 This is a schematic diagram of the structure of the cargo quantity identification device in the fourth embodiment of the present application;
[0027] Figure 5 It is a structural diagram of the electronic device in Example 5 of the present application. DETAILED DESCRIPTION
[0028] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present application and are not intended to limit the present application. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions of the present application, not all of the structures.
[0029] Example 1
[0030] Figure 1 This is a flowchart of a method for identifying the quantity of goods provided in the first embodiment of the present application. This embodiment can be applied to image recognition of goods images to obtain the quantity of goods in the image, and involves the field of image recognition and deep learning technology. The method can be executed by a device for identifying the quantity of goods, which can be implemented in software and / or hardware, and is preferably configured in an electronic device, such as a computer device or server. Figure 1 As shown, the method specifically includes:
[0031] S101. Perform image feature extraction on a cargo image using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers.
[0032] The quantity recognition model is a neural network model with a multi-layer convolutional structure pre-trained using machine learning methods. The training process for the quantity recognition model can begin by obtaining a large number of images of goods to be identified as training samples. The goods in the images are manually labeled, for example, by using rectangular boxes to mark the cross-sections of the goods. The training samples are then input into the model, which predicts the goods in the training samples. The error is calculated based on the predicted value and the labeled value of the training samples. The gradient of the error with respect to the convolution parameters is then calculated to optimize the model parameters. Through the iteration and optimization of this process, the final convolutional neural network-based quantity recognition model is obtained.
[0033] S102. Determine a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of the object candidate frames.
[0034] The trained quantity recognition model can extract features from the cargo image, and each convolutional layer can output image features, such as object categories and the positioning coordinates of object candidate frames. Based on these image features, the model can identify candidate frames for each object in the cargo image. Among them, the object category can include target and background, the target refers to the object in the foreground of the image, and the background refers to the background part of the image. It can be understood that the cargo to be identified as an object in the image has a category of target. The positioning coordinates are used to locate the object candidate frame, for example, they can include the X, Y coordinates of the upper left corner and the X, Y coordinates of the lower right corner of the object candidate frame. Since the candidate frame is usually rectangular, the candidate frame can be located according to the X, Y coordinates of the upper left corner and the lower right corner of the candidate frame. Of course, the positioning coordinates can also be the coordinates of other points on the candidate frame, and the embodiments of the present application do not impose any restrictions on this.
[0035] In the embodiment of the present application, the image features output by at least two convolutional layers in the model are selected to determine a set of object candidate frames in the cargo image. The advantages of this are: since the sizes of the object candidate frames corresponding to the image features output by different convolutional layers are different, the sizes of the objects of the target category are also different; considering that the cargo in the cargo image is not necessarily of the same size, and even if the cargo is of the same size, the front and back distances of the cargo displayed in the image may be inconsistent due to the way it is placed, that is, the size of the cargo displayed in the image is inconsistent; therefore, by selecting the image features output by at least two convolutional layers in the model to determine the set of object candidate frames, objects of different positions and sizes in the cargo image can be obtained, thereby making the model have a higher recall rate and improving the accuracy of the model in identifying the quantity of cargo.
[0036] In one embodiment, the number of convolutional layers in the quantity recognition model is greater than 2, and the at least two convolutional layers include the last convolutional layer and at least one convolutional layer before it. Exemplarily, the number of convolutional layers in the quantity recognition model can be 7, so the image features output by the 5th, 6th, and 7th convolutional layers can be selected. Of course, in other embodiments, the image features output by the 4th, 5th, 6th, and 7th convolutional layers, or the image features output by the 6th and 7th convolutional layers can also be selected. This can be configured according to actual needs, and the embodiments of the present application do not impose any restrictions on this.
[0037] S103 : Screen and remove duplicate object candidate frames in the object candidate frame set according to the object categories in the image features and the positioning coordinates of the object candidate frames.
[0038] Each object candidate frame in the obtained object candidate frame set is derived from image features output by at least two convolutional layers. Some of these image features are target objects, while others are background objects. Therefore, object candidate frames need to be filtered based on object type. Furthermore, image features output by different convolutional layers may point to the same object, so duplicate object candidate frames need to be removed.
[0039] Specifically, we can retain the object candidate frames corresponding to the image features of the target object type. Then, among the filtered object candidate frames, we determine which ones have a high degree of overlap. These highly overlapping object candidate frames refer to the same object, and we can then remove duplicates. For example, we can determine the positional relationship between different object candidate frames based on their positioning coordinates, and then determine their degree of overlap based on this positional relationship.
[0040] S104: The number of target object candidate frames obtained after screening and deduplication is used as the number of goods in the goods image.
[0041] After screening and deduplication, target object candidate frames are obtained. The object pointed to in each target object candidate frame is the goods in the final identified image. Therefore, the number of target object candidate frames can be used as the number of goods in the goods image.
[0042] In the technical solution of the embodiment of the present application, a neural network model with multi-layer convolution is used to extract image features, and the image features output by at least two convolution layers in the model are selected to determine the set of object candidate frames in the cargo image. Then, through screening and deduplication, the target object candidate frames that are finally identified are determined from the set of object candidate frames, and the number of target object candidate frames is used as the number of cargo in the cargo image. Among them, since the image features output by at least two convolution layers in the model are selected, the model can identify objects of different positions and sizes in the cargo image, thereby discovering multiple targets in the image, accurately identifying the cargo in the image, and then improving the accuracy of the model in identifying the number of cargoes. In addition, the present application realizes cargo recognition based on a multi-layer convolutional neural network, which can extract features of different degrees and shapes of targets in the image, rather than just edge features and binary features, making the model highly robust compared to the existing technology; at the same time, it can also be applied to the recognition of cargo of different types and shapes, making it easier to expand in the future.
[0043] Example 2
[0044] Figure 2 This is a flow chart of the method for identifying the quantity of goods provided in Example 2 of this application. This example is further optimized based on the above example. Figure 2 As shown, the method includes:
[0045] S201. Extract image features from a cargo image using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers.
[0046] S202. Determine a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of the object candidate frames.
[0047] In one embodiment, a deconvolution layer is connected after at least one convolutional layer in a quantity recognition model, and the image features output by at least two convolutional layers are deconvolutional image features. The deconvolution layer is configured to expand the original image features output by the connected convolutional layer through deconvolution and add the expanded image features to the image features output by the convolutional layer preceding the connected convolutional layer.
[0048] For example, if a quantity recognition model has seven convolutional layers, at least one of the 3rd to 7th convolutional layers can be connected to a deconvolution layer. However, since the image feature size output by the second convolutional layer is too large, adding deconvolution would easily lead to a high computational load and reduced computational efficiency. Therefore, deconvolution is not required after the second convolutional layer.
[0049] Taking the deconvolution layer connected to the 7th convolution layer as an example, the 7th convolution layer outputs an original image feature through feature extraction, and then the original image feature is expanded through deconvolution. The expanded image feature is added to the image feature output by the 6th convolution layer, and the result of the addition is used as the image feature finally output by the 7th convolution layer, that is, the image feature selected in the embodiment of the present application.
[0050] The purpose of adding a deconvolution layer is to reorganize the features extracted by the convolutional layer through the deconvolution structure and add them to the image features output by the convolutional layer above it, thereby obtaining more detailed image features. In other words, the bottom-level features can be combined with the top-level features and used for prediction. This allows the model to more accurately identify the goods in the goods image and mark the location of the goods in the image with candidate boxes, thereby improving the model's accuracy in goods recognition.
[0051] S203: Select an object candidate frame whose object type is a target from the object candidate frame set.
[0052] S204: Determine the area and position of the candidate object box whose object type is the target according to the positioning coordinates.
[0053] S205 : Determine the degree of overlap between different object candidate frames based on their areas and positions, and remove duplicate object candidate frames whose object type is a target.
[0054] In step S204, the background is removed from the object candidate frame set, and then in step S205, duplicate object candidate frames representing the same object are removed to obtain the final identified Mubao object candidate frames. In specific implementation, a coincidence threshold can be configured. If the coincidence degree of any two object candidate frames is greater than the threshold based on position and area, the two object candidate frames are considered to correspond to the same object.
[0055] S206 : The number of target object candidate frames obtained after screening and deduplication is used as the number of goods in the goods image.
[0056] The technical solution of the embodiment of the present application uses a neural network model with multiple layers of convolution to extract image features, and selects image features output by at least two convolutional layers in the model to determine a set of object candidate frames in the cargo image, so that the model can identify objects of different positions and sizes in the cargo image, thereby discovering multiple targets in the image. Then, through screening and deduplication, the final identified target object candidate frames are determined from the object candidate frame set, and the number of target object candidate frames is used as the number of goods in the cargo image, thereby improving the recognition accuracy of the cargo quantity. At the same time, by adding a deconvolution network to the model structure, the model can detect objects of different positions and sizes in the image, further improving the accuracy of cargo recognition in the image, giving the model a higher recall rate, and thereby improving the accuracy of cargo quantity recognition.
[0057] Example 3
[0058] Figure 3 This is a flow chart of the method for identifying the quantity of goods provided in Example 3 of this application. This example is further optimized based on the above examples. Figure 3 As shown, the method includes:
[0059] S301. Extract image features from a cargo image using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers.
[0060] S302. Determine a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of the object candidate frames.
[0061] S303 : Screen and remove duplicate object candidate frames in the object candidate frame set according to the object categories in the image features and the positioning coordinates of the object candidate frames.
[0062] S304: The number of target object candidate frames obtained after screening and deduplication is used as the number of goods in the goods image.
[0063] S305 : Determine the center point of each target object candidate box according to the positioning coordinates of the target object candidate box.
[0064] S306: Display the center point of the target object candidate box on the cargo image.
[0065] To visualize cargo quantity identification and facilitate management, the corresponding management system can display the cargo image along with the center point of each identified target object candidate box, for example, by marking the center point with a green dot. These displayed center points can directly indicate whether cargo identification is accurate, whether any cargo has been missed, or whether the center point is located at a location other than the cargo to be identified.
[0066] S307 : In response to a click operation on the center point of the target object candidate frame displayed on the cargo image, reduce the quantity of cargo in the cargo image.
[0067] S308 : In response to a click operation on an object other than the target object represented by the center point on the cargo image, the quantity of the cargo in the cargo image is increased.
[0068] If errors occur in the identified goods, corrections can be made through the above steps S307 and S308. For example, if the identified center point is not a piece of goods, the manager can click on the displayed green center point to remove it, thereby reducing the number of goods in the goods image. If there are still goods that have not been identified, the manager can directly click on the location of the unidentified goods on the goods image. Based on the clicked location, the goods can be identified and the number of goods in the goods image can be increased. At the same time, the center point of the cross-section of the manually selected goods can also be displayed on the image. Through this interactive process, not only can the identification results of the goods and their quantity be intuitively viewed, but the initial identification results can also be corrected, improving the efficiency of goods management.
[0069] Furthermore, the method of the present embodiment further includes: obtaining new image training samples based on the click operation and the reduced or increased quantity of goods; and updating the quantity recognition model using the new image training samples. In other words, after correcting the initial recognition results, new training data can be generated based on the corrections to continuously improve the model's accuracy and recall. This allows for continuous iteration and training to optimize the model.
[0070] The technical solution of the embodiment of the present application uses a neural network model with multiple layers of convolution to extract image features, which is highly robust. Furthermore, the model selects image features output by at least two convolutional layers to determine a set of candidate object frames in the cargo image. This allows the model to identify objects of different locations and sizes in the cargo image, thereby discovering multiple targets in the image and improving the model's recall rate. The final identified target object candidate frames are then determined from the set of candidate object frames through screening and deduplication, and the number of target object candidate frames is used as the number of goods in the cargo image, thereby improving the accuracy of cargo quantity recognition. Furthermore, human-computer interaction can be used to correct the recognition results, further obtaining a more accurate cargo quantity and improving cargo management efficiency.
[0071] Example 4
[0072] Figure 4 This is a schematic diagram of the structure of the cargo quantity recognition device in this embodiment. This embodiment can be applied to image recognition of cargo images and obtain the quantity of cargo in the image, which involves the field of image recognition and deep learning technology. The device can implement the cargo quantity recognition method described in any embodiment of this application. Figure 4 As shown, the device 400 specifically includes:
[0073] An image feature extraction module 401 is used to extract image features from a cargo image using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers;
[0074] an object candidate frame set determination module 402, configured to determine a set of object candidate frames in the cargo image based on image features output by at least two convolutional layers in the quantity recognition model, wherein the image features include at least object categories and location coordinates of object candidate frames;
[0075] a screening and deduplication module 403, configured to screen and deduplication the object candidate frames in the object candidate frame set according to the object categories and the positioning coordinates of the object candidate frames in the image features;
[0076] The cargo quantity determination module 404 is configured to use the number of target object candidate frames obtained after the screening and deduplication as the cargo quantity in the cargo image.
[0077] Optionally, a deconvolution layer is connected after at least one convolution layer in the quantity recognition model, and the image features output by the at least two convolution layers are image features after deconvolution;
[0078] The deconvolution layer is used to expand the original image features output by the connected convolution layer through deconvolution, and add the expanded image features to the image features output by the convolution layer before the connected convolution layer.
[0079] Optionally, the number of convolutional layers in the quantity recognition model is greater than 2, and the at least two convolutional layers include the last convolutional layer and at least one convolutional layer before it.
[0080] Optionally, the object categories include target and background;
[0081] Accordingly, the screening and deduplication module 403 includes:
[0082] A screening unit, configured to select an object candidate frame whose object type is a target from the object candidate frame set;
[0083] an area and position determining unit, configured to determine the area and position of an object candidate frame whose object type is a target according to the positioning coordinates;
[0084] The deduplication unit is configured to determine the degree of overlap between different object candidate frames according to the areas and positions, and to dedupe the object candidate frames whose object type is a target.
[0085] Optionally, the device further includes a display module, specifically configured to:
[0086] Determining the center point of each target object candidate frame according to the positioning coordinates of the target object candidate frame;
[0087] The center point of the target object candidate box is displayed on the cargo image.
[0088] Optionally, the device further includes a first correction module, configured to:
[0089] In response to a click operation on a center point of the target object candidate frame displayed on the cargo image, the quantity of cargo in the cargo image is reduced.
[0090] Optionally, the device further includes a second correction module, configured to:
[0091] In response to a click operation on an object other than the target object represented by the center point on the cargo image, the quantity of cargo in the cargo image is increased.
[0092] Optionally, the device further includes a model training and updating module, configured to:
[0093] Obtaining new image training samples based on the click operation and the reduced or increased quantity of goods;
[0094] The quantity recognition model is updated using the new image training sample.
[0095] Optionally, the positioning coordinates include the X and Y coordinates of the upper left corner and the X and Y coordinates of the lower right corner of the object candidate box.
[0096] The cargo quantity identification device provided in the embodiments of the present application can execute the cargo quantity identification method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects of the execution method.
[0097] Example 5
[0098] Figure 5 This is a structural diagram of an electronic device provided in Example 5 of the present application. Figure 5 A block diagram of an exemplary electronic device 12 suitable for implementing embodiments of the present application is shown. Figure 5 The electronic device 12 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0099] like Figure 5 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).
[0100] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0101] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0102] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 5 Not shown, often called a "hard drive"). Although Figure 5Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.
[0103] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.
[0104] The electronic device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface 22. Furthermore, the electronic device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0105] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the cargo quantity identification method provided in the embodiment of the present application.
[0106] Example 6
[0107] Embodiment 6 of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for identifying the quantity of goods provided in the embodiment of the present application is implemented.
[0108] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer-readable media.Computer-readable media can be computer-readable signal media or computer-readable storage media.Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof.More specific examples (non-exhaustive list) of computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof.In this document, computer-readable storage media can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0109] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0110] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0111] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0112] In addition, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for identifying the quantity of goods provided in any of the above embodiments.
[0113] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for identifying the quantity of goods, characterized in that: include: Extracting image features from the cargo image using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers; Determine a set of object candidate frames in the cargo image according to image features output by at least two convolutional layers in the quantity recognition model, wherein the image features at least include object categories and positioning coordinates of object candidate frames; According to the object category in the image feature and the positioning coordinates of the object candidate boxes, the object candidate boxes in the object candidate box set are screened and duplicated; The number of target object candidate frames obtained after the screening and deduplication is used as the number of goods in the goods image; The object categories include target and background; Correspondingly, the screening and deduplication of the object candidate boxes in the object candidate box set according to the object category in the image feature and the positioning coordinates of the object candidate boxes includes: Selecting an object candidate frame whose object category is a target from the object candidate frame set; Determine the area and position of the object candidate box whose object category is the target according to the positioning coordinates; Determine the overlap between different object candidate frames according to the area and position, and remove duplicate object candidate frames whose object category is a target; The determining, based on the area and the position, the overlap between different object candidate frames, and removing duplicate object candidate frames whose object category is a target, includes: An overlap threshold is configured, and according to the area and the position, it is determined that the overlap of any two object candidate frames is greater than the overlap threshold, and duplicate object candidate frames whose object category is the target are removed.
2. The method according to claim 1, characterized in that In the quantity recognition model, at least one convolution layer is further connected to a deconvolution layer, and the image features output by the at least two convolution layers are image features after deconvolution; The deconvolution layer is used to expand the original image features output by the convolution layer connected to it through deconvolution, and add the expanded image features to the image features output by the convolution layer before the convolution layer connected to it.
3. The method according to claim 1, characterized in that The number of convolutional layers in the quantity recognition model is greater than 2, and the at least two convolutional layers include the last convolutional layer and at least one convolutional layer before it.
4. The method according to claim 1, characterized in that: Also includes: Determine the center point of each target object candidate frame according to the positioning coordinates of the target object candidate frame; The center point of the target object candidate box is displayed on the cargo image.
5. The method according to claim 4, characterized in that Also includes: In response to a click operation on a center point of the target object candidate frame displayed on the cargo image, the quantity of cargo in the cargo image is reduced.
6. The method according to claim 4, characterized in that Also includes: In response to a selection operation on an object other than the target object represented by the center point on the cargo image, the quantity of cargo in the cargo image is increased.
7. The method according to claim 5 or 6, characterized in that: Also includes: According to the click operation and the quantity of goods reduced or increased, a new image training sample is obtained; The quantity recognition model is updated using the new image training sample.
8. The method according to claim 1, characterized in that The positioning coordinates include the upper left corner X, Y coordinates, and the lower right corner X, Y coordinates of the object candidate box.
9. A cargo quantity identification device, characterized in that: include: An image feature extraction module, used to extract image features from cargo images using a pre-trained quantity recognition model, wherein the quantity recognition model has multiple convolutional layers; An object candidate frame set determination module, configured to determine a set of object candidate frames in the cargo image according to image features output by at least two convolutional layers in the quantity recognition model, wherein the image features at least include object categories and location coordinates of object candidate frames; A screening and deduplication module, configured to screen and deduplication the object candidate frames in the object candidate frame set according to the object categories in the image features and the positioning coordinates of the object candidate frames; A goods quantity determination module, used to use the number of target object candidate frames obtained after the screening and deduplication as the number of goods in the goods image; The screening and deduplication module includes: A screening unit, configured to select an object candidate frame whose object category is a target from the object candidate frame set; An area and position determination unit, configured to determine the area and position of an object candidate box whose object category is a target according to the positioning coordinates; a deduplication unit, configured to determine the overlap between different object candidate frames according to the areas and positions, and to deduplication the object candidate frames whose object category is a target; The determining, based on the area and the position, the overlap between different object candidate frames, and removing duplicate object candidate frames whose object category is a target, includes: An overlap threshold is configured, and according to the area and the position, it is determined that the overlap of any two object candidate frames is greater than the overlap threshold, and duplicate object candidate frames whose object category is the target are removed.
10. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the cargo quantity identification method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for identifying the quantity of goods as described in any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for identifying the quantity of goods according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for identifying number of targets, and computer readable storage medium
CN108921105A
Steel size and quantity identification method based on deep learning, intelligent equipment and storage medium
CN110929756A