Method for realizing stereoscopic warehouse material visual inventory based on deep learning

By using deep learning object detection algorithms in the stereo library, the number of parts is automatically identified, and the problems of inefficiency and reliability that rely on manual labor in the material inventory of the stereo library are solved, and unmanned management and low-cost material inventory are realized.

CN120494676APending Publication Date: 2025-08-15SAIC MAXUS AUTOMOTIVE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411899814.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The inventory of materials in existing three-dimensional libraries relies on manual operations, resulting in inefficiency and inability to quantify reliability and unmanned management.

Method used

The object detection algorithm based on deep learning is used to monitor the Liku part box through the camera, and the material image recognition is used by the Pytorch framework Yolov5-6.0 object detection algorithm, a material encoding library is created, and the number of parts in the library is detected in real time through the detect.py script, and feedback to the Liku system database.

Benefits of technology

It realizes automatic identification of the number of parts, reduces labor time waste, improves data reliability, reduces equipment losses, and realizes unmanned identification and low-cost management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494676A_ABST
    Figure CN120494676A_ABST
Patent Text Reader

Abstract

The invention discloses a method for realizing visual inventory of materials in a stereoscopic warehouse based on deep learning, and the method comprises the steps: monitoring parts in a part box of the stereoscopic warehouse through a camera, uploading a shot material picture to an industrial personal computer, and taking a target detection algorithm of installing a Pytorch frame in the industrial personal computer as a core, carrying out model training on the material pictures acquired in the previous step, creating a material code library according to model classification, acquiring real-time material codes, traversing the material code library, calling a corresponding detect.py script after matching is completed, carrying out quantity detection on the current pictures, and feeding back a detection result to a database of a three-dimensional system. The method has the advantages that waste of a large amount of labor hours is avoided, data reliability is high, unmanned recognition can be achieved, implementation cost is low, and equipment loss is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for realizing visual inventory of materials in a stereoscopic warehouse based on deep learning, and belongs to the technical field of automobile manufacturing. Background Art

[0002] At present, all inventory forms in the material storage area of ​​stereoscopic warehouses in the industry are completed manually, which wastes man-hours and is inefficient. Moreover, the reliability of manual inventory cannot be quantified and depends entirely on human consciousness. The visual inventory solution based on deep learning solves this problem very well. It is not only highly reliable, but also eliminates the need for special inventory. Because the stereoscopic warehouse has the characteristics of a closed environment, it only needs to identify the quantity at the entrance, so there is no need for a second inventory. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide a method for realizing visual inventory of materials in a stereoscopic warehouse based on deep learning, breaking through the original manual inventory function, using a deep learning-based target detection algorithm as the core to realize automatic identification of the number of parts, and solving the problem that when counting parts in a stereoscopic warehouse, the stacker takes a long time to retrieve materials and wastes labor costs.

[0004] In order to achieve the above objectives, the specific technical solution of the present invention is as follows: A method for realizing visual inventory of materials in a stereoscopic warehouse based on deep learning, characterized in that it includes the following specific steps:

[0005] Step 1: Monitor the parts in the parts box of the library through a camera, and upload the captured material pictures to an industrial computer with a Python compiler installed in it;

[0006] Step 2: Using the target detection algorithm of the Pytorch framework installed in the industrial computer as the core, perform model training on the material images collected in the previous steps.

[0007] Step 3. Create a material coding library by model classification, collect real-time material codes and traverse the material coding library. After matching, call the corresponding detect.py script. The detect.py script will perform quantity detection on the current image and feed back the detection results to the library system database.

[0008] Furthermore, in step 1, the parts are all packed in a uniform material box, so the relative position states are all uniform.

[0009] Furthermore, in step 1, the industrial computers are installed with Python compilers, the program compiler uses Python 3.8 version, the IDE uses PyCharm, the target detection algorithm is based on Yolov 5-6.0 of the Pytorch framework, and the material coding libraries code, code1, code2... are created according to model classification.

[0010] Furthermore, in step 2, when the material box reaches the entrance of the vertical warehouse, the vertical warehouse WCS system takes a photo of the material box and the PLC feeds back the photo completion signal "photo_comple" to the python main program script. At the same time, the python main program script reads the material tag information to obtain the material number "num". Python receives the "photo_comple" signal and starts to traverse the created material code library code with the obtained material number "num". When the current material code is matched in the code, the corresponding "detect.py" script is called to start the photo recognition process and write the quantity into the database

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] The system of the present invention is based on the Yolov5-6.0 target detection algorithm as the core. By making training and test data sets for each part, the collected images are segmented, rotated, translated, mosaiced, etc. to enhance the image data to enrich the data set. Then, the image data is convolved, pooled, upsampled, and spliced to finally obtain feature data of different scales of 32x32, 16x16, and 8x8. According to the detection frame manually marked in advance, the coordinate loss value, confidence loss value, and category loss value of the prediction frame and the marked frame are calculated by gradient descent and back propagation. The optimization w weight value and b bias value are not counted, and the final loss is minimized to obtain a model. Then, the material library code is established according to the model classification, the real-time material code is collected and traversed through the code material library. After matching, the corresponding detect.py script is called to detect the parts image at the entrance in real time, and the detection results are fed back to the library system database. This has the advantages of avoiding the waste of a large amount of man-hours, high data reliability, unmanned recognition, low implementation cost, and reduced equipment loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Schematic diagram of the architecture of the present invention.

[0014] Figure 2 Schematic diagram of the deep learning architecture in the present invention.

[0015] Figure 3 This is an architectural diagram of the Focus structure of the Backbone network in the present invention.

[0016] Figure 4 This is an architectural diagram of the CSP structure of the Backbone network in the present invention.

[0017] Figure 5 This is an architectural diagram of the SPP structure of the Backbone network in the present invention.

[0018] Figure 6 This is an architectural diagram of the Neck network in the present invention. DETAILED DESCRIPTION

[0019] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings:

[0020] This embodiment proposes a method for realizing visual inventory of materials in a vertical warehouse based on deep learning. Its architecture is as follows: Figure 1 As shown, by installing a camera at the entrance of the vertical warehouse, the parts in the parts box are monitored in real time, and the industrial computers are installed with Python compilers. In this embodiment, the compiler uses Python 3.8 version, the IDE uses PyCharm, and the target detection algorithm is Yolov5-6.0 based on the Pytorch framework.

[0021] 1. Create a material coding library code, code1, code2... according to model classification.

[0022] 2. Create a data set for parts and use labeling to annotate the data.

[0023] 3. Use the dataset for model training.

[0024] A. Design deep learning architecture

[0025] The architecture is as Figure 2 As shown, it is divided into four parts: Input, Backbone network, Neck network, and Prediction.

[0026] Input: Data enhancement, adaptive anchor box calculation, adaptive image scaling, and image initialization processing.

[0027] Among them, data enhancement is: randomly select four pictures, resize them and splice them together to enrich the data set.

[0028] Adaptive image scaling: In common object detection algorithms, different images have different lengths and widths, so a common approach is to scale the original image to a standard size before feeding it into the detection network. In this project, img_size is set to 640×640×3.

[0029] Backbone network: including Focus structure, CBL structure, CSP structure, and SPP structure, mainly used for feature extraction.

[0030] The Focus structure is as follows: the original 640×640×3 image input is sliced (sampling every other column) to become a 320×320×12 feature map. After concatenation (Concat), it undergoes a CBL operation and finally becomes a 320×320×64 feature map. Figure 3 shown.

[0031] The CBL structure is: use nn.Conv2d+nn.BatchNorm2d+nn.SiLU for convolution, batch normalization, and SiLU activation function.

[0032] The CSP structure is: Figure 4 As shown in the figure, the CSP structure divides the original input into two branches. One branch first passes through CBL, then through the residual structure (Resunit: borrowing the residual structure in the resent network to allow the network to be built deeper), and then performs a convolution; the other branch directly performs convolution; then the two branches are concat, so that the input and output are the same size, and then pass through BN (normal distribution), and then perform SiLU activation, so that the model can learn more features.

[0033] SPP structure: First, the input channels are halved through a standard convolution module, and then maxpooling with kernel-size of 5, 9, and 13 is performed respectively. The three results are concat with the unoperated data. The number of channels after the final merging is twice the original, such as Figure 5 shown.

[0034] Neck: The purpose is to better integrate the features extracted by the previous network. The rest of the network mainly uses FPN+PAN for downsampling and upsampling, and gives three feature maps of different scales for prediction, such as Figure 6 shown.

[0035] Prediction: Calculates category loss, confidence loss, and positioning loss, and ultimately obtains the prediction result.

[0036] B. Training Model

[0037] 2.1 Create training and test images, preferably with the target object photographed from various angles, locally and globally.

[0038] 2.2 Create labels for training and testing images. This project uses labelimg software for image recognition.

[0039] 2.3 Set the parameters of the model training script train.py: specify the initial weight file, model configuration file, and data configuration file, parser.add_argument('--weights', type = str, default = ", help = 'initialweightspath')

[0040] parser.add_argument('--cfg',type=str,default='models / yolov5s.yaml',

[0041] help='model.yaml path')

[0042] parser.add_argument('--data',type=str,default=ROOT / 'data / tools.yaml',

[0043] help='dataset.yaml path')

[0044] parser.add_argument('--epochs',type=int,default=300)

[0045] parser.add_argument('--batch-size',type=int,default=64,help='totalbatchsize for all GPUs')

[0046] parser.add_argument('--imgsz','--img','--img-size',type=int,default=640,help='train,val image size(pixels)')

[0047] Among them, weights: initial weight file, since it is a new category learning, it is set to empty.

[0048] Cfg: Model configuration file. This project uses yolov5s.yaml.

[0049] Data: Data configuration file, which defines the training and testing paths, number of categories, and category names of the images required by the model. The specific configuration is as follows:

[0050] train:.. / SAIC_MAXUS_VI / tool / images / train

[0051] val:.. / SAIC_MAXUS_VI / tool / images / val test:nc:18

[0052] names:

[0053] ['C00091495','C00239460','C00239490','C00239790','C00239830','C00239840','C0

[0054] 0244530','C00244550','C00244560','C00244630','C00244640','C00244670','C00244

[0055] 680','C00244710','C00244720','C00402002','C00478920','C00478930']#class

[0056] names

[0057] epochs: The number of times the model is trained. This project sets it to 300.

[0058] Batch-size: The number of images that the GPU training model processes at one time. It depends on the GPU memory. In this project, it is set to 64.

[0059] Imgsz: The size of the original image to be converted. This project sets it to 640x640.

[0060] Create a training and test image set: Create a tool folder in the SAIC_MAXUS_VI folder, and create two folders, images and labels, within the images folder. Create two folders, train and val, within the images folder. Place training images in the train folder and test images in the val folder. To ensure accuracy, this project uses 30 training images and 10 test images.

[0061] 4. Deploy the model on the upper system server of the library. First, create a material coding library: this project is named code, and the corresponding material codes of the model training images are input by category in the code.txt file.

[0062] Python environment connects to Siemens S-1500PLC

[0063] while True:

[0064] def set_connection_type(self,connection_type):

[0065] result=self.library.Cli_SetConnectionType(self.pointer,c_uint16(connection_type))

[0066] if result!=0:

[0067] raise Snap7Exception("The paramter wasinvalid")

[0068] plcObj=snap7.client.Client()

[0069] client.set_connection_type(2)

[0070] plcObj.connect('192.168.10.230',0,1)

[0071] Because the snap7 library uses PG mode connection by default, if you use Ethernet connection, you need to set set_connection_type() to 2-10.

[0072] plcObj.connect('192.168.10.230',0,1) sets the IP address, rack number, and CPU slot number respectively.

[0073] After reading the PLC photo completion signal and material coding information, the code library is traversed. After matching is OK, the detect.py script is called.

[0074] photo_comple=bool.from_bytes(data[0:1],byteorder='big')

[0075] tag=data[10:264].decode(encoding="ascii")

[0076] num=tag

[0077] if photo_comple == 1:

[0078] file=open('code.txt') reads the code library code

[0079] forlinein file.readlines(): Use the line variable to traverse each line of material code

[0080] curLine = line.strip() to match

[0081] if num == curLine: If it matches, call the following path script

[0082] os.system(r"python C:\Users\Administrator\Desktop\yolov5-6.0\detect.py")

[0083] break to exit the loop

[0084] detect.py parameter settings

[0085] parser.add_argument('--weights',nargs='+',type=str,default="best.pt",help='model path(s)')parser.add_argument('--source',type=str,default=ROOT / 'data / images',help='file / dir / URL / glob,0for webcam')

[0086] weights: Load the trained model best.pt.

[0087] source: Set the WCS save image path.

[0088] The main.py script calls detect.py to run, identify the real-time image data of the library entrance and obtain the detection results.

[0089] After the test is completed, the data will be saved in the result.txt file. Writing in the w+ mode can replace the original data so that the text always only saves the current data.

[0090] file2=open("result.txt",'w+')

[0091] file2.write(f"{n}\n")

[0092] file2.close()

[0093] WCS reads the text data and stores it in the WCS local database.

[0094] 5. When the material box reaches the entrance of the vertical warehouse, the vertical warehouse WCS system takes a photo of the material box and the PLC feeds back the photo completion signal "photo_comple" to the Python main program script. At the same time, the Python main program script reads the material tag information to obtain the material number "num". Python receives the "photo_comple" signal and starts to traverse the created material code library code with the obtained material number "num". When the current material code is matched in the code, the corresponding "detect.py" script is called to start identifying the photo and writing the quantity into the database.

[0095] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the specific embodiments described above. The specific embodiments and descriptions in the specification are merely intended to further illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for visual inventory of materials in a three-dimensional warehouse based on deep learning, characterized by: The specific steps include: Step 1: Monitor the parts in the parts box of the library through a camera, and upload the captured material pictures to an industrial computer with a Python compiler installed in it; Step 2: Using the target detection algorithm of the Pytorch framework installed in the industrial computer as the core, perform model training on the material images collected in the previous steps. Step 3. Create a material coding library by model classification, collect real-time material codes and traverse the material coding library. After matching, call the corresponding detect.py script. The detect.py script will perform quantity detection on the current image and feed back the detection results to the library system database.

2. The method for implementing visual inventory of materials in a three-dimensional warehouse based on deep learning according to claim 1 is characterized in that: In step 1, the parts are all packed in a uniform material box, so the relative position states are all uniform.

3. The method for implementing visual inventory of materials in a three-dimensional warehouse based on deep learning according to claim 1 is characterized in that: In step 1, the Python compiler is installed on the industrial computer, the program compiler uses Python 3.8 version, the IDE uses PyCharm, the target detection algorithm is based on Yolov 5-6.0 of the Pytorch framework, and the material coding libraries code, code1, code2... are created according to model classification.

4. The method for implementing visual inventory of materials in a three-dimensional warehouse based on deep learning according to claim 3 is characterized by: In step 2, when the material box reaches the entrance of the vertical warehouse, the vertical warehouse WCS system takes a photo of the material box and the PLC feeds back the photo completion signal "photo_comple" to the Python main program script. At the same time, the Python main program script reads the material tag information to obtain the material number "num". Python receives the "photo_comple" signal and starts to traverse the created material code library code with the obtained material number "num". When the current material code is matched in the code, the corresponding "detect.py" script is called to start identifying the photo and writing the quantity into the database.