Intelligent warehouse management method and system and medium
By using the deep learning object detection model in the warehousing management system to detect monitoring video frames, generate target trajectories and determine in-house and out-house events, the problem of poor flexibility and accuracy of cargo out-of-house events recognition in the prior art is solved, and efficient and accurate event recognition is achieved.
Patent Information
- Application Number
- CN202311680932.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the flexibility and accuracy of cargo outbound/incoming events are poor, and the target detection algorithm is prone to problems of repeated reporting and inaccurate event type identification, while radio frequency identification technology cannot identify personnel in and outbound events.
The deep learning object detection model is used to detect the monitoring video frames captured by the camera device, generate the target trajectory, and determine the in-store and exit events based on the trajectory, and send them to the monitoring terminal. This method can detect equipment targets and pedestrian targets simultaneously without installing identification equipment, improving the flexibility and accuracy of event recognition.
Through the use of deep learning object detection model, the problem of repeated events is avoided. A target trajectory can only generate in-store events once, which improves the accuracy and flexibility of event recognition, and can accurately identify in-store event types.
Smart Images

Figure CN120126064A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to an intelligent warehouse management method, system and medium. Background Art
[0002] With the development of Internet technology and social progress, the intelligent management of warehouse areas has become increasingly popular, and the identification of goods outbound / inbound events is particularly important for warehouse management.
[0003] In the prior art, target detection algorithms or radio frequency identification (RFID) technology are usually used to identify goods outbound / inbound events. When using a target detection algorithm for event identification, the warehouse management system can use the target detection algorithm to identify the video images captured by a camera. If a detection target is identified, it can be considered that a goods outbound / inbound event has occurred. When using RFID technology for event identification, staff can install RFID radio frequency devices on equipment such as forklift trays in advance. If the warehouse management system receives the signal emitted by the RFID radio frequency device, it can be considered that a goods outbound / inbound event has occurred. However, both of these event identification methods have certain defects: the target detection algorithm can identify the detection target, but when there are a large number of goods and personnel or equipment repeatedly perform handling operations, there will be a large number of repeated reports, and the target detection algorithm cannot identify whether the target is outbound or inbound, the event identification is not flexible, and the identification accuracy is poor. Although RFID technology can accurately identify the inbound / outbound events related to equipment, it cannot identify when personnel perform inbound / outbound events, and staff need to install RFID radio frequency devices on the equipment in advance, and it cannot identify when new equipment appears, the event identification is not flexible, and the identification accuracy is poor.
[0004] Therefore, an intelligent warehouse management solution that can improve the flexibility and accuracy of goods outbound / inbound event identification is needed. Summary of the Invention
[0005] This application provides an intelligent warehouse management method, system and medium to solve the technical problem of poor flexibility and accuracy in the identification of existing goods outbound / inbound events.
[0006] In a first aspect, this application provides an intelligent warehouse management method, including:
[0007] Obtain the monitoring video sent by the camera device, and parse to obtain the real-time video frame corresponding to the monitoring video, where the camera device is used to capture video images of the warehouse area;
[0008] Use a deep learning object detection model to detect the real-time video frame to detect whether there is an identification target in the real-time video frame, where the identification target is a device target or a pedestrian target;
[0009] If there is an identification target, generate a corresponding target trajectory according to the video frame in which the identification target is detected;
[0010] After the identification target becomes inactive, generate a corresponding inbound / outbound event according to the target trajectory and send the inbound / outbound event to the monitoring terminal.
[0011] In a possible implementation manner, the generating a corresponding target trajectory according to the video frame in which the identification target is detected specifically includes:
[0012] Obtain the tracking targets in the current tracking queue and determine whether there is a first target that matches the identification target among the tracking targets;
[0013] If there is, merge and fuse the identification target with the first target, and generate a corresponding target trajectory according to the target video frame in which the first target is detected;
[0014] If not, use the identification target as a new tracking target and put it into the current tracking queue for tracking.
[0015] In a possible implementation manner, the generating a corresponding target trajectory according to the target video frame in which the first target is detected specifically includes:
[0016] Determine each target video frame in which the first target is detected, and determine the trajectory points corresponding to the first target in each target video frame;
[0017] Connect the trajectory points corresponding to the first target in each target video frame in the order of the generation sequence of the target video frames to generate the target trajectory corresponding to the first target.
[0018] In a possible implementation manner, the determining whether there is a first target that matches the identification target among the tracking targets specifically includes:
[0019] For each tracking target, determine the first video frame with the closest time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first trajectory point of the tracking target in the first video frame and the second trajectory point of the identification target in the real-time video frame; determine whether the distance between the first trajectory point and the second trajectory point is less than a distance threshold; if so, the tracking target is the first target that matches the identification target; if not, the tracking target does not match the identification target;
[0020] Alternatively,
[0021] For each tracking target, determine the first video frame that is closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first feature of the tracking target in the first video frame and the second feature of the recognition target in the real-time video frame; determine whether the feature similarity between the first feature and the second feature is greater than the similarity threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
[0022] Alternatively,
[0023] For each tracking target, determine the first video frame that is closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first image of the tracking target in the first video frame and the second image of the recognition target in the real-time video frame; determine whether the image overlap degree between the first image and the second image is greater than the overlap degree threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
[0024] In a possible implementation manner, after the recognition target is deactivated, generating a corresponding inbound / outbound event according to the target trajectory specifically includes:
[0025] Determine the latest generated target trajectory point in the target trajectory and determine the generation time of the video frame corresponding to the target trajectory point;
[0026] Determine the duration between the generation time of the video frame and the current time, and determine whether the duration is greater than the duration threshold;
[0027] If not, the recognition target is not deactivated, and continue to execute the step of parsing the real-time video frame corresponding to the surveillance video;
[0028] If so, the recognition target is deactivated, and generate a corresponding inbound / outbound event according to the target trajectory.
[0029] In a possible implementation manner, generating a corresponding inbound / outbound event according to the target trajectory specifically includes:
[0030] Determine the position of the warehouse door of the storage area, and determine the event type according to the warehouse door position and the target trajectory, where the event type includes an outbound event and an inbound event;
[0031] Determine the target type of the identified target, and determine the goods information corresponding to the identified target according to the target type, where the target type includes equipment type and pedestrian type;
[0032] Generate a corresponding inbound / outbound event according to one or more of the identified target, the event type, the goods information, and the target trajectory.
[0033] In a possible implementation manner, the determining the goods information corresponding to the identified target according to the target type specifically includes:
[0034] If the target type of the identified target is the equipment type, use the deep learning recognition model for the number of goods pallets to recognize the video frame where the identified target is detected, so as to determine the number of goods pallets corresponding to the identified target, and determine the goods information corresponding to the identified target according to the number of goods pallets;
[0035] If the target type of the identified target is the pedestrian type, use the deep learning recognition model for whether a person is holding goods to recognize the video frame where the identified target is detected, so as to determine whether the identified target is holding goods, and determine the goods information corresponding to the identified target according to whether the identified target is holding goods;
[0036] Among them, the deep learning recognition model for the number of goods pallets is trained with multiple first images as samples, and the first images are images of the equipment holding different numbers of goods pallets respectively; the deep learning recognition model for whether a person is holding goods is trained with multiple second images as samples, and the second images are images of a person holding goods and a person not holding goods.
[0037] In a second aspect, the present application provides an intelligent warehouse management system, including:
[0038] An acquisition module, configured to acquire the monitoring video sent by the camera device, and parse to obtain the real-time video frame corresponding to the monitoring video, where the camera device is used to capture video images of the warehouse area;
[0039] A processing module, configured to use the deep learning target detection model to detect the real-time video frame to detect whether there is an identified target in the real-time video frame, where the identified target is an equipment target or a pedestrian target; if there is an identified target, generate a corresponding target trajectory according to the video frame where the identified target is detected; after the identified target is deactivated, generate a corresponding inbound / outbound event according to the target trajectory, and send the inbound / outbound event to the monitoring terminal.
[0040] In a third aspect, the present application provides another intelligent warehouse management system, including: a processor and a memory communicatively connected to the processor;
[0041] The memory stores computer-executable instructions;
[0042] The processor executes the computer-executable instructions stored in the memory to implement the above method.
[0043] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above method when executed by a processor.
[0044] In a fifth aspect, the present application provides a computer program product including a computer program, which implements the above method when executed by a processor.
[0045] The intelligent warehousing management method, system and medium provided by the present application can obtain the monitoring video sent by a camera device, and parse the real-time video frames corresponding to the monitoring video. The camera device is used to capture video images of a warehousing area; use a deep learning object detection model to detect the real-time video frames to detect whether there is an identification target in the real-time video frames, and the identification target is a device target or a pedestrian target; if there is an identification target, generate a corresponding target trajectory according to the video frames in which the identification target is detected; after the identification target is deactivated, generate a corresponding inbound / outbound event according to the target trajectory, and send the inbound / outbound event to a monitoring terminal. The method of the present application uses a deep learning object detection model to detect the identification target, can detect both device targets and pedestrian targets at the same time, and does not require installing identification devices on device targets or pedestrian targets, improving the flexibility and accuracy of identifying inbound / outbound events of goods. Further, use a deep learning object detection model to detect whether there is an identification target in the real-time video frames, and after detecting the identification target, generate the target trajectory corresponding to the identification target according to all the video frames in which the identification target is detected, and after the identification target is deactivated, generate a corresponding inbound / outbound event according to the target trajectory. Through such a setting, corresponding inbound / outbound events are generated according to the target trajectory, avoiding the problem of repeated reporting of events caused by the target appearing multiple times during the detection process. Only one inbound / outbound event can be generated for one target trajectory. In addition, the target trajectory can be used to accurately identify whether the target is inbound or outbound, further improving the flexibility and accuracy of identifying inbound / outbound events of goods. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0047] Figure 1 It is a system architecture diagram of an embodiment of the present application;
[0048] Figure 2 Flow chart of the intelligent warehousing management method according to an embodiment of the present application;
[0049] Figure 3 Flow chart of the intelligent warehousing management method according to another embodiment of the present application;
[0050] Figure 4 Structural schematic diagram of the intelligent warehousing management system according to an embodiment of the present application;
[0051] Figure 5 Structural schematic diagram of the intelligent warehousing management system according to another embodiment of the present application.
[0052] Reference numerals: 1, camera device; 2, intelligent warehousing management system; 3, monitoring terminal.
[0053] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be given later. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0054] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0055] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0056] It should also be noted that the intelligent warehousing management method, system, and medium of the present application can be used in the field of data processing, and can also be used in any field other than the field of data processing, such as the field of artificial intelligence, the field of intelligent warehousing, etc. The application fields of the intelligent warehousing management method, system, and medium of the present application are not limited.
[0057] First, the nouns involved in the present application are explained:
[0058] Radio Frequency Identification (RFID) technology, whose principle is non-contact data communication between a reader and a tag to achieve the purpose of identifying a target.
[0059] In the prior art, target detection algorithms or Radio Frequency Identification (RFID) technology are usually used to identify goods outbound / inbound events. When using a target detection algorithm for event identification, a warehouse management system can use the target detection algorithm to identify the video images captured by a camera. If a detection target is identified, it can be considered that a goods outbound / inbound event has occurred. When using RFID technology for event identification, staff can install an RFID radio device on a device such as a forklift pallet in advance. If the warehouse management system receives a signal emitted by the RFID radio device, it can be considered that a goods outbound / inbound event has occurred.
[0060] However, both of these event identification methods have certain defects: The target detection algorithm can identify the detection target, but when there are a large number of goods and personnel or equipment repeatedly carry out handling work, there will be a large number of repeated reports, and the target detection algorithm cannot identify whether the target is outbound or inbound, so the event identification is not flexible and the identification accuracy is poor. Although RFID technology can accurately identify equipment-related inbound / outbound events, it cannot identify when personnel carry out inbound / outbound events, and staff need to install the RFID radio device on the equipment in advance and cannot identify when new equipment appears, so the event identification is not flexible and the identification accuracy is poor.
[0061] Based on this technical problem, the inventive concept of this application lies in: how to provide an intelligent warehouse management method that can improve the flexibility and accuracy of goods outbound / inbound event identification.
[0062] Specifically, it is possible to obtain the surveillance video sent by the imaging device and parse it to obtain the real-time video frames corresponding to the surveillance video. The imaging device is used to capture video images of the storage area; use the deep learning object detection model to detect the real-time video frames to detect whether there are recognition targets in the real-time video frames. The recognition targets are equipment targets or pedestrian targets; if there are recognition targets, generate corresponding target trajectories based on the video frames in which the recognition targets are detected; after the recognition targets are deactivated, generate corresponding inbound / outbound events based on the target trajectories, and send the inbound / outbound events to the monitoring terminal. The method of this application uses the deep learning object detection model to detect recognition targets, which can detect equipment targets or pedestrian targets at the same time, and there is no need to install recognition devices on the equipment targets or pedestrian targets, improving the flexibility and accuracy of identifying inbound / outbound events of goods. Further, use the deep learning object detection model to detect whether there are recognition targets in the real-time video frames, and after detecting the recognition targets, generate the target trajectory corresponding to the recognition target based on all the video frames in which the recognition targets are detected, and after the recognition targets are deactivated, generate corresponding inbound / outbound events based on the target trajectories. Through such settings, corresponding inbound / outbound events are generated based on the target trajectories, avoiding the problem of repeated reporting of events caused by the target appearing multiple times during the detection process. Only one inbound / outbound event can be generated for one target trajectory. In addition, the target trajectory can be used to accurately identify whether the target is inbound or outbound, further improving the flexibility and accuracy of identifying inbound / outbound events of goods.
[0063] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the drawings.
[0064] Figure 1 is a system architecture diagram of an embodiment of this application, as Figure 1 shown, 1 is the imaging device, 2 is the intelligent warehouse management system, and 3 is the monitoring terminal. The imaging device 1 set in the storage area can capture the surveillance video within its shooting area and send the surveillance video to the intelligent warehouse management system 2. After receiving it, the intelligent warehouse management system 2 parses the surveillance video to obtain the corresponding real-time video frames; uses the deep learning object detection model to detect the real-time video frames to detect whether there are recognition targets in the real-time video frames; if there are recognition targets, generate corresponding target trajectories based on the video frames in which the recognition targets are detected; after the recognition targets are deactivated, generate corresponding inbound / outbound events based on the target trajectories, and send the inbound / outbound events to the monitoring terminal 3.
[0065] Embodiment 1
[0066] Figure 2 The figure is a flowchart of an intelligent warehouse management method according to an embodiment of the present application. In this embodiment, the execution subject is an intelligent warehouse management system to illustrate the intelligent warehouse management method. As Figure 2 shown, the intelligent warehouse management method may include the following steps:
[0067] S101: Obtain the monitoring video sent by the camera device, and parse it to obtain the real-time video frames corresponding to the monitoring video.
[0068] In this embodiment, the camera device can be used to capture video images of the warehouse area. The camera device can be a camera or other devices capable of capturing video images. The warehouse area can be a warehouse or other areas capable of storing goods. The camera device can be installed at any position in the warehouse area as long as it can capture the video images of the goods entering and leaving the warehouse area.
[0069] In this embodiment, the intelligent warehouse management system can be a server or a part of the server. The camera device can be communicatively connected to the intelligent warehouse management system and upload the captured monitoring video to the intelligent warehouse management system in real time. After receiving the monitoring video, the intelligent warehouse management system can parse the monitoring video to obtain the corresponding real-time video frames.
[0070] In this embodiment, the specific method of parsing the monitoring video to obtain the real-time video can refer to the prior art and will not be elaborated here.
[0071] S102: Use the deep learning object detection model to detect the real-time video frames to detect whether there are recognition targets in the real-time video frames.
[0072] In this embodiment, the recognition target can be a device target or a pedestrian target.
[0073] In this embodiment, those skilled in the art can pre-set the recognition target according to actual needs, and there are no specific restrictions on the setting of the recognition target here. Exemplarily, the recognition targets set in a certain warehouse can be forklifts and pedestrians.
[0074] In this embodiment, the deep learning object detection model in the above step S102 can adopt the object detection model in the prior art, as long as it is a deep learning model trained using the object detection algorithm, and there are no specific restrictions here.
[0075] In this embodiment, the specific process of using the deep learning object detection model to detect the real-time video frames to detect whether there are recognition targets in the real-time video frames can refer to the process of object detection using the object detection algorithm in the prior art and will not be elaborated here.
[0076] S103: If there is an identified target, generate a corresponding target trajectory based on the video frame in which the identified target is detected.
[0077] In this embodiment, if there is no identified target, the above-mentioned step S101 of obtaining the surveillance video sent by the imaging device and parsing the real-time video frame corresponding to the surveillance video may be continuously executed until an identified target is detected.
[0078] In a possible implementation manner, generating a corresponding target trajectory based on the video frame in which the identified target is detected in the above-mentioned step S103 may include:
[0079] S1031: Obtain the tracking targets in the current tracking queue and determine whether there is a first target that matches the identified target among the tracking targets.
[0080] S1032: If there is, merge and fuse the identified target with the first target, and generate a corresponding target trajectory based on the target video frame in which the first target is detected.
[0081] S1033: If there is no such target, use the identified target as a new tracking target and put it into the current tracking queue for tracking.
[0082] In this implementation manner, the tracking targets in the current tracking queue may be identified targets detected in multiple video frames before the real-time video frame. That is, when an identified target A is first detected, it will be placed in the tracking queue and tracked as a tracking target until the identified target A cannot be tracked within a certain period of time (the identified target A becomes inactive), and then the identified target A will be deleted from the tracking queue and this event ends. When the identified target A is identified again, the identified target A can be put back into the tracking queue to start the next event.
[0083] In this implementation manner, if the first target matches the identified target, it means that the first target and the identified target are the same target, and the target (and trajectory) can be merged and fused.
[0084] In this implementation manner, since the imaging device is continuously capturing the surveillance video in real time and continuously parsing the video frames obtained from the surveillance video, identified targets may also be detected in multiple video frames before the real-time video frame. These identified targets may be the same or different. To avoid duplicate reporting of events and improve the accuracy of event identification, these identified targets can be classified and fused, that is, the targets that can be matched are merged and fused, and the targets that cannot be matched are used as new tracking targets and put into the current tracking queue for tracking.
[0085] In a possible implementation, generating a corresponding target trajectory based on the target video frame of the first target in step S1032 above may include: determining each target video frame in which the first target is detected, and determining the trajectory points corresponding to the first target in each target video frame; concatenating the trajectory points corresponding to the first target in each target video frame in the order of generation of the target video frames to generate the target trajectory corresponding to the first target.
[0086] In this implementation, when generating the target trajectory corresponding to the first target, all the trajectory points corresponding to the first target can be obtained first, and then the trajectory points corresponding to the first target in each target video frame are concatenated in the order of generation of the target video frames to obtain the target trajectory. Of course, after obtaining a new trajectory point, the new trajectory point can also be concatenated with the previous trajectory points. After concatenating to the last generated trajectory point, the target trajectory can be obtained.
[0087] In this implementation, by concatenating the trajectory points corresponding to the first target in each target video frame in the order of generation of the target video frames, the target trajectory corresponding to the first target can be generated simply and accurately.
[0088] In a possible implementation, determining whether there is a first target in the tracking target that matches the recognition target in step S1031 above may include: for each tracking target, determining the first video frame closest to the real-time video frame in the video frame in which the tracking target is detected; obtaining the first trajectory point of the tracking target in the first video frame and the second trajectory point of the recognition target in the real-time video frame; determining whether the distance between the first trajectory point and the second trajectory point is less than the distance threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
[0089] In this implementation, the distance threshold may be the maximum moving distance of the same target in two adjacent video frames. When the distance between the first trajectory point and the second trajectory point is greater than the distance threshold, it can be considered that the recognition target and the first target are not the same target, that is, they do not match. Those skilled in the art can flexibly set the distance threshold according to the specific recognition target and no limitation is made here.
[0090] In this implementation, by using the distance between the first trajectory point and the second trajectory point of the target in the two video frames closest in time, it can be simply and accurately determined whether the recognition target and the first target are the same target, that is, whether they match.
[0091] Alternatively, determining whether there is a first target that matches the recognition target in step S1031 above may further include: for each tracking target, determining a first video frame that is closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtaining a first feature of the tracking target in the first video frame and a second feature of the recognition target in the real-time video frame; determining whether the feature similarity between the first feature and the second feature is greater than a similarity threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
[0092] In this embodiment, the first feature may be an image feature of the tracking target extracted from the first video frame, and the second feature may be an image feature of the recognition target extracted from the real-time video frame. The extraction methods of the first feature and the second feature may refer to the image feature extraction methods in the prior art and will not be elaborated here. Similarly, the calculation method of the feature similarity between the first feature and the second feature may also refer to the feature similarity calculation methods in the prior art and will not be elaborated here.
[0093] In this embodiment, those skilled in the art can flexibly set the similarity threshold according to the actual situation and no restrictions are imposed here.
[0094] In this embodiment, by using the feature similarity between the first feature and the second feature of the target in the two video frames that are closest in time, it is possible to simply and accurately determine whether the recognition target and the first target are the same target, that is, whether they match.
[0095] Alternatively, determining whether there is a first target that matches the recognition target in step S1031 above may also include: for each tracking target, determining a first video frame that is closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtaining a first image of the tracking target in the first video frame and a second image of the recognition target in the real-time video frame; determining whether the image overlap degree between the first image and the second image is greater than an overlap degree threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
[0096] In this embodiment, the overlap degree threshold may be the minimum overlap degree of the images of the same target in two adjacent video frames. When the image overlap degree between the first image and the second image is greater than the overlap degree threshold, it can be considered that the recognition target and the first target are the same target, that is, they match. Those skilled in the art can flexibly set the overlap degree threshold according to the specific recognition target and no restrictions are imposed here.
[0097] In this embodiment, by using the image overlap degree between the first image and the second image of the target in the two most recent video frames, it is possible to simply and accurately determine whether the recognized target and the first target are the same target, that is, whether they match.
[0098] S104: After the recognized target is deactivated, generate corresponding inbound / outbound events according to the target trajectory, and send the inbound / outbound events to the monitoring terminal.
[0099] In this embodiment, for the specific implementation of generating corresponding inbound / outbound events according to the target trajectory after the recognized target is deactivated in step S104 above, please refer to Embodiment 2.
[0100] In this embodiment, after a certain recognized target is detected for the last time, if the recognized target is not detected within a preset time interval, it can be considered that the recognized target is deactivated, and the event of this recognized target has ended and can be reported. If the recognized target is not deactivated, it means that the recognized target can still be tracked, the event has not ended, and it can still be reported.
[0101] In this embodiment, using a deep learning object detection model to detect recognized targets can simultaneously detect device targets or pedestrian targets, and there is no need to install recognition devices on device targets or pedestrian targets, which improves the flexibility and accuracy of cargo inbound / outbound event recognition. Further, use a deep learning object detection model to detect whether there are recognized targets in real-time video frames, and after detecting a recognized target, generate a target trajectory corresponding to the recognized target according to all video frames in which the recognized target is detected, and after the recognized target is deactivated, generate corresponding inbound / outbound events according to the target trajectory. Through such a setting, corresponding inbound / outbound events are generated according to the target trajectory, avoiding the problem of repeated reporting of events caused by the target appearing multiple times during the detection process. Only one inbound / outbound event can be generated for one target trajectory. In addition, using the target trajectory can accurately identify whether the target is inbound or outbound, further improving the flexibility and accuracy of cargo inbound / outbound event recognition.
[0102] The following takes a specific Embodiment 2 to elaborate in detail on the implementation manner of generating corresponding inbound / outbound events according to the target trajectory after the recognized target is deactivated in step S104 of the above Embodiment 1.
[0103] Embodiment 2
[0104] Figure 3 It is a flowchart of an intelligent warehousing management method according to another embodiment of the present application. In this embodiment, the intelligent warehousing management method is described with the execution subject being an intelligent warehousing management system. As Figure 3 shown, the intelligent warehousing management method may include the following steps:
[0105] S201: Determine the latest generated target trajectory point in the target trajectory, and determine the generation time of the video frame corresponding to the target trajectory point.
[0106] In this embodiment, the latest generated target trajectory point in the target trajectory may be the last generated trajectory point in the target trajectory, that is, the trajectory point generated by the video frame when the recognition target is last detected.
[0107] S202: Determine the duration between the generation time of the video frame and the current time, and determine whether the duration is greater than the duration threshold.
[0108] In this embodiment, those skilled in the art can flexibly set the duration threshold according to the actual situation, and no limitation is made here.
[0109] S203: If not, it is recognized that the target is not deactivated, and continue to execute the step of parsing the real-time video frame corresponding to the monitored video.
[0110] In this embodiment, if it is recognized that the target is not deactivated, the step of parsing the real-time video frame corresponding to the monitored video can be continued, and the new generated video frame can be used to continue tracking the trajectory of the recognition target.
[0111] S204: If so, it is recognized that the target is deactivated, and generate a corresponding inbound / outbound event according to the target trajectory.
[0112] In a possible implementation manner, the generating a corresponding inbound / outbound event according to the target trajectory in step S204 above may include:
[0113] S2041: Determine the position of the warehouse door in the storage area, and determine the event type according to the position of the warehouse door and the target trajectory. The event type includes an outbound event and an inbound event.
[0114] S2042: Determine the target type of the recognition target, and determine the cargo information corresponding to the recognition target according to the target type. The target type includes an equipment type and a pedestrian type.
[0115] S2043: Generate a corresponding inbound / outbound event according to one or more of the recognition target, the event type, the cargo information, and the target trajectory.
[0116] In this implementation manner, the position of the warehouse door in the storage area may be pre-stored by the staff in the intelligent warehouse management system, or may be obtained by the intelligent warehouse management system after parsing the monitored video sent by the camera device and identifying the position of the warehouse door in the video frame.
[0117] In this embodiment, since the target trajectory is obtained by connecting each trajectory point in the order of generation time, the moving direction of the recognition target can be determined through the target trajectory. Therefore, according to the position of the warehouse door and the target trajectory, it is possible to simply and accurately determine whether the recognition target is leaving or entering the warehouse, so as to determine the event type, further improving the accuracy of event recognition. Further, since the recognition target may be an equipment target or a pedestrian target, it is also necessary to determine the target type of the recognition target and determine the corresponding cargo information of the recognition target according to the target type, thereby improving the accuracy of the event generated according to the cargo information.
[0118] In a possible implementation manner, determining the cargo information corresponding to the recognition target according to the target type in step S2042 may include:
[0119] S21: If the target type of the recognition target is an equipment type, use the deep learning recognition model for the number of cargo pallets to recognize the video frame detecting the recognition target to determine the number of cargo pallets corresponding to the recognition target, and determine the cargo information corresponding to the recognition target according to the number of cargo pallets.
[0120] S22: If the target type of the recognition target is a pedestrian type, use the deep learning recognition model for a person carrying goods to recognize the video frame detecting the recognition target to determine whether the recognition target is carrying goods, and determine the cargo information corresponding to the recognition target according to whether the recognition target is carrying goods.
[0121] Among them, the deep learning recognition model for the number of cargo pallets is trained with a plurality of first images as samples, and the first images are images of the equipment carrying different numbers of cargo pallets; the deep learning recognition model for a person carrying goods is trained with a plurality of second images as samples, and the second images are images of a person carrying goods and a person not carrying goods.
[0122] In this embodiment, the number of cargo pallets may be the quantity obtained by a pallet of equipment such as a forklift at one time. Generally, the number of cargo pallets may be 0.5 pallet, 1 pallet, 1.5 pallets, 2 pallets, etc.
[0123] Exemplarily, the first images may be images of a forklift carrying different numbers of cargo pallets such as 0 pallets (not carrying goods), 0.5 pallet, 1 pallet, 1.5 pallets, 2 pallets, etc. The number of images corresponding to different numbers of cargo pallets can be set arbitrarily, and can be the same or different. The second images may be images of multiple persons carrying goods and not carrying goods respectively.
[0124] In this embodiment, since the recognition target may be an equipment target or a pedestrian target, different deep learning recognition models corresponding to different target types can be used to specifically determine the cargo information corresponding to the recognition target, improving the accuracy of cargo information determination. Further, by using images of equipment carrying different numbers of pallets of goods to train the deep learning recognition model for the number of pallets of goods, the trained deep learning recognition model for the number of pallets of goods can learn the ability to recognize the number of pallets of goods in the recognition image; by using images of a person holding goods and a person not holding goods to train the deep learning recognition model, the trained deep learning recognition model for a person holding goods can learn the ability to recognize whether a pedestrian is holding goods, further improving the accuracy of determining the cargo information of the corresponding type of target.
[0125] In this embodiment, the time duration between the generation time of the video frame corresponding to the target trajectory point and the current time is the time duration during which the recognition target has not been detected. If this time duration is greater than the time duration threshold, it can be considered that the recognition target is deactivated, and this event of this recognition target has ended. The corresponding inbound / outbound event can be generated according to the target trajectory. If this time duration is not greater than the time duration threshold, it indicates that the recognition target has not been deactivated and the event has not ended. The step of parsing the real-time video frame corresponding to the monitored video can be continued, that is, the recognition target can be continuously tracked. By determining whether the recognition target is deactivated and generating the corresponding inbound / outbound event according to the target trajectory after the recognition target is deactivated, the problem of repeated reporting of events caused by the target appearing multiple times during the detection process can be avoided. Only one inbound / outbound event can be generated for one target trajectory, improving the flexibility and accuracy of the recognition of cargo outbound / inbound events.
[0126] Next, a specific embodiment is used to elaborate on the intelligent warehousing management method of the present application.
[0127] Embodiment III
[0128] In a specific embodiment, a camera is installed in a certain warehouse. The intelligent warehousing management system uses the monitored video captured by the camera to recognize the cargo outbound / inbound events in this warehouse. The specific intelligent warehousing management process is as follows:
[0129] First step, the camera captures the monitored video of this warehouse and sends the monitored video to the intelligent warehousing management system. The intelligent warehousing management system parses the received monitored video to obtain the corresponding real-time video frame.
[0130] Second step, the intelligent warehousing management system uses the deep learning target detection model to detect the real-time video frame and detects that there is a recognition target in the real-time video frame: a forklift.
[0131] In the third step, the intelligent warehousing management system obtains the tracking targets in the current tracking queue, determines the first target in the tracking targets that matches the identification target (forklift), merges and fuses the identification target with the first target, and generates a corresponding target trajectory based on the target video frame in which the forklift is detected.
[0132] In the fourth step, the intelligent warehousing management system determines the newly generated target trajectory point in the target trajectory, determines the generation time of the video frame corresponding to the target trajectory point, determines the duration between the generation time of the video frame and the current time, and determines that the duration is greater than the duration threshold, that is, the identification target is deactivated.
[0133] In the fifth step, the intelligent warehousing management system determines the position of the warehouse door in the warehousing area, and determines that the event type is an inbound event based on the position of the warehouse door and the target trajectory.
[0134] In the sixth step, the intelligent warehousing management system uses the deep learning recognition model of the number of goods pallets to recognize the video frame in which the identification target is detected, so as to determine the number of goods pallets corresponding to the identification target, and determine the goods information corresponding to the identification target according to the number of goods pallets.
[0135] In the seventh step, the intelligent warehousing management system generates a corresponding inbound / outbound event based on the identification target (forklift), event type (inbound event), goods information (number of goods pallets), and target trajectory, and sends the inbound / outbound event to the monitoring terminal.
[0136] Figure 4 It is a schematic structural diagram of the intelligent warehousing management system according to an embodiment of the present application. As Figure 4 shown, the intelligent warehousing management system includes: an acquisition module 41, configured to acquire the monitoring video sent by the imaging device and parse it to obtain the real-time video frame corresponding to the monitoring video, and the imaging device is used to capture the video image of the warehousing area; a processing module 42, configured to use the deep learning target detection model to detect the real-time video frame to detect whether there is an identification target in the real-time video frame, and the identification target is a device target or a pedestrian target; if there is an identification target, a corresponding target trajectory is generated according to the video frame in which the identification target is detected; after the identification target is deactivated, a corresponding inbound / outbound event is generated according to the target trajectory, and the inbound / outbound event is sent to the monitoring terminal. In an implementation manner, the description of the specific functions implemented by the intelligent warehousing management system can refer to steps S101-S104 in Embodiment 1 and steps S201-S204 in Embodiment 2, which will not be elaborated here.
[0137] Figure 5 It is a schematic structural diagram of the intelligent warehousing management system according to another embodiment of the present application. As Figure 5As shown in the figure, the intelligent warehousing management system includes: a processor 101 and a memory 102 communicatively connected to the processor 101; the memory 102 stores computer-executable instructions; the processor 101 executes the computer-executable instructions stored in the memory 102 to implement the steps of the intelligent warehousing management method in the above-mentioned method embodiments.
[0138] The intelligent warehousing management system can be independent or a part of a server, and the processor 101 and the memory 102 can use the existing hardware of the server.
[0139] In the above intelligent warehousing management system, the memory 102 and the processor 101 are directly or indirectly electrically connected to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines, such as being connected through a bus. The memory 102 stores computer-executable instructions for implementing the data access control method, including at least one software function module that can be stored in the memory 102 in the form of software or firmware. The processor 101 executes various functional applications and data processing by running the software programs and modules stored in the memory 102.
[0140] The memory 102 can be, but is not limited to, a random access memory (Random Access Memory, abbreviated as RAM), a read-only memory (Read Only Memory, abbreviated as ROM), a programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), an erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), an electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, abbreviated as EEPROM), etc. Among them, the memory 102 is used to store programs, and the processor 101 executes the programs after receiving the execution instructions. Further, the software programs and modules in the above memory 102 may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide a running environment for other software components.
[0141] The processor 101 may be an integrated circuit chip with the ability to process signals. The aforementioned processor 101 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0142] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the steps of the method embodiments of the present application.
[0143] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method embodiments of the present application.
[0144] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0145] Further, it should be noted that although the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0146] It should be understood that the above device embodiments are only illustrative, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0147] In addition, unless otherwise specified, each functional unit / module in the various embodiments of the present application may be integrated into one unit / module, may exist physically as individual units / modules, or may be integrated together with two or more units / modules. The above integrated unit / module may be implemented in the form of hardware or in the form of a software program module.
[0148] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. The technical features of the above embodiments may be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0149] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the appended claims.
[0150] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An intelligent warehousing management method, characterized in that, it includes: Obtain the monitoring video sent by the camera device, and parse to obtain the real-time video frames corresponding to the monitoring video, where the camera device is used to capture video images of the warehousing area; Use a deep learning object detection model to detect the real-time video frames to detect whether there is a recognition target in the real-time video frames, and the recognition target is a device target or a pedestrian target; If there is a recognition target, generate a corresponding target trajectory according to the video frame in which the recognition target is detected; After the recognition target becomes inactive, generate a corresponding inbound / outbound event according to the target trajectory, and send the inbound / outbound event to the monitoring terminal.
2. The method according to claim 1, characterized in that, The step of generating a corresponding target trajectory according to the video frame in which the recognition target is detected specifically includes: Obtain the tracking targets in the current tracking queue, and determine whether there is a first target in the tracking targets that matches the recognition target; If there is, merge and fuse the recognition target with the first target, and generate a corresponding target trajectory according to the target video frame in which the first target is detected; If not, use the recognition target as a new tracking target and put it into the current tracking queue for tracking.
3. The method according to claim 2, characterized in that, The step of generating a corresponding target trajectory according to the target video frame in which the first target is detected specifically includes: Determine each target video frame in which the first target is detected, and determine the trajectory points corresponding to the first target in each target video frame; Connect the trajectory points corresponding to the first target in each target video frame in the order of generation of the target video frames to generate the target trajectory corresponding to the first target.
4. The method according to claim 3, characterized in that, The step of determining whether there is a first target in the tracking targets that matches the recognition target specifically includes: For each tracking target, determine the first video frame closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first trajectory point of the tracking target in the first video frame and the second trajectory point of the recognition target in the real-time video frame; determine whether the distance between the first trajectory point and the second trajectory point is less than the distance threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target; Or, For each tracking target, determine the first video frame closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first feature of the tracking target in the first video frame and the second feature of the recognition target in the real-time video frame; determine whether the feature similarity between the first feature and the second feature is greater than the similarity threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target; Or, For each tracking target, determine the first video frame that is closest in time to the real-time video frame among the video frames in which the tracking target is detected; obtain the first image of the tracking target in the first video frame and the second image of the recognition target in the real-time video frame; determine whether the image overlap degree between the first image and the second image is greater than the overlap degree threshold; if so, the tracking target is the first target that matches the recognition target; if not, the tracking target does not match the recognition target.
5. The method according to any one of claims 1-4, wherein, after the recognition target is deactivated, generating a corresponding inbound / outbound event according to the target trajectory specifically includes: determining the latest generated target trajectory point in the target trajectory and determining the generation time of the video frame corresponding to the target trajectory point; determining the duration between the generation time of the video frame and the current time and determining whether the duration is greater than the duration threshold; if not, the recognition target is not deactivated, and continue to execute the step of parsing to obtain the real-time video frame corresponding to the surveillance video; if so, the recognition target is deactivated, and generate a corresponding inbound / outbound event according to the target trajectory.
6. The method according to claim 4, wherein, generating a corresponding inbound / outbound event according to the target trajectory specifically includes: determining the position of the warehouse door in the storage area and determining the event type according to the warehouse door position and the target trajectory, the event type including an outbound event and an inbound event; determining the target type of the recognition target and determining the cargo information corresponding to the recognition target according to the target type, the target type including an equipment type and a pedestrian type; generating a corresponding inbound / outbound event according to one or more of the recognition target, the event type, the cargo information, and the target trajectory.
7. The method according to claim 6, wherein, determining the cargo information corresponding to the recognition target according to the target type specifically includes: if the target type of the recognition target is the equipment type, use the deep learning recognition model for the number of cargo pallets to recognize the video frame in which the recognition target is detected to determine the number of cargo pallets corresponding to the recognition target, and determine the cargo information corresponding to the recognition target according to the number of cargo pallets; if the target type of the recognition target is the pedestrian type, use the deep learning recognition model for a person carrying goods to recognize the video frame in which the recognition target is detected to determine whether the recognition target is carrying goods, and determine the cargo information corresponding to the recognition target according to whether the recognition target is carrying goods; wherein, the deep learning recognition model for the number of cargo pallets is trained with a plurality of first images as samples, and the first image is an image of an equipment carrying different numbers of cargo pallets; the deep learning recognition model for a person carrying goods is trained with a plurality of second images as samples, and the second image is an image of a person carrying goods and a person not carrying goods.
8. An intelligent warehouse management system, wherein, including: An acquisition module, configured to acquire a surveillance video sent by a camera device, and parse to obtain a real-time video frame corresponding to the surveillance video, where the camera device is configured to capture video images of a storage area; A processing module, configured to detect the real-time video frame by using a deep learning object detection model to detect whether there is an identification target in the real-time video frame, where the identification target is a device target or a pedestrian target; if there is an identification target, generate a corresponding target trajectory according to the video frame in which the identification target is detected; After the identification target is deactivated, generate a corresponding inbound / outbound event according to the target trajectory, and send the inbound / outbound event to a monitoring terminal.
9. An intelligent warehouse management system, characterized in that it includes a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.