Pallet position state detection method, device, equipment and medium
By capturing video with monitoring equipment and utilizing computer vision technology, the status of storage locations can be automatically detected, solving the problems of low efficiency and high cost in existing technologies and achieving efficient monitoring of storage location status.
Patent Information
- Application Number
- CN202111101884.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-09-18
AI Technical Summary
Existing storage location detection technologies are inefficient when detecting items without RFID tags and empty storage locations, and the redundancy of visual sensors increases costs and makes deployment cumbersome.
By capturing surveillance video through monitoring equipment, obtaining cargo location identification information, cropping and correcting cargo location images, extracting features to determine cargo location status, and using computer vision technology to achieve automatic detection of cargo location status.
It reduces the cost of cargo location status detection, improves detection efficiency, and enables real-time monitoring of cargo location status within a spatial area.
Smart Images

Figure CN115841640B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment and medium for detecting the status of cargo locations. Background Technology
[0002] In the process of information management in a factory, monitoring the storage status of goods in the factory is an essential step. Current storage location monitoring technology can utilize Radio Frequency Identification (RFID) technology. RFID can identify items with RFID tags in a storage location; if an item with an RFID tag is identified in a storage location, it indicates that the item is placed there. However, when some storage locations in a factory contain items without RFID tags, and others are empty, RFID technology cannot detect the storage status of these locations. This necessitates using alternative methods to monitor these locations, or manual methods to obtain the storage status, resulting in low storage location monitoring efficiency.
[0003] In addition, current location detection technology can also use vision sensors. By installing vision sensors on each location in the factory, one vision sensor corresponds to one location. The redundant video sensors increase the cost of location detection. Furthermore, when the location of any location in the factory changes, the vision sensor corresponding to that location also needs to be redeployed, which is cumbersome and results in low efficiency of either not being detected or not being detected. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for detecting the status of cargo locations, which can reduce the cost of detecting cargo location status and improve the efficiency of detecting cargo location status.
[0005] One embodiment of this application provides a method for detecting the status of a cargo location, including:
[0006] Obtain the target video frame from the surveillance video captured by the monitoring equipment, and obtain the cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the surveillance video, where M is a positive integer;
[0007] Based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, M cargo location cropped images are obtained from the target video frame. The positions of the cargo locations contained in the M cargo location cropped images are corrected to obtain M cargo location area images.
[0008] Obtain the features of the M storage locations corresponding to each storage location area image, and determine the target storage location status of the M storage locations in the target video frame based on the storage location area features; the target storage location status is used to indicate the occupancy status of the items in the M storage locations.
[0009] One embodiment of this application provides a cargo location status detection device, including:
[0010] The first acquisition module is used to acquire target video frames from the surveillance video captured by the monitoring equipment and to acquire the cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the surveillance video, where M is a positive integer;
[0011] The cargo location cropping module is used to obtain M cargo location cropping images from the target video frame based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, and to perform position correction on the cargo locations contained in the M cargo location cropping images to obtain M cargo location area images.
[0012] The storage location detection module is used to acquire the storage location area features corresponding to M storage location area images, and determine the target storage location status corresponding to the M storage locations in the target video frame based on the storage location area features; the target storage location status is used to indicate the occupancy status of the items in the M storage locations.
[0013] The status of the target storage location includes its occupied status;
[0014] The device also includes:
[0015] The location pair determination module is used to determine two locations as a target location pair if there are two adjacent locations among M locations, both of which are occupied.
[0016] The item detection module is used to acquire the image of the area to be detected containing the target location pair in the target video frame, perform item detection on the image of the area to be detected, and obtain the item detection result corresponding to the image of the area to be detected.
[0017] The anomaly alert module is used to generate anomaly alert information for the target storage location pair when the item detection result is an item overlap result; the item overlap result is used to indicate that the target storage location pair is occupied by the same item.
[0018] The item detection module includes:
[0019] The region cropping unit is used to determine the cropping region associated with the target location pair in the target video frame based on the vertex pixel coordinates corresponding to the target location pair, and to determine the pixels covered by the cropping region as the region image to be detected.
[0020] The object shape determination unit is used to acquire the object edge features in the image of the region to be detected, and to determine the object edge shape in the image of the region to be detected based on the object edge features;
[0021] The storage location shape determination unit is used to determine the edge shape of each storage location in the target storage location pair based on the vertex pixel coordinates of each storage location in the target storage location pair.
[0022] The item detection result determination unit is used to determine the item detection result corresponding to the image of the area to be detected as the item overlap result when the edge shape of the item overlaps with the edge shape of each storage location in the target storage location pair.
[0023] The first acquisition module includes:
[0024] The frame segmentation processing unit is used to acquire the surveillance video captured by the monitoring equipment, perform frame segmentation processing on the surveillance video, and obtain the video frame sequence corresponding to the surveillance video.
[0025] The background template acquisition unit is used to acquire N background templates corresponding to the monitoring device, and the applicable time ranges for each of the N background templates; the N background templates are obtained by modeling the background of images captured by the monitoring device in different time periods, where N is a positive integer;
[0026] The video frame extraction unit is used to determine the i-th video frame as the target video frame when the shooting time of the i-th video frame falls within the applicable time range corresponding to the j-th background template, and the i-th video frame is inconsistent with the j-th background template. The i-th video frame belongs to the video frame sequence, and the j-th background template belongs to N background templates. i is a positive integer less than or equal to the number of video frames in the video frame sequence, and j is a positive integer less than or equal to N.
[0027] The storage location trimming module includes:
[0028] The rectangular region determination unit is used to obtain the vertex pixel coordinates of the kth storage location among M storage locations in the storage location identification information, and determine the target rectangular region in the target video frame based on the vertex pixel coordinates of the kth storage location; k is a positive integer less than or equal to M.
[0029] The video frame segmentation unit is used to segment the target video frame according to the target rectangular region to obtain the k-th location cropped image containing the pixels covered by the target rectangular region; the k-th location cropped image belongs to M location cropped images;
[0030] The position correction unit is used to determine the position transformation matrix corresponding to the cropped image of the kth storage location based on the vertex pixel coordinates corresponding to the kth storage location, and to perform position correction on the cropped image of the kth storage location based on the position transformation matrix to obtain the storage location area image corresponding to the kth storage location.
[0031] The device also includes:
[0032] The second acquisition module is used to acquire the physical space coordinate information corresponding to the kth storage location, as well as the device extrinsic and intrinsic parameters corresponding to the monitoring device. The device extrinsic parameters are used to characterize the position of the monitoring device in physical space, and the device intrinsic parameters are determined by the optical center and the focal length of the monitoring device.
[0033] The vertex coordinate determination module is used to determine the vertex pixel coordinates corresponding to the kth storage location based on physical space coordinate information, equipment external parameters, and equipment internal parameters.
[0034] The cargo location identification generation module is used to generate cargo location identification information corresponding to the monitoring device based on the correspondence between the vertex pixel coordinates corresponding to the kth cargo location and the kth cargo location.
[0035] The storage location detection module includes:
[0036] The first feature extraction unit is used to divide the k-th storage location image among M storage location area images into L image blocks, and obtain the local region features corresponding to the L image blocks respectively; k is a positive integer less than or equal to M, and L is a positive integer;
[0037] The feature concatenation unit is used to concatenate the local region features corresponding to L image blocks to obtain the cargo location region features corresponding to the kth cargo location region image.
[0038] The first classification unit is used to input the features of the storage area corresponding to the kth storage area image into the first classifier, and to identify the features of the storage area in the kth storage area image in the first classifier to obtain the target storage location status corresponding to the storage location in the kth storage area image.
[0039] The first feature extraction unit includes:
[0040] The image segmentation subunit is used to divide the pixels contained in the kth storage area image among M storage area images to obtain an image unit set; each image unit in the image unit set contains the same number of pixels;
[0041] The gradient histogram statistics subunit is used to calculate the gradient histogram for each image unit in the image unit set based on the gradient of the pixels contained in each image unit.
[0042] The local feature acquisition subunit is used to combine image units in the image unit set into image blocks to obtain L image blocks. The gradient histograms corresponding to the image units in the d-th image block are concatenated to obtain the local region features corresponding to the d-th image block. The d-th image block belongs to L image blocks, where d is a positive integer less than or equal to L.
[0043] The first classification unit includes:
[0044] The prediction distance determination subunit is used to input the storage area features corresponding to the kth storage area image into the first classifier, and determine the prediction distance between the classification hyperplane and the storage area features corresponding to the kth storage area image in the first classifier.
[0045] The location status determination subunit is used to determine the target location status of the location in the k-th location area image as occupied if the prediction distance is greater than the classification threshold.
[0046] The aforementioned cargo location status determination subunit is also used to determine the target cargo location status corresponding to the cargo location in the k-th cargo location area image as an idle state if the prediction distance is less than the classification threshold.
[0047] The storage location detection module includes:
[0048] The second feature extraction unit is used to input the kth storage area image from the M storage area images into the image recognition model, and to obtain the storage area features corresponding to the kth storage area image in the image recognition model.
[0049] The matching degree acquisition unit is used to obtain the first matching degree between the occupied attribute feature and the storage area feature corresponding to the image of the kth storage area, and the second matching degree between the idle attribute feature and the storage area feature corresponding to the image of the kth storage area, through the second classifier in the image recognition model.
[0050] The second classification unit is used to determine the target storage location status of the storage location in the kth storage location area image as occupied if the first matching degree is greater than the second matching degree.
[0051] The second classification unit is further configured to determine that the target storage location corresponding to the storage location in the kth storage location area image is in an idle state if the first matching degree is less than the second matching degree.
[0052] The second feature extraction unit includes:
[0053] The convolutional subunit is used to input the kth storage location image from M storage location images into the image recognition model. Based on the convolutional layer in the image recognition model, the kth storage location image is convolved to obtain the storage location convolutional features.
[0054] The residual subunit is used to perform residual convolution processing on the location convolution features based on the residual layer in the image recognition model to obtain the location residual features;
[0055] The location feature acquisition unit is used to generate the location area features corresponding to the k-th location area image based on the location convolution features and location residual features.
[0056] The device also includes:
[0057] The location status update module is used to update the historical location status in the management system to the target location status corresponding to the kth location when the target location status corresponding to the kth location out of M locations is inconsistent with the historical location status corresponding to the kth location in the management system. The historical location status refers to the location status determined based on historical video frames. The shooting time of the historical video frames is earlier than the shooting time of the target video frames, and k is a positive integer less than or equal to M.
[0058] The device also includes:
[0059] The training sample acquisition module is used to acquire positive sample region images and negative sample region images from historical surveillance videos captured by the monitoring equipment; the positive sample region images include occupied labels, and the negative sample region images include idle labels.
[0060] The negative sample feature extraction module is used to obtain the positive sample region features corresponding to the positive sample region image and the negative sample region features corresponding to the negative sample region image.
[0061] The positive sample feature extraction module is used to determine the first hyperplane in the first classifier based on the positive sample region features and occupied labels, and to determine the second hyperplane in the first classifier based on the negative sample region features and idle labels; the first hyperplane is parallel to the second hyperplane.
[0062] The hyperplane optimization module is used to determine the classification hyperplane in the first classifier based on the first hyperplane and the second hyperplane when the interval distance between the first hyperplane and the second hyperplane is optimized to the maximum value; the distance between the classification hyperplane and the first hyperplane is equal to the distance between the classification hyperplane and the second hyperplane.
[0063] One aspect of this application provides a computer device, including a memory and a processor. The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method provided in one aspect of this application.
[0064] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having a processor performs the method provided in one aspect of this application.
[0065] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the above aspect.
[0066] This application embodiment can capture surveillance video using monitoring equipment and obtain target video frames from the surveillance video for detecting the status of cargo locations. Based on the vertex pixel coordinates of each cargo location in the surveillance video, M cargo location cropped images can be cropped from the target video frames. By performing position correction on the M cargo location cropped images, M cargo location area images can be obtained. Then, feature extraction can be performed on each cargo location area image to obtain the cargo location area features corresponding to each cargo location area image. By classifying the cargo location area features, the status of the target cargo locations corresponding to the M cargo locations in the target video frame can be determined. As can be seen, monitoring equipment can collect real-time video of the entire space area. Based on the pre-set cargo location identification information, the corresponding cargo location cropped image is extracted from the video frames of the monitoring video. By performing a series of operations such as position correction, feature extraction, and feature classification on the cargo location cropped image, the cargo location status corresponding to each cargo location is obtained. In other words, the monitoring equipment can realize the monitoring of cargo locations in the entire space area, which can reduce the cost of cargo location status detection. The monitoring video captured by the monitoring equipment can detect the cargo location status of each cargo location in the space area in real time, which can improve the efficiency of cargo location status detection. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 This is a schematic diagram of the structure of a cargo location status detection system provided in an embodiment of this application;
[0069] Figure 2 This is a schematic diagram of a cargo location status detection scenario provided in an embodiment of this application;
[0070] Figure 3 This is a flowchart illustrating a method for detecting the status of a cargo location provided in an embodiment of this application;
[0071] Figure 4 This is a schematic diagram illustrating the status of a target storage location that integrates multiple monitoring devices, as provided in an embodiment of this application.
[0072] Figure 5 This is a schematic diagram illustrating an update of the storage location status provided in an embodiment of this application;
[0073] Figure 6 This is a flowchart illustrating a method for detecting the status of a cargo location provided in an embodiment of this application;
[0074] Figure 7 This is a schematic diagram of a cargo location position correction for a cargo location cropping image provided in an embodiment of this application;
[0075] Figure 8 This is a flowchart illustrating the training process of a classifier provided in an embodiment of this application;
[0076] Figure 9 This is a schematic diagram of a cargo location status detection process provided in an embodiment of this application;
[0077] Figure 10 This is a schematic diagram of an article detection method provided in an embodiment of this application;
[0078] Figure 11 This is a schematic diagram of the structure of a cargo location status detection device provided in an embodiment of this application;
[0079] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0080] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0081] This application relates to Computer Vision (CV) technology. Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to using cameras and computers to replace human eyes for target recognition, measurement, and other machine vision tasks, and further performing image processing to make the computer-processed images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to establish artificial intelligence systems capable of extracting information from images or multi-dimensional data. This application specifically relates to image recognition technology under computer vision. It can use a camera to capture overhead surveillance video of a factory, extract images of storage areas from the video frames, and perform image recognition on the storage area images to identify the status of the storage locations. The storage location status can include occupied and idle states, where occupied state indicates that the storage location in the storage area image is occupied by an item (or goods), and idle state indicates that the storage location in the storage area image is not occupied by an item (or storage location).
[0082] Please see Figure 1 , Figure 1 This is a schematic diagram of a cargo location status detection system provided in an embodiment of this application. Figure 1 As shown, the storage location status detection system may include monitoring device 10a, server 10c, and user terminal 10d. Monitoring device 10a can refer to a device used to capture images of the monitored area 10b. The monitoring device 10a may include ordinary cameras, high-pole cameras, LCD wall-mounted cameras, etc. The monitored area 10b may include factory areas, material warehouses, factory workshops, etc. The number of monitoring devices in the storage location status detection system can be one or more. These monitoring devices are deployed at some specific locations in the physical space area (e.g., monitored area 10b) to simultaneously capture video content from different angles within that space area. The video content collected by the monitoring devices is called the monitoring video associated with monitored area 10b. The surveillance video captured by the monitoring equipment can be transmitted to server 10c. Server 10c can perform image processing on the surveillance video captured by monitoring equipment 10a, identifying the status of each storage location in a single video frame, that is, whether each storage location in a single video frame is occupied by an item. If a storage location is found to be occupied by an item, it indicates that the storage location status is occupied; if a storage location is found to be empty, it indicates that the storage location status is vacant. Server 10c can associate and store each storage location and its corresponding storage location status. When the administrator of monitoring area 10b can query the storage location status of all storage locations in monitoring area 10b through user terminal 10d, they can monitor the storage location occupancy status of monitoring area 10b in real time. Figure 1The cargo location status detection system shown can remotely monitor the monitoring area 10b, which can improve the speed of acquiring cargo location status; the number of user terminals in the cargo location status detection system can be one or more, and there is no limit to the number of user terminals here.
[0083] Server 10c can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. User terminal 10d can include: smartphones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches, smart bracelets, etc.), and smart TVs, etc., all intelligent terminals with data query functions. Figure 1 As shown, the monitoring device 10a can interact with the server 10c, and the server 10c can also interact with the user terminal 10d. Optionally, when the user terminal 10d in the storage location status detection system integrates image processing functions, the monitoring device 10a can also directly transmit the captured monitoring video to the user terminal 10d, whereby the user terminal 10d performs image processing on the monitoring video to obtain the storage location status corresponding to each storage location in the video frame.
[0084] Please see Figure 2 , Figure 2 This is a schematic diagram of a storage location status detection scenario provided in an embodiment of this application. Taking a factory as an example, this embodiment describes the storage location status detection process in a factory area. To enable real-time monitoring of the factory, implement remote control, and combine information management to comprehensively understand data during the manufacturing process, one or more cameras can be configured in the factory area. These cameras can monitor the storage location status from different directions. It should be noted that some locations for placing items can be pre-set in the factory area; these can be referred to as storage locations. Some storage locations in the factory area are already occupied, while others are vacant.
[0085] like Figure 2The factory area 20c shown may include workers, production equipment 20j, and some pre-set storage locations (e.g., storage location 1, storage location 2, storage location 3, ..., storage location 13, etc.). This factory area 20c may also be equipped with cameras 20a and 20b, which can monitor the status of the storage locations within the factory area 20c from two different directions. Using two cameras simultaneously to capture the storage location status in the factory area 20c can solve the problem of storage locations obstructing each other. Each camera can be responsible for a sub-area within the factory area 20c, and a sub-area may include one or more storage locations. After cameras 20a and 20b capture the monitoring video of their respective sub-areas, they can transmit the captured video to a server (as described above). Figure 1 The server 10c shown can collect surveillance videos of the factory area 20c through cameras 20a and 20b, such as surveillance video 20d of the sub-area under the responsibility of camera 20a, and surveillance video 20e of the sub-area under the responsibility of camera 20b. Surveillance videos 20d and 20e are videos of the same duration.
[0086] After acquiring surveillance videos 20d and 20e, the server needs to synchronize them to ensure that the status of all storage locations within factory area 20c is detected at the same time, which is beneficial for the effectiveness of factory storage location management. The server can use the same method to perform frame segmentation processing on surveillance videos 20d and 20e, obtaining video frame sequence A corresponding to surveillance video 20d and video frame sequence B corresponding to surveillance video 20e. The number of video frames in video frame sequence A and video frame sequence B is the same, and the video frames in video frame sequence A correspond one-to-one with the video frames in video frame sequence B in chronological order. When the second video frame in video frame sequence A is selected to detect the storage location status, the second video frame in video frame sequence B will also be selected to detect the storage location status. By combining the storage location status detection results of the second video frames in the above two video frame sequences, the storage location status of all storage locations in factory area 20c at the same time (the time when the above cameras 20a and 20b respectively capture the second video frame) can be obtained.
[0087] It should be noted that the location status detection process is the same for any video frame in video frame sequence A and video frame sequence B. The following description uses any video frame 20f in monitoring video 20e as an example to illustrate the location status detection process for video frame 20f. Since camera 20b may only be responsible for monitoring a sub-area of factory area 20c, video frame 20f in monitoring video 20e captured by camera 20b only contains some locations in factory area 20c. For example, video frame 20f may include locations 5, 6, 7, and locations 11 to 13. The server can extract the location area image corresponding to a single location from video frame 20f based on the configured location identification information. This location identification information can refer to the vertex pixel coordinates corresponding to each location in the monitoring video captured by camera 20b. Since the position of camera 20b in factory area 20c is fixed, the position of each storage location in each video frame of monitoring video 20e also remains unchanged. That is, for camera 20b, the storage location identification information only needs to be configured once, and there is no need to configure the storage location identification information for each video frame in the monitoring video. Of course, if the position of camera 20b in factory area 20c is adjusted, the sub-area it is responsible for shooting changes, and then the storage location identification information needs to be reconfigured for camera 20b.
[0088] like Figure 2As shown, the server can obtain the vertex pixel coordinates of the storage location 10 from the storage location identification information corresponding to the camera 20b. Based on the vertex pixel coordinates corresponding to the storage location 10, the server can extract the region image 20g containing the storage location 10 from the video frame 20f. Then, the server can perform feature extraction on the region image 20g to obtain feature 1 in the region image 20g. By classifying feature 1, the storage location status detection result of the region image 20g is obtained as follows: the storage location is empty, that is, the storage location status of the region image 20g is idle. Similarly, the server can obtain the vertex pixel coordinates of storage location 6 from the storage location identification information corresponding to camera 20b. Based on the vertex pixel coordinates of storage location 6, the server can extract the region image 20h containing storage location 10 from video frame 20f. Then, feature extraction can be performed on region image 20h to obtain feature 2. By classifying feature 2, the storage location status detection result of region image 20h is obtained as: the storage location is occupied, i.e., the storage location status of region image 20h is occupied. Based on the same operation, the server can obtain the storage location status detection result corresponding to each storage location in video frame 20f. Simultaneously, the server can also obtain associated video frames from monitoring video 20d that have the same shooting time as video frame 20f, and based on the same operation as video frame 20f, obtain the storage location status detection result corresponding to each storage location in the associated video frames. By combining the storage location status detection results corresponding to each storage location in video frame 20f with the storage location status detection results corresponding to each storage location in the associated video frames, the global storage location status detection results in factory area 20c can be obtained. The global storage location status detection results are shown in area 20i. The storage location status detection results of storage locations 1, 2, 5, 8, 10, 11, and 13 in factory area 20c are all empty, while the storage location status detection results of storage locations 3, 4, 6, 7, 9, and 12 are all occupied.
[0089] Optionally, the server can store the location status detection results for each storage location in factory area 20c. When the location status detection results in subsequent video frames of monitoring videos 20d and 20e differ from the stored results, the stored results can be updated. Using these two cameras, the factory area can be monitored in real time. By processing the video frames in the monitoring video, the location status in the factory area can be obtained in real time, reducing both the cost and efficiency of location status detection.
[0090] Please see Figure 3 , Figure 3This is a schematic flowchart illustrating a method for detecting the status of a storage location provided in an embodiment of this application. It can be understood that this method for detecting the status of a storage location can be executed by a computer device, which can be a server (e.g., the one described above). Figure 1 The server 10c shown is either a user terminal (e.g., the one mentioned above). Figure 1 The user terminal 10d shown may be a computing application (including program code), but this is not specifically limited. Figure 3 As shown, the method for detecting the status of a cargo location may include the following steps S101-S103:
[0091] Step S101: Obtain the target video frame from the monitoring video captured by the monitoring equipment, and obtain the cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the monitoring video, where M is a positive integer.
[0092] Specifically, when a user wants to monitor the real-time status of cargo locations within a certain spatial area, one or more monitoring devices can be configured in that area (e.g., the aforementioned...). Figure 2 The cameras 20a and 20b in the corresponding embodiments can monitor the space area in real time through one or more configured monitoring devices. If a single monitoring device can monitor all storage locations in the entire space area without any obstruction, then only one monitoring device is needed for the entire space area. If a single monitoring device monitors the space area and some storage locations are obstructed, multiple monitoring devices can be configured to monitor the storage locations in the space area simultaneously from different directions. One monitoring device can be responsible for monitoring a sub-area of the space area and can ensure that the storage locations in the sub-area are not obstructed. The aforementioned space area can be a factory area, storage warehouse, factory production area, etc.
[0093] Computer equipment can acquire surveillance video of a spatial area through one or more monitoring devices, and obtain surveillance video captured by one or more monitoring devices respectively (e.g., the above). Figure 2(As shown in the corresponding embodiments, monitoring video 20e and monitoring video 20d). Alternatively, one or more monitoring videos configured in the spatial area can transmit the monitoring video of their respective responsible sub-areas to the computer device, thereby allowing the computer device to acquire the monitoring videos transmitted by one or more monitoring devices. When there is only one monitoring device, the computer device can acquire the monitoring video captured by that monitoring device and perform cargo location status detection on the video frames in that monitoring video. When there are multiple monitoring devices, the computer device can acquire the monitoring videos captured by multiple monitoring devices respectively, and can synchronously perform cargo location status detection on the video frames in the monitoring videos captured by multiple monitoring devices. For example, it can acquire video frames with the same shooting time from the monitoring videos captured by multiple monitoring devices respectively, and perform the same cargo location status detection on multiple video frames with the same shooting time. That is, the processing procedure of the computer device for the monitoring video captured by each monitoring device is the same. For ease of understanding, the following description uses the monitoring video captured by one monitoring device as an example.
[0094] After acquiring the surveillance video captured by the monitoring equipment, the computer equipment can perform frame-by-frame processing to obtain the corresponding video frame sequence. In other words, the surveillance video can be divided into individual images (i.e., video frames), and a video frame sequence can be generated according to the temporal order of each video frame in the surveillance video. The computer equipment can select target video frames from the surveillance video at a fixed acquisition frequency for performing cargo location status detection. For example, cargo location status detection can be performed once every 6 video frames in the video frame sequence, that is, a video frame is selected as the target video frame every 5 video frames; or, a video frame can be extracted from the video frame sequence every 1 second as the target video frame for performing cargo location status detection, etc.
[0095] Optionally, the computer device can also acquire N background templates corresponding to the monitoring device, and the applicable time ranges corresponding to each of the N background templates. These N background templates are obtained by modeling the backgrounds of images captured by the monitoring device in different time periods, where N is a positive integer, such as 1, 2, ... . When the capture time of the i-th video frame falls within the applicable time range corresponding to the j-th background template, and the i-th video frame is inconsistent with the j-th background template, the i-th video frame is determined as the target video frame. The i-th video frame belongs to the video frame sequence, and the j-th background template belongs to the N background templates, where i is a positive integer less than or equal to the number of video frames in the video frame sequence, and j is a positive integer less than or equal to N. The computer device can use low-complexity algorithms such as background modeling algorithms as pre-judgments to select the target video frame from the video frame sequence. Computer equipment can use background modeling algorithms to model the background of videos captured by surveillance equipment. Since lighting changes throughout the 24 hours of a day, different background templates can be built for different time periods. For example, a background template can be built every hour, in which case the number of background templates N is 24; or a background template can be built every two hours, in which case the number of background templates N is 12. The computer equipment can sequentially detect whether there are differences between each video frame in the video frame sequence and its corresponding background template. For example, for any video frame in the video frame sequence (such as the i-th video frame), the shooting time of the i-th video frame can be obtained. When the time frame falls within the applicable time range of the j-th background template, such as the shooting time of the i-th video frame being 13:23:20 on September 1, 20xx, and the applicable time range of the j-th background template being from 13:00:00 to 15:59:59, the i-th video frame can be compared with the j-th background template. If there is a difference between the i-th video frame and the j-th background template, the i-th video frame can be identified as the target video frame. If the i-th video frame is consistent with the j-th background template, the next video frame (such as the (i+1)-th video frame) can be judged. Any video frame in the video frame sequence that differs from its corresponding background template can be identified as the target video frame. The background modeling algorithm used in this application may include, but is not limited to, the average background modeling algorithm, the Gaussian background modeling algorithm, and non-parametric background. This application does not limit the background modeling algorithm used.
[0096] It should be noted that the above methods of selecting target video frames from a video frame sequence based on acquisition frequency or using background templates are merely examples. Other methods can also be used to select target video frames from a video frame sequence. For example, each video frame in the video frame sequence can be used to perform cargo location status detection, meaning that each video frame can be used as a target video frame. This application does not limit the selection process of target video frames.
[0097] To determine the location of a cargo location contained in a target video frame, cargo location identification information associated with the monitoring equipment can be obtained. This cargo location identification information can include the vertex pixel coordinates of M cargo locations in the monitoring video within the sub-area monitored by the monitoring equipment. M can be a positive integer, such as 1, 2, ... The locations of the cargo locations and the monitoring equipment in the spatial area are usually fixed. Therefore, the pixel positions (e.g., cargo location vertices) of the cargo locations in the images captured by the monitoring equipment can be marked. The marked positions are the vertex pixel coordinates of the cargo locations. Based on the vertex pixel coordinates of the M cargo locations, the aforementioned cargo location identification information can be combined. In other words, the cargo location identification information can be pre-configured vertex coordinates used to mark the location of the cargo locations. For example, if the cargo location is quadrilateral, the cargo location identification information can include the pixel coordinates of the four vertices corresponding to that cargo location; if the cargo location is triangular, the cargo location identification information can include the pixel coordinates of the three vertices corresponding to that cargo location, etc. This application does not limit the shape of the cargo location. It should be noted that in practical applications, the number of storage locations included in the storage location identification information corresponding to the monitoring device can be less than or equal to the number of storage locations in the target video frame actually captured by the monitoring device. For example, if a storage location in the target video frame is a newly added storage location and its vertex pixel coordinates have not yet been labeled, or if a storage location in the target video frame can be captured by other monitoring devices and the storage location identification information corresponding to the other monitoring devices has already labeled the storage location, then there is no need to label it again in the storage location identification information corresponding to the monitoring device, and there is no need to perform repeated storage location status detection on it subsequently.
[0098] Optionally, the vertex pixel coordinates of the M cargo locations in the real physical space can be calculated using the calibration information of the monitoring equipment. For example, the coordinates of the M cargo locations in the real physical space can be obtained, i.e., physical space coordinates, and then converted into pixel coordinates. For any one of the M cargo locations (e.g., the k-th cargo location, where k is a positive integer less than M), the computer equipment can obtain the physical space coordinates of the k-th cargo location, as well as the extrinsic and intrinsic parameters of the monitoring equipment. The extrinsic parameters characterize the position of the monitoring equipment in the physical space, and the intrinsic parameters can be determined by the optical center and the focal length (distance from the optical center to the image plane) of the monitoring equipment. Then, based on the physical space coordinates, extrinsic and intrinsic parameters, the vertex pixel coordinates of the k-th cargo location can be determined. Based on the correspondence between the vertex pixel coordinates of the k-th cargo location and the k-th cargo location itself, the cargo location identification information corresponding to the monitoring equipment can be generated.
[0099] In one or more embodiments, the physical space coordinate information corresponding to the k-th storage location among the M storage locations can refer to the coordinates of the k-th storage location in the world coordinate system (also known as the measurement coordinate system). The world coordinate system is a three-dimensional rectangular coordinate system, which can be used as a reference to describe the spatial position of the monitoring equipment and the k-th storage location. The position of the world coordinate system can be freely determined according to actual needs. The coordinate transformation process may also involve the camera coordinate system, image coordinate system, and pixel coordinate system corresponding to the monitoring equipment. The camera coordinate system is also a three-dimensional rectangular coordinate system, with the origin located at the optical center of the lens of the monitoring equipment. The horizontal and vertical axes are parallel to the two sides of the image plane, and the vertical axis is the optical axis of the lens, perpendicular to the image plane. The image coordinate system can be a two-dimensional planar coordinate system, parallel to the imaging plane, with the origin at the center of the image (such as the center of the video frame in the monitoring video). The pixel coordinate system is in pixels, with the origin at the upper left corner of the image (such as the upper left corner of the video frame in the monitoring video).
[0100] After the computer equipment obtains the physical space coordinate information corresponding to the k-th storage location, the coordinate transformation process is described using the physical space coordinate information of any vertex P of the k-th storage location as an example; assuming that the physical space coordinate information of vertex P in the world coordinate system is P c Coordinate P c It can be represented as a column vector [X] c Y c Z c ] T , [ ] T Represented as the transpose of a vector, the coordinates of vertex P in the camera coordinate system are Pi. w Coordinate P w It can be represented as a column vector [X] w Y w Z w ] T Coordinate P c and coordinates P w The coordinate transformation between the two coordinates can be achieved using a rotation matrix R and a translation vector t, where the rotation matrix R can be a... The matrix, the translation vector t can be a The matrix represents the rotation and translation along the horizontal, vertical, and axial directions, respectively. The coordinate P... c This can be represented as Pc = P w R+t, specifically as shown in the following formula (1):
[0101] = + (1)
[0102] in, Represented as a rotation matrix R, This is represented as a translation vector t. For ease of calculation, the above formula (1) can be changed to homogeneous coordinate form, as shown in the following formula (2):
[0103] = = (2)
[0104] in, This can be referred to as the extrinsic parameters of the monitoring equipment. Assume the coordinates of vertex P in the image coordinate system (such as the imaging plane) are [x, y]. T Then the coordinates are [x, y]. T With coordinates [X c Y c Z c ] T The relationship can be shown in the following formula (3):
[0105] = (3)
[0106] Where f is the focal length of the monitoring device. Assume the coordinates of vertex P in the pixel coordinate system are [u, v]. T The origin of the pixel coordinate system is at the top left corner of the video frame, with the u-axis pointing horizontally to the left and the v-axis pointing vertically downwards. The origin of the image coordinate system is at the center of the image. There is a translation (c) between the image coordinate system and the pixel coordinate system. x , c y Assuming that the length of a pixel in a video frame along the horizontal axis (e.g., the x-axis) of the image coordinate system is α, and the length along the vertical axis (e.g., the y-axis) is β, then when vertex P is transformed from the image coordinate system to the pixel coordinate system, it is scaled by a factor of α and translated by a factor of c along the horizontal axis. x Scale by β times and translate by c in the horizontal direction y The conversion process can be shown in the following formula (4):
[0107] = (4)
[0108] Among them, f x =αf,f y =βf,f x This can be referred to as the normalized focal length on the u-axis of the pixel coordinate system, f. y This can be referred to as the normalized focal length on the v-axis of the pixel coordinate system. This can be referred to as the device intrinsic parameters of the monitoring equipment. According to the above formulas (2) to (4), the process of vertex P transforming from the world coordinate system to the pixel coordinate system can be determined as shown in the following formula (5):
[0109] (5)
[0110] Wherein, the coordinates of vertex P in the pixel coordinate system are [u, v] T That is, the vertex pixel coordinates of the k-th storage location P in the monitoring video. Similarly, the vertex pixel coordinates of all vertices of the M storage locations in the monitoring video can be calculated using the above formula (5). Based on the vertex pixel coordinates of the M storage locations, storage location identification information for the monitoring video can be generated.
[0111] It is understood that the location identification information in this application can be achieved by marking the location of the location in the monitoring video, by performing coordinate transformation on the physical spatial coordinates of the location, or by other methods. This application does not limit the method of determining the location identification information. The coordinate transformation process described above is only an example in the embodiments of this application, and this application should also include other transformation forms besides the coordinate transformation forms described above.
[0112] Step S102: Based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, obtain M cargo location cropped images from the target video frame, perform position correction on the cargo locations contained in the M cargo location cropped images, and obtain M cargo location area images.
[0113] Specifically, the computer equipment can determine the position of each cargo location in the target video frame based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information. That is, it can locate the position of each cargo location in the target video frame. Based on the position of each cargo location in the target video frame, the target video frame can be cropped to obtain M cargo location cropped images. Since the shooting angle of the monitoring equipment may cause tilting or distortion of the cargo locations in the cropped images, the computer equipment can perform position correction on the cargo locations in the area detection image and adjust the position-corrected area cropped image to the target size, obtaining cargo location area images corresponding to each cargo location, i.e., M cargo location area images, one cargo location corresponding to one cargo location cropped image, and one cargo location cropped image corresponding to one cargo location area image. The methods for position correction of the cargo location cropped images may include, but are not limited to: tilt correction based on projection, image correction based on Hough transform, image correction based on linear fitting, and perspective transformation. This application does not limit the method of position correction.
[0114] For example, suppose the location identification information includes the vertex pixel coordinates of three locations (M can be 3) in the monitoring video, represented as location A, location B, and location C. The vertex pixel coordinates corresponding to location A include the pixel coordinates of vertex 1, vertex 2, vertex 3, and vertex 4. Based on these four vertex pixel coordinates, a rectangular area can be determined. By cropping the target video frame based on this rectangular area, a cropped image of location A can be obtained. Then, by correcting the position of location A in the cropped image, the location area image corresponding to location A can be obtained. The vertex pixel coordinates corresponding to location B include the pixel coordinates of vertex 5, vertex 6, vertex 7, and vertex 8. The vertex pixel coordinates corresponding to location C include the pixel coordinates of vertex 9, vertex 10, vertex 11, and vertex 12. Based on the same processing procedure, the location area images corresponding to location B and location C can be obtained.
[0115] Step S103: Obtain the features of the M storage locations corresponding to the images of the M storage locations respectively, and determine the status of the target storage locations corresponding to the M storage locations in the target video frame based on the features of the storage locations; the status of the target storage locations is used to indicate the occupancy status of the items in the M storage locations.
[0116] Specifically, the computer equipment can acquire images of M storage locations and extract features from each, obtaining the storage location features corresponding to each of the M images. These features can then be classified to determine the status of each of the M storage locations in the target video frame. This status indicates the occupancy of items in the M storage locations within the target video frame. The storage location status includes an occupied status and an idle status. An occupied status indicates that the storage location is occupied by items when the monitoring equipment captures the target video frame, while an idle status indicates that the storage location is empty when the monitoring equipment captures the target video frame. The feature extraction methods used in this application may include, but are not limited to: Scale-invariant feature transform (SIFT) algorithm, Speeded Up Robust Features (SURF) algorithm, Histogram of Oriented Gradient (HGP) algorithm, Difference of Gaussian function, and deep learning algorithms such as convolutional neural network models (e.g., LeNet network model, AlexNet network model, etc.). Of course, this application may also design deep learning models according to the actual needs of the cargo location status detection scenario. The feature classification methods used in this application may include, but are not limited to: AdaBoost algorithm (an iterative algorithm), Support Vector Machine (SVM), neural network algorithm, perceptron algorithm, Softmax classifier, and logistic classifier. This application does not limit the feature extraction methods and classification methods used.
[0117] In one or more embodiments, the feature extraction and feature classification processes of the storage location area images are described using a convolutional neural network model and a softmax classifier as examples. For any one of the M storage location area images (such as the k-th storage location area image), the computer device can input the k-th storage location area image into an image recognition model. Through this image recognition model, the storage location area features corresponding to the k-th storage location area image can be obtained. This image recognition model can be a pre-trained convolutional neural network model, which can be used to extract the storage location area features from the storage location area image. Optionally, when the image recognition model includes convolutional layers and residual layers (e.g., the image recognition model can be a residual convolutional network), the computer device inputs the image of the kth storage location area into the image recognition model. Based on the convolutional layers in the image recognition model, the image of the kth storage location area is subjected to convolutional processing to obtain the storage location convolutional features corresponding to the image of the kth storage location area. Then, based on the residual layers in the image recognition model, the above-mentioned storage location convolutional features are subjected to residual convolutional processing to obtain the storage location residual features. Based on the above-mentioned storage location convolutional features and storage location residual features, the storage location area features corresponding to the image of the kth storage location area can be generated. The image recognition model can contain one or more convolutional layers, and one or more residual layers can be connected after the convolutional layers. The convolutional features of the cargo location can refer to the image features output by the last convolutional layer in the one or more convolutional layers of the image recognition model, and the residual features of the cargo location can refer to the image features output by the last residual layer in the one or more residual layers of the image recognition model. By fusing the convolutional features and the residual features of the cargo location (e.g., feature splicing, feature merging, etc.), the fused cargo location features are obtained. Here, the fused cargo location features can be called the cargo location area features corresponding to the k-th cargo location area image.
[0118] After extracting the features of the storage area corresponding to the k-th storage area image, the second classifier in the image recognition model (e.g., the softmax classifier mentioned above) can be used to obtain the first matching degree between the occupancy attribute feature and the features of the storage area corresponding to the k-th storage area image, and the second matching degree between the idle attribute feature and the features of the storage area corresponding to the k-th storage area image. If the first matching degree is greater than the second matching degree, the target storage area state corresponding to the storage area in the k-th storage area image is determined to be in an occupied state; if the first matching degree is less than the second matching degree, the target storage area state corresponding to the storage area in the k-th storage area image is determined to be in an idle state. The aforementioned occupancy attribute features can refer to attribute features unique to images of storage locations that are in an occupied state, and the vacancy attribute features can refer to attribute features unique to images of storage locations that are in an vacant state. The aforementioned first matching degree and second matching degree can refer to the probability values output by the second classifier for different storage location states. That is, the first matching degree refers to the probability that the storage location feature corresponding to the k-th storage location image belongs to the occupied state, and the second matching degree refers to the probability that the storage location feature corresponding to the k-th storage location image belongs to the vacant state. In the second classifier, the matching degree between the storage area feature corresponding to the k-th storage location image and two attribute features (occupancy attribute feature and freehold attribute feature) can be identified. By comparing the first matching degree and the second matching degree, it can be determined which of the two attribute features is more similar to the storage area feature corresponding to the k-th storage location image. The larger the matching degree, the more similar the storage area feature corresponding to the k-th storage location image is to that attribute feature. For example, if the first matching degree is greater than the second matching degree, it means that the storage area feature corresponding to the k-th storage location image is more similar to the occupancy attribute feature, so the target storage location state corresponding to the k-th storage location image (i.e., the k-th storage location) can be determined to be in an occupied state. If the first matching degree is less than the second matching degree, it means that the storage area feature corresponding to the k-th storage location image is more similar to the freehold attribute feature, so the target storage location state corresponding to the k-th storage location can be determined to be in a free state. Based on the same execution operation, the target storage location states corresponding to the storage locations in M storage location images can be obtained.
[0119] It should be noted that if multiple monitoring devices are used for real-time monitoring simultaneously, the computer equipment needs to process the monitoring videos captured by the multiple monitoring devices synchronously. For example, when monitoring devices 1, 2, and 3 are used to monitor a certain spatial area at the same time, the monitoring videos captured by monitoring devices 1, 2, and 3 have the same video duration. The computer equipment can perform the same processing on the monitoring videos captured by the above three monitoring devices, such as video frame segmentation, storage location cropping, storage location position correction, and storage location status detection. When the first video frame in the monitoring video captured by a monitoring device is selected as the target video frame, the first video frame in the monitoring video captured by monitoring device 2 and the first video frame in the monitoring video captured by monitoring device 3 can both be used as target video frames. By integrating the target storage location status detected in the target video frames of each monitoring video, the target storage location status of all storage locations in the spatial area at the same time can be obtained. Optionally, the computer equipment can process surveillance videos captured by multiple monitoring devices in parallel, and can also process M storage location area images extracted from the same target video frame in parallel. By adopting parallel processing, the storage location status detection speed for surveillance videos can be improved.
[0120] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating the status of a target storage location integrating multiple monitoring devices, as provided in an embodiment of this application. Figure 4 As shown, taking a factory production workshop as an example, two monitoring devices can be configured in the factory production workshop to simultaneously capture images from two different angles. These two monitoring devices are monitoring device 30a and monitoring device 30c. In order to capture each storage location in the factory production area more clearly, the storage locations to be monitored by the two monitoring devices can be pre-set. For example, for two storage locations with an adjacent relationship (such as storage location 12 and storage location 13), when storage location 12 is occupied by a large number of items and the items in storage location 12 are relatively tall, storage location 13 may be blocked by the items in storage location 12. Therefore, different monitoring devices can be used to monitor storage locations with an adjacent relationship. That is, the two monitoring devices can cross-monitor storage locations with an adjacent relationship. Therefore, a surveillance video captured by a monitoring device may contain all or part of the storage locations in the factory's production workshop. When performing storage location status detection on the target video frame captured by the monitoring device, it is only necessary to detect part of the storage locations contained in the target video frame. Of course, the storage location identification information corresponding to the monitoring device may also only contain the vertex pixel coordinates of this part of the storage locations.
[0121] like Figure 4As shown, the computer device extracts video frame 1 from the monitoring video captured by monitoring device 30a, and performs location status detection on video frame 1 to obtain output result 30b. The monitoring device 30a captures video frame 1 at time a. This video frame 1 contains all the locations in the factory production workshop. Based on the location identification information corresponding to the monitoring device 30a, the location status detection on this video frame 1 can obtain the target location status corresponding to locations 1 to 7, 8, 10, 12, 14, 16, 17, 19, 21, and 23, respectively. The target location status corresponding to these locations is shown in output result 30b, which is the location status of the above locations at time a. Similarly, the computer equipment can extract video frame 2 captured at time a from the monitoring video of the monitoring equipment 30c. Based on the cargo location identification information corresponding to the monitoring equipment 30c, the cargo location status can be detected on the video frame 2 to obtain the target cargo location status corresponding to cargo location 9, cargo location 11, cargo location 13, cargo location 15, cargo location 18, cargo location 20, cargo location 22, and cargo location 24 respectively. The target cargo location status corresponding to these cargo locations is shown in the output result 30d.
[0122] Furthermore, after the computer equipment obtains the location status output results of each monitoring device at time a (i.e., output results 30b and 30d), it can integrate output results 30b and 30d to obtain the location detection result 30e of the entire factory production workshop. The location detection result 30e includes the target location status corresponding to each location in the factory production workshop.
[0123] Optionally, different identification information can be used in the location inspection result 30e to represent different target location statuses, such as... Figure 4 As shown, text is used as identification information to distinguish between idle and occupied states; or a solid dot can be used to represent the occupied state and a hollow circle to represent the idle state; or a square can be used to represent the occupied state and a triangle to represent the idle state, etc.; this application does not limit the identification information for the status of the storage space.
[0124] Optionally, after acquiring the target storage location status of all storage locations in the spatial area at the same time (the capture time of the target video frame), the computer equipment can compare the target storage location status with the historical storage location status stored in the management system. For storage locations where the historical and target storage location statuses are inconsistent, the storage location status is updated to match the target storage location status. If a storage location's historical storage location status is "idle" but its current target storage location status in the target video frame is "occupied," then the historical storage location status for that location needs to be updated in the management system to show it as "occupied." The aforementioned management system can refer to a system used to manage information such as storage location status, equipment usage, and worker operation of equipment in the aforementioned spatial area. The information stored in this management system can be updated in real time, which is beneficial for managers to monitor the situation in the spatial area in real time. For example, when the target storage location status corresponding to the kth storage location among the M storage locations mentioned above is inconsistent with the historical storage location status corresponding to the kth storage location in the management system, the historical storage location status can be updated to the target storage location status corresponding to the kth storage location in the management system. The historical storage location status refers to the storage location status determined based on historical video frames. The shooting time of the historical video frames is earlier than the shooting time of the target video frames. When the target storage location status of the kth storage location is inconsistent with the historical storage location status, it means that the storage location status of the kth storage location has changed between the shooting time of the historical video frames and the shooting time of the target video frames.
[0125] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating an update of the storage location status provided in an embodiment of this application. For example... Figure 5As shown in Figure 40a, the target storage location status obtained after the computer device performs storage location status detection on the target video frame is as follows: Of the nine storage locations contained in the target video frame, storage locations 3 to 5 and storage location 9 are in an occupied state, while the remaining storage locations are in an idle state. However, in the storage location status storage information 40b in the management system, the historical storage location status of storage locations 4 and 5 is in an occupied state, while the historical storage location status of the remaining storage locations is in an idle state. Clearly, the target storage location status (occupied status) of storage location 3 in storage location status detection result 40a is inconsistent with the historical storage location status (idle status) of storage location 3 in storage location status storage information 40b. Therefore, it is necessary to update the historical storage location status of storage location 3 in storage location status storage information 40b. In addition, the target storage location status (occupied status) of storage location 9 in storage location status detection result 40a is inconsistent with the historical storage location status (idle status) of storage location 9 in storage location status storage information 40b. Therefore, it is necessary to update the historical storage location status of storage location 9 in storage location status storage information 40b to obtain the updated storage location status storage information 40c.
[0126] In this embodiment, monitoring equipment can collect real-time video of the spatial area. Based on pre-set cargo location identification information, a cargo location cropped image corresponding to each cargo location is extracted from the video frames of the monitoring video. By performing a series of operations such as position correction, feature extraction, and feature classification on the cargo location cropped image, the target cargo location status corresponding to each cargo location is obtained. That is, the monitoring equipment can realize the monitoring of cargo locations in the entire spatial area, which can reduce the cost of cargo location status detection. The monitoring video captured by the monitoring equipment can detect the cargo location status of each cargo location in the spatial area in real time, which can improve the efficiency of cargo location status detection. In the process of managing the spatial area, the cargo location status corresponding to each cargo location can be detected in real time, improving the speed and accuracy of cargo location status detection.
[0127] Please see Figure 6 , Figure 6 This is a flowchart illustrating a method for detecting the status of a cargo location provided in an embodiment of this application. It is understood that this method can be executed by a computer device, which can be a server, a user terminal, or a computer application (including program code), and is not specifically limited thereto. Figure 6 As shown, the method for detecting the status of a cargo location may include the following steps S201-S210:
[0128] Step S201: Obtain the target video frame from the monitoring video captured by the monitoring equipment, and obtain the cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the monitoring video, where M is a positive integer.
[0129] The specific implementation process of step S201 can be found in the above. Figure 3 Step S101 in the corresponding embodiment will not be described again here.
[0130] Step S202: In the location identification information, obtain the vertex pixel coordinates corresponding to the kth location among the M locations, and determine the target rectangular region in the target video frame based on the vertex pixel coordinates of the kth location; k is a positive integer less than or equal to M.
[0131] Specifically, the computer equipment can obtain the vertex pixel coordinates of the kth storage location out of M storage locations from the storage location identification information corresponding to the monitoring equipment. Based on the vertex pixel coordinates of the kth storage location, a target rectangular region is determined in the target video frame, where M can be the number of storage locations included in the target video frame, such as M can take values of 1, 2, ..., and k can be a positive integer less than or equal to M. The computer equipment determines the location of the kth storage location in the target video frame based on the vertex pixel coordinates of the kth storage location. Since the location of the kth storage location in the target video frame may be an irregular polygon, to facilitate subsequent cropping of the target video frame, a target rectangular region containing the kth storage location can be determined based on the vertex pixel coordinates of the kth storage location. For example, when the k-th storage location includes 4 vertices, the k-th storage location can correspond to the pixel coordinates of the 4 vertices, which are (u1, v1), (u2, v2), (u3, v3), and (u4, v4). The computer device can determine the maximum and minimum horizontal coordinate values from u1, u2, u3, and u4 of the above 4 vertex pixel coordinates, and determine the maximum and minimum vertical coordinate values from v1, v2, v3, and v4 of the 4 vertex pixel coordinates. Then, based on the maximum horizontal coordinate value, the minimum horizontal coordinate value, the maximum vertical coordinate value, and the minimum vertical coordinate value, the target rectangular area can be determined. The coordinates of the 4 vertices of the target rectangular area can be represented as (maximum horizontal coordinate value, maximum vertical coordinate value), (maximum horizontal coordinate value, minimum vertical coordinate value), (minimum horizontal coordinate value, minimum vertical coordinate value), and (minimum horizontal coordinate value, maximum vertical coordinate value).
[0132] Step S203: Segment the target video frame according to the target rectangular region to obtain the kth cargo location cropped image containing the pixels covered by the target rectangular region; the kth cargo location cropped image belongs to M cargo location cropped images.
[0133] Specifically, the computer device can segment the target video frame according to the target rectangular region, and crop the k-th storage location corresponding to the target video frame. This storage location crop image can be called the k-th storage location crop image, which can contain the pixels covered by the target rectangular region in the target video frame.
[0134] Step S204: Based on the vertex pixel coordinates corresponding to the kth storage location, determine the position transformation matrix corresponding to the cropped image of the kth storage location, and perform position correction on the cropped image of the kth storage location based on the position transformation matrix to obtain the storage location area image corresponding to the kth storage location.
[0135] Specifically, based on the vertex pixel coordinates corresponding to the kth storage location, the computer equipment can determine the polygonal region of the kth storage location in the target video frame. The shooting angle of the monitoring equipment may cause the polygonal region of the kth storage location to be irregular (such as tilting, deformation, etc.). Therefore, the computer equipment can perform position correction on the kth storage location in the cropped image of the kth storage location to obtain the storage location region image corresponding to the kth storage location.
[0136] The computer device can calculate the position transformation matrix corresponding to the k-th storage location based on the vertex pixel coordinates corresponding to the k-th storage location and the vertex coordinates of a pre-set fixed polygon size. Then, a perspective transformation method can be used to stretch the irregular polygon of the k-th storage location in the cropped image of the k-th storage location to a regular polygon based on the aforementioned position transformation matrix, thus obtaining the storage location area image corresponding to the k-th storage location. Here, the regular polygon can be a rectangle, square, etc. Based on the same processing procedure, storage location area images corresponding to M storage locations can be obtained. Each storage location area image can be an image with a fixed size, and the storage locations contained in the storage location area image are stretched into regular polygons. The aforementioned position transformation matrix can refer to the matrix used to correct the position of the storage locations in the cropped image of the k-th storage location, such as stretching an irregular quadrilateral in the cropped image of the k-th storage location to a rectangle.
[0137] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating the correction of the cargo location position in a cargo location cropping image provided in an embodiment of this application. For example... Figure 7As shown in the figure, after the computer device obtains the target video frame 50a, taking the storage location 1 in the target video frame 50a as an example, the computer device can obtain the 4 vertex pixel coordinates corresponding to the storage location 1 from the storage location identification information, such as the pixel coordinates of vertex A, the pixel coordinates of vertex B, the pixel coordinates of vertex C, and the pixel coordinates of vertex D. According to the pixel coordinates corresponding to vertex A, vertex B, vertex C, and vertex D respectively, the target rectangular area 50b can be determined in the target video frame 50a. By segmenting the target video frame 50a according to the target rectangular area 50b, the storage location cropping area 50c corresponding to the storage location 1 can be obtained. Furthermore, the position transformation matrix for the storage location cropping area 50c can be calculated based on the pixel coordinates of the above 4 vertices. Based on this position transformation matrix, the storage location 1 in the storage location cropping area 50c is stretched to obtain the storage location area image 50d corresponding to the storage location 1.
[0138] Step S205: Divide the k-th storage location area image among the M storage location area images into L image blocks, and obtain the local area features corresponding to the L image blocks respectively; k is a positive integer less than or equal to M, and L is a positive integer.
[0139] Specifically, in the embodiments of the present application, taking the Histogram of Oriented Gradients (HOG) algorithm as an example, the feature extraction process of the M storage location area images is described. For any one of the M storage location area images (such as the k-th storage location area image), the computer device can divide the k-th storage location area image into L image blocks, and obtain the local area features corresponding to the L image blocks respectively according to the gradients of the pixel points included in each image block. L is the number of image blocks into which the k-th storage location area image is divided, and L is a positive integer.
[0140] Furthermore, the computer device can divide the pixel points included in the k-th storage location area image to obtain an image unit set. The number of pixel points included in each image unit in this image unit set is the same. For example, the size of each image unit in the image unit set can be , or can be etc.; according to the gradients of the pixel points included in each image unit, the gradient histogram corresponding to each image unit in the image unit set is statistically calculated; the image units in the image unit set are combined into image blocks to obtain L image blocks. For example, every 4 adjacent image units in a "tian" - shaped structure in the k-th storage location area image form an image block. Based on all the image units in the image unit set, L image blocks can be combined; furthermore, the gradient histograms corresponding to the image units in the d-th image block can be concatenated to obtain the local area feature corresponding to the d-th image block; where the d-th image block belongs to the L image blocks, and d is a positive integer less than or equal to L.
[0141] For example, in the process of feature extraction from the k-th storage location area image, assuming the size of the k-th storage location area image is , divide the pixels in the k-th storage location area image into an image unit (cell), that is, the k-th storage location area image can be divided into image units, that is, the above image unit set can include 128 image units; 4 adjacent image units in a "field" structure can be combined into an image block (block), and image blocks can be obtained (here L can take the value of 105). The computer device can project the gradient directions of all pixels in a single image unit to form a gradient histogram corresponding to each image unit respectively. For example, 9 direction bins can be preset in advance, and each 20 degrees can correspond to a direction bin. The directions from 0 degrees to 180 degrees and from 180 degrees to 360 degrees can be classified and divided by the method of equal angles to generate the gradient histogram corresponding to each image unit respectively; by concatenating the data of the gradient histograms of the 4 image units included in each image block, the local area feature corresponding to each image block can be obtained. For example, if the gradient histogram of a single image unit is a 9-dimensional vector, then each image block can extract a 36-dimensional vector, and the 36-dimensional vector at this time can be called the local area feature of the image block. It should be noted that the sizes of the above image units and image blocks are only examples, and this application does not limit the sizes of image units and image blocks.
[0142] Step S206: Concatenate the local area features corresponding to the L image blocks to obtain the storage location area feature corresponding to the k-th storage location area image.
[0143] Specifically, the computer device can concatenate the local area features corresponding to the L image blocks according to the pixel point order in the target video frame to obtain the storage location area feature corresponding to the k-th storage location area image. As in the previous example, if the local area feature of each image block is a 36-dimensional feature vector, then it can be determined that the storage location area feature corresponding to the k-th storage location area image is dimensional feature.
[0144] Step S207: Input the storage location area feature corresponding to the k-th storage location area image into the first classifier, and identify the storage location area feature of the k-th storage location area image in the first classifier to obtain the target storage location state corresponding to the storage location in the k-th storage location area image.
[0145] Specifically, this application uses Support Vector Machine (SVM) as an example to describe the classification process of the storage area features corresponding to M storage area images. For the storage area features of any storage area image (such as the k-th storage area image) among the M storage area images, the computer device can input the storage area features corresponding to the k-th storage area image into a first classifier (such as a support vector machine). In the first classifier, the predicted distance between the classification hyperplane and the storage area features corresponding to the k-th storage area image can be determined. If the predicted distance is greater than the classification threshold, the target storage location state corresponding to the storage location in the k-th storage area image is determined to be occupied. If the predicted distance is less than the classification threshold, the target storage location state corresponding to the storage location in the k-th storage area image is determined to be idle. The classification threshold can be 0, and the classification hyperplane in the first classifier is obtained through training.
[0146] Before using a Support Vector Machine (SVM) to classify the aforementioned cargo location area features, the SVM needs to be trained. This SVM can be based on a maximum margin learning method. The maximum margin learning strategy is margin maximization, which can be solved by solving an optimal algorithm of convex quadratic programming. The computer equipment can acquire positive and negative sample area images from historical surveillance videos captured by the monitoring equipment. The positive sample area images may include occupancy labels, and the negative sample area images may include free labels. Then, it can acquire the positive sample area features corresponding to the positive sample area images, and the negative sample area features corresponding to the negative sample area images. Based on the positive sample area features and occupancy labels, a first hyperplane in the first classifier is determined, and based on the negative sample area features and free labels, a second hyperplane in the first classifier is determined. The first hyperplane is parallel to the second hyperplane. When the margin between the first and second hyperplanes is optimized to its maximum value, a classification hyperplane in the first classifier is determined based on the first and second hyperplanes. The distance between the classification hyperplane and the first hyperplane is equal to the distance between the classification hyperplane and the second hyperplane. In other words, support vector machines are used to separate positive and negative sample regions in an image and find the optimal segmentation plane, which can be called the classification hyperplane.
[0147] Please see Figure 8 , Figure 8 This is a flowchart illustrating the training process of a classifier provided in an embodiment of this application. Figure 8As shown, the training process of the classifier can be implemented through steps S11 to S14. In step S11, the computer device can acquire historical monitoring videos captured by the monitoring equipment and extract positive sample region images and negative sample region images from the historical monitoring videos. The positive sample region images can carry an occupancy label to indicate that the storage locations in these positive sample region images are occupied, such as positive sample region images 60a, 60b, and 60c, etc.; the negative sample region images can carry an idle label to indicate that the storage locations in these negative sample region images are not occupied, such as negative sample region images 60d, 60e, and 60f, etc.
[0148] In step S12, the computer device can perform position correction on the cargo locations in each positive sample region image and each negative sample region image, obtaining position-corrected positive sample region images and position-corrected negative sample region images. Then, step S13 can be executed to extract features from both the position-corrected positive and negative sample region images, obtaining the positive sample region features of the position-corrected positive sample region image and the negative sample region features of the position-corrected negative sample region image. Based on the negative and positive sample region features, step S14 is then executed, i.e., the classifier is trained. The classifier here can be a support vector machine, such as the first classifier mentioned above.
[0149] The purpose of training a support vector machine (SVM) is to separate negative and positive sample region features and find the optimal separating hyperplane, i.e., the classification hyperplane mentioned above. This separating hyperplane must not only separate negative and positive sample region features but also maximize the margin between the two classes of sample region features. Negative and positive sample region features are collectively referred to as sample region features, and are denoted as {x}. i y i}, i = 1, 2, 3, ..., S; where x i y is the feature of the i-th sample region. i ={+1, -1} are the category labels for binary classification, where +1 represents the occupied label and -1 represents the idle label. Assume the hyperplane of the support vector machine is as shown in the following formula (6):
[0150] (6)
[0151] Here, w can be the normal vector of the hyperplane, and b is the offset of the hyperplane.
[0152] For linearly separable regions, two hyperplanes can be found such that for the features of the positive sample region, Features of negative sample regions The interval between the two parallel hyperplanes is The two parallel hyperplanes here are the first and second hyperplanes in the first classifier mentioned above. To maximize the margin, we need to make... Minimum. Therefore, the optimization model can be represented by the following formula (7):
[0153] (7)
[0154] Optionally, for the linearly inseparable case, the analytical expressions for its normal vector w and offset b cannot be obtained. Therefore, it is necessary to improve upon the above formula (7), and the improved optimization model is shown in the following formula (8):
[0155]
[0156] (8)
[0157] For any given sample region feature, The value is always greater than or equal to 0, and c is a parameter. For linearly separable or inseparable cases, the optimal hyperplane of the support vector machine can be obtained by solving the above formula (7) or formula (8).
[0158] Please see Figure 9 , Figure 9 This is a schematic diagram of a cargo location status detection process provided in an embodiment of this application. For example... Figure 9 As shown, the cargo location status detection process can be implemented through steps S21 to S25. In step S21, after the computer device acquires the monitoring video captured by the monitoring equipment, it can extract the target video frame from the monitoring video and extract the area cropped images corresponding to each cargo location from the target video frame. Since the cargo locations in the area cropped images can be irregular polygons, step S22 can be executed to perform image position correction on each area cropped image, obtaining the cargo location area image corresponding to each cargo location in the target video frame. Then, step S23 can be executed to extract features from the cargo location area images to obtain the cargo location area features corresponding to each cargo location area image. The first classifier, which has been trained as described above, is used to identify the cargo location area features, i.e., to calculate the predicted distance between the cargo location area features and the optimal hyperplane in the first classifier. Based on the predicted distance, the target cargo location status corresponding to each cargo location in the target video frame is output.
[0159] Step S208: If two target storage locations among the M storage locations are both occupied, then the two storage locations are identified as a target storage location pair.
[0160] Specifically, in practical applications, an item that could be stored in one storage location might be mistakenly placed between two adjacent locations due to human error, resulting in the item occupying two locations and thus wasting space. Therefore, after detecting the target storage location status of M storage locations, the computer equipment can further detect adjacent storage locations that are also occupied to ensure the accuracy of the target storage location status. In other words, if two adjacent storage locations among the M storage locations are both occupied, these two adjacent storage locations can be called a target storage location pair; that is, two adjacent storage locations that are both occupied are considered the target storage location pair that need to be further detected.
[0161] Step S209: Obtain the image of the area to be detected containing the target location pair in the target video frame, perform item detection on the image of the area to be detected, and obtain the item detection result corresponding to the image of the area to be detected.
[0162] Specifically, the computer equipment can extract the image of the area to be detected where the target storage location is located from the target video frame. By performing item detection on the image of the area to be detected, the edge shape of the items in the target storage location can be obtained. Based on the edge shape of the items, the item detection result corresponding to the image of the area to be detected can be determined. The item detection result can include item overlap result and item non-overlap result. Item overlap result is used to indicate that the target storage location is occupied by the same item, and item non-overlap result is used to indicate that the target storage location is occupied by different items.
[0163] Specifically, the computer equipment can determine the cropping region associated with the target storage location pair in the target video frame based on the vertex pixel coordinates corresponding to the target storage location pair, and define the pixels covered by the cropping region as the area to be detected. Then, it can acquire the edge features of the items in the area to be detected and determine the edge shape of the items in the area to be detected based on these features. Based on the vertex pixel coordinates corresponding to each storage location in the target storage location pair, it can determine the edge shape of each storage location in the target storage location pair. When the edge shape of the item overlaps with the edge shape of each storage location in the target storage location pair, the item detection result corresponding to the area to be detected is determined to be an item overlap result. When the edge shape of the item overlaps only with the edge shape of one storage location in the target storage location pair, the item detection result corresponding to the area to be detected is determined to be an item non-overlap result, meaning the item is placed in only one storage location and does not occupy any additional storage locations.
[0164] Step S210: When the item detection result is an item overlap result, generate an abnormal prompt message for the target storage location pair; the item overlap result is used to indicate that the target storage location pair is occupied by the same item.
[0165] Specifically, when the item detection result indicates overlapping items, an anomaly alert can be generated for the target storage location pair. This alert reminds the user to check the item storage status in the target storage location pair and adjust the items stored there in a timely manner to improve storage location utilization. When the item detection result indicates non-overlapping items, it can be determined that the items in the target storage location pair are placed normally and there is no anomaly.
[0166] Please see Figure 10 , Figure 10 This is a schematic diagram of an article detection method provided in an embodiment of this application. Figure 10 As shown, if the target storage locations 1 and 2 in the target video frame are both occupied and have an adjacent position relationship, then storage locations 1 and 2 can be called a target storage location pair. The detection area image 70a containing storage locations 1 and 2 is obtained by cropping from the target video frame. By performing item detection on the detection area image 70a, the edge shape of the item in the detection area image 70a is obtained as edge shape 70b. Since this edge shape 70b overlaps with the edge shape of storage location 1 and the edge shape of storage location 2, it can be determined that the item placed in storage location 1 and storage location 2 is the same. That is, the item detection result of the detection area image 70a is the item overlap result, and an abnormal prompt message is generated for storage locations 1 and 2.
[0167] Optionally, if the target storage locations 3 and 4, which are adjacent in the target video frame, are both occupied, then storage locations 3 and 4 can also be referred to as a target storage location pair. A detection area image 70c containing storage locations 3 and 4 can be cropped from the target video frame. By performing item detection on the detection area image 70c, the edge shapes of the items in the detection area image 70c are obtained as edge shape 70d and edge shape 70e. Since edge shape 70d only overlaps with the edge shape of storage location 3, and edge shape 70e only overlaps with the edge shape of storage location 4, it can be determined that the items placed in storage locations 3 and 4 are different. That is, the item detection result of the detection area image 70c is that the items do not overlap.
[0168] It should be noted that steps S208 to S210 above are only one optional embodiment provided by this application. After obtaining the target storage location status corresponding to each of the M storage locations, the computer device may stop executing steps S208 to S210 and directly compare the target storage location status with the historical storage location status stored in the management system to realize the storage location status update process; or, after obtaining the target storage location status corresponding to each of the M storage locations, it may continue executing steps S208 to S210 above, suspend the update of the historical storage location status of the target storage location pair corresponding to the abnormal prompt information, and update the historical storage location status of the remaining storage locations.
[0169] In this embodiment, monitoring equipment can collect real-time video of the spatial area. Based on pre-set cargo location identification information, a cargo location cropped image corresponding to each cargo location is extracted from the video frames of the monitoring video. By performing a series of operations such as position correction, feature extraction, and feature classification on the cargo location cropped image, the target cargo location status corresponding to each cargo location is obtained. That is, the monitoring equipment can realize the monitoring of cargo locations in the entire spatial area, which can reduce the cost of cargo location status detection. The monitoring video captured by the monitoring equipment can detect the cargo location status of each cargo location in the spatial area in real time, which can improve the efficiency of cargo location status detection. In the process of managing the spatial area, the cargo location status corresponding to each cargo location can be detected in real time, improving the speed and accuracy of cargo location status detection.
[0170] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a cargo location status detection device provided in an embodiment of this application. Figure 11 As shown, the storage location status detection device 1 may include: a first acquisition module 11, a storage location cutting module 12, and a storage location detection module 13;
[0171] The first acquisition module 11 is used to acquire target video frames from the monitoring video captured by the monitoring equipment and acquire cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the monitoring video, where M is a positive integer;
[0172] The cargo location cropping module 12 is used to obtain M cargo location cropping images from the target video frame based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, and to perform position correction on the cargo locations contained in the M cargo location cropping images to obtain M cargo location area images.
[0173] The storage location detection module 13 is used to acquire the storage location area features corresponding to the M storage location area images respectively, and determine the target storage location status corresponding to the M storage locations in the target video frame according to the storage location area features; the target storage location status is used to indicate the occupancy status of the items in the M storage locations.
[0174] The specific functional implementation methods of the first acquisition module 11, the storage location cutting module 12, and the storage location detection module 13 can be found in the above description. Figure 3 Steps S101-S103 in the corresponding embodiments will not be described again here.
[0175] In one or more embodiments, the target storage location status includes an occupied status;
[0176] The storage location status detection device 1 may also include: a storage location pair determination module 14, an item detection module 15, and an anomaly prompting module 16;
[0177] The location pair determination module 14 is used to determine the two locations as a target location pair if there are two adjacent locations among the M locations, both of which are occupied.
[0178] The item detection module 15 is used to acquire an image of the area to be detected containing the target location pair in the target video frame, perform item detection on the image of the area to be detected, and obtain the item detection result corresponding to the image of the area to be detected.
[0179] The anomaly alert module 16 is used to generate anomaly alert information for the target storage location pair when the item detection result is an item overlap result; the item overlap result is used to indicate that the target storage location pair is occupied by the same item.
[0180] The specific implementation methods of the location matching module 14, the item detection module 15, and the anomaly alert module 16 can be found in the above description. Figure 6 Steps S208-S210 in the corresponding embodiments will not be described again here.
[0181] In one or more embodiments, the item detection module 15 may include: a region clipping unit 151, an object shape determination unit 152, a storage location shape determination unit 153, and an item detection result determination unit 154.
[0182] The region cropping unit 151 is used to determine the cropping region associated with the target cargo location pair in the target video frame based on the vertex pixel coordinates corresponding to the target cargo location pair, and to determine the pixels covered by the cropping region as the region image to be detected.
[0183] The object shape determination unit 152 is used to acquire the object edge features in the image of the region to be detected, and determine the object edge shape in the image of the region to be detected based on the object edge features;
[0184] The storage location shape determination unit 153 is used to determine the edge shape of each storage location in the target storage location pair based on the vertex pixel coordinates corresponding to each storage location in the target storage location pair.
[0185] The item detection result determination unit 154 is used to determine the item detection result corresponding to the image of the area to be detected as the item overlap result when the edge shape of the item overlaps with the edge shape of each storage location in the target storage location pair.
[0186] The specific functional implementation methods of the region clipping unit 151, the object shape determination unit 152, the storage location shape determination unit 153, and the item detection result determination unit 154 can be found above. Figure 6 Step S209 in the corresponding embodiment will not be described again here.
[0187] In one or more embodiments, the first acquisition module 11 may include: a frame processing unit 111, a background template acquisition unit 112, and a video frame extraction unit 113.
[0188] The frame segmentation processing unit 111 is used to acquire the monitoring video captured by the monitoring equipment, perform frame segmentation processing on the monitoring video, and obtain the video frame sequence corresponding to the monitoring video.
[0189] Background template acquisition unit 112 is used to acquire N background templates corresponding to the monitoring device, and the applicable time ranges corresponding to the N background templates respectively; the N background templates are obtained by modeling the background of the images captured by the monitoring device in different time periods, and N is a positive integer;
[0190] The video frame extraction unit 113 is used to determine the i-th video frame as the target video frame when the shooting time of the i-th video frame belongs to the applicable time range corresponding to the j-th background template and the i-th video frame is inconsistent with the j-th background template; the i-th video frame belongs to the video frame sequence, the j-th background template belongs to N background templates, i is a positive integer less than or equal to the number of video frames in the video frame sequence, and j is a positive integer less than or equal to N.
[0191] The specific functional implementation methods of the frame processing unit 111, the background template acquisition unit 112, and the video frame extraction unit 113 can be found above. Figure 3 Step S101 in the corresponding embodiment will not be described again here.
[0192] In one or more embodiments, the cargo location trimming module 12 may include: a rectangular area determination unit 121, a video frame segmentation unit 122, and a position correction unit 123;
[0193] The rectangular region determination unit 121 is used to obtain the vertex pixel coordinates of the kth storage location among M storage locations from the storage location identification information, and determine the target rectangular region in the target video frame based on the vertex pixel coordinates of the kth storage location; k is a positive integer less than or equal to M.
[0194] The video frame segmentation unit 122 is used to segment the target video frame according to the target rectangular region to obtain the kth location cropped image containing the pixels covered by the target rectangular region; the kth location cropped image belongs to M location cropped images;
[0195] The position correction unit 123 is used to determine the position transformation matrix corresponding to the cropped image of the kth storage location based on the vertex pixel coordinates corresponding to the kth storage location, and to perform position correction on the cropped image of the kth storage location based on the position transformation matrix to obtain the storage location area image corresponding to the kth storage location.
[0196] The specific functional implementation methods of the rectangular region determination unit 121, the video frame segmentation unit 122, and the position correction unit 123 can be found above. Figure 6 Steps S202-S204 in the corresponding embodiments will not be described again here.
[0197] In one or more embodiments, the cargo location status detection device 1 may further include: a second acquisition module 17, a vertex coordinate determination module 18, and a cargo location identifier generation module 19;
[0198] The second acquisition module 17 is used to acquire the physical space coordinate information corresponding to the kth storage location, as well as the device extrinsic parameters and device intrinsic parameters corresponding to the monitoring device; the device extrinsic parameters are used to characterize the position of the monitoring device in physical space, and the device intrinsic parameters are determined by the optical center and the focal length of the monitoring device;
[0199] Vertex coordinate determination module 18 is used to determine the vertex pixel coordinates corresponding to the kth storage location based on physical space coordinate information, equipment external parameters and equipment internal parameters;
[0200] The cargo location identification generation module 19 is used to generate cargo location identification information corresponding to the monitoring device based on the correspondence between the vertex pixel coordinates corresponding to the kth cargo location and the kth cargo location.
[0201] The specific functional implementation methods of the second acquisition module 17, the vertex coordinate determination module 18, and the cargo location identifier generation module 19 can be found in the above description. Figure 3 Step S101 in the corresponding embodiment will not be described again here.
[0202] In one or more embodiments, the cargo location detection module 13 may include: a first feature extraction unit 131, a feature concatenation unit 132, and a first classification unit 133;
[0203] The first feature extraction unit 131 is used to divide the kth storage location image in M storage location area images into L image blocks and obtain the local region features corresponding to the L image blocks respectively; k is a positive integer less than or equal to M, and L is a positive integer;
[0204] The feature concatenation unit 132 is used to concatenate the local region features corresponding to L image blocks to obtain the cargo location region features corresponding to the kth cargo location region image.
[0205] The first classification unit 133 is used to input the cargo location area features corresponding to the kth cargo location area image into the first classifier, and to identify the cargo location area features of the kth cargo location area image in the first classifier to obtain the target cargo location status corresponding to the cargo location in the kth cargo location area image.
[0206] The first feature extraction unit 131 may include: an image segmentation subunit 1311, a gradient histogram statistics subunit 1312, and a local feature acquisition subunit 1313.
[0207] Image segmentation subunit 1311 is used to segment the pixels contained in the kth storage area image among M storage area images to obtain an image unit set; each image unit in the image unit set contains the same number of pixels;
[0208] The gradient histogram statistics subunit 1312 is used to count the gradient histogram corresponding to each image unit in the image unit set based on the gradient of the pixels contained in each image unit.
[0209] The local feature acquisition subunit 1313 is used to combine the image units in the image unit set into image blocks to obtain L image blocks. The gradient histograms corresponding to the image units in the d-th image block are concatenated to obtain the local region features corresponding to the d-th image block. The d-th image block belongs to L image blocks, where d is a positive integer less than or equal to L.
[0210] The first classification unit 133 may include: a prediction distance determination subunit 1331 and a cargo location status determination subunit 1332;
[0211] The prediction distance determination subunit 1331 is used to input the storage area features corresponding to the kth storage area image into the first classifier, and determine the prediction distance between the classification hyperplane and the storage area features corresponding to the kth storage area image in the first classifier.
[0212] The storage location status determination subunit 1332 is used to determine the target storage location status corresponding to the storage location in the k-th storage location area image as occupied if the prediction distance is greater than the classification threshold.
[0213] The aforementioned cargo location status determination subunit 1332 is also used to determine the target cargo location status corresponding to the cargo location in the k-th cargo location area image as an idle state if the prediction distance is less than the classification threshold.
[0214] Optionally, the location detection module 13 may also include: a second feature extraction unit 134, a matching degree acquisition unit 135, and a second classification unit 136;
[0215] The second feature extraction unit 134 is used to input the kth storage location image from the M storage location area images into the image recognition model, and obtain the storage location area features corresponding to the kth storage location area image in the image recognition model.
[0216] The matching degree acquisition unit 135 is used to acquire, through the second classifier in the image recognition model, the first matching degree between the occupied attribute feature and the cargo area feature corresponding to the kth cargo area image, and the second matching degree between the idle attribute feature and the cargo area feature corresponding to the kth cargo area image.
[0217] The second classification unit 136 is used to determine the target storage location status of the storage location in the kth storage location area image as occupied if the first matching degree is greater than the second matching degree.
[0218] The second classification unit is further configured to determine that the target storage location corresponding to the storage location in the kth storage location area image is in an idle state if the first matching degree is less than the second matching degree.
[0219] The second feature extraction unit 134 may include: a convolution subunit 1341, a residual subunit 1342, and a cargo location feature acquisition unit 1343.
[0220] The convolutional subunit 1341 is used to input the kth storage location image from M storage location images into the image recognition model, and perform convolution processing on the kth storage location image according to the convolutional layer in the image recognition model to obtain the storage location convolutional features.
[0221] The residual subunit 1342 is used to perform residual convolution processing on the location convolution features based on the residual layer in the image recognition model to obtain the location residual features;
[0222] The location feature acquisition unit 1343 is used to generate the location area features corresponding to the k-th location area image based on the location convolution features and location residual features.
[0223] The specific functional implementation of the first feature extraction unit 131, the feature concatenation unit 132, and the first classification unit 133 can be found above. Figure 6The specific functional implementation methods of steps S205-S207, the second feature extraction unit 134, the matching degree acquisition unit 135, the second classification unit 136, and the sub-units contained in each unit in the corresponding embodiment can be found in the above description. Figure 3 Step S103 in the corresponding embodiment will not be described again here. Specifically, when the first feature extraction unit 131, the feature concatenation unit 132, and the first classification unit 133 are performing their respective operations, the second feature extraction unit 134, the matching degree acquisition unit 135, and the second classification unit 136 all pause their operations; when the second feature extraction unit 134, the matching degree acquisition unit 135, and the second classification unit 136 are performing their respective operations, the first feature extraction unit 131, the feature concatenation unit 132, and the first classification unit 133 all pause their operations.
[0224] In one or more embodiments, the storage location status detection device 1 may further include a storage location status update module 20.
[0225] The storage location status update module 20 is used to update the historical storage location status in the management system to the target storage location status corresponding to the kth storage location when the target storage location status corresponding to the kth storage location in the M storage locations is inconsistent with the historical storage location status corresponding to the kth storage location in the management system. The historical storage location status refers to the storage location status determined based on historical video frames. The shooting time of the historical video frames is earlier than the shooting time of the target video frames, and k is a positive integer less than or equal to M.
[0226] The specific implementation of the storage location status update module 20 can be found in the above description. Figure 3 Step S103 in the corresponding embodiment will not be described again here.
[0227] In one or more embodiments, the cargo location status detection device 1 may further include: a training sample acquisition module 21, a negative sample feature extraction module 22, a positive sample feature extraction module 23, and a hyperplane optimization module 24.
[0228] The training sample acquisition module 21 is used to acquire positive sample region images and negative sample region images from historical surveillance videos captured by the monitoring equipment; the positive sample region images include occupied labels, and the negative sample region images include idle labels.
[0229] The negative sample feature extraction module 22 is used to obtain the positive sample region features corresponding to the positive sample region image and to obtain the negative sample region features corresponding to the negative sample region image.
[0230] The positive sample feature extraction module 23 is used to determine the first hyperplane in the first classifier based on the positive sample region features and occupied labels, and to determine the second hyperplane in the first classifier based on the negative sample region features and idle labels; the first hyperplane is parallel to the second hyperplane.
[0231] The hyperplane optimization module 24 is used to determine the classification hyperplane in the first classifier based on the first hyperplane and the second hyperplane when the interval distance between the first hyperplane and the second hyperplane is optimized to the maximum value; the distance between the classification hyperplane and the first hyperplane is equal to the distance between the classification hyperplane and the second hyperplane.
[0232] The specific implementation methods of the training sample acquisition module 21, negative sample feature extraction module 22, positive sample feature extraction module 23, and hyperplane optimization module 24 can be found above. Figure 6 Step S207 in the corresponding embodiment will not be described again here.
[0233] In this embodiment, monitoring equipment can collect real-time video of the spatial area. Based on pre-set cargo location identification information, a cargo location cropped image corresponding to each cargo location is extracted from the video frames of the monitoring video. By performing a series of operations such as position correction, feature extraction, and feature classification on the cargo location cropped image, the target cargo location status corresponding to each cargo location is obtained. That is, the monitoring equipment can realize the monitoring of cargo locations in the entire spatial area, which can reduce the cost of cargo location status detection. The monitoring video captured by the monitoring equipment can detect the cargo location status of each cargo location in the spatial area in real time, which can improve the efficiency of cargo location status detection. In the process of managing the spatial area, the cargo location status corresponding to each cargo location can be detected in real time, improving the speed and accuracy of cargo location status detection.
[0234] Further, please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 12 As shown, the computer device 1000 can be a user terminal, for example, the one described above. Figure 1 The user terminal 10d in the corresponding embodiment can also be a server, for example, as described above. Figure 1The server 10c in the corresponding embodiment will not be limited here. For ease of understanding, this application takes a computer device as a user terminal as an example. The computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1004 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The memory 1005 may optionally be at least one storage device located remotely from the aforementioned processor 1001. Figure 12 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0235] The network interface 1004 in the computer device 1000 can also provide network communication functions, and the optional user interface 1003 can also include a display screen and a keyboard. Figure 12 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to achieve:
[0236] Obtain the target video frame from the surveillance video captured by the monitoring equipment, and obtain the cargo location identification information associated with the monitoring equipment; the cargo location identification information includes the vertex pixel coordinates of M cargo locations in the surveillance video, where M is a positive integer;
[0237] Based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, M cargo location cropped images are obtained from the target video frame. The positions of the cargo locations contained in the M cargo location cropped images are corrected to obtain M cargo location area images.
[0238] Obtain the features of the M storage locations corresponding to each storage location area image, and determine the target storage location status of the M storage locations in the target video frame based on the storage location area features; the target storage location status is used to indicate the occupancy status of the items in the M storage locations.
[0239] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 and Figure 6The description of the storage location status detection method in any corresponding embodiment can also be executed as described above. Figure 7 The description of the cargo location status detection device 1 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0240] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned cargo location status detection device 1. The computer program includes program instructions, and when the processor executes the program instructions, it can execute the aforementioned... Figure 3 and Figure 6 The description of the cargo location status detection method in any corresponding embodiment is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed and executed on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network. These multiple computing devices distributed across multiple locations and interconnected via a communication network can constitute a blockchain system.
[0241] Furthermore, it should be noted that this application also provides a computer program product or computer program, which may include computer instructions, which may be stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, causing the computer device to perform the aforementioned actions. Figure 3 and Figure 6 The description of the storage location status detection method in any corresponding embodiment is already provided, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program products or computer program embodiments involved in this application, please refer to the description of the method embodiments of this application.
[0242] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0243] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0244] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0245] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0246] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method for detecting the status of a cargo location, characterized in that, include: Obtain target video frames from the surveillance video captured by the monitoring equipment, and obtain the cargo location identification information associated with the monitoring equipment; The cargo location identification information includes the vertex pixel coordinates of M cargo locations in the monitoring video, where M is a positive integer; Based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, M cargo location cropped images are obtained from the target video frame, and the cargo locations contained in the M cargo location cropped images are positionally corrected to obtain M cargo location area images. Obtain the cargo location area features corresponding to the M cargo location area images respectively, and determine the target cargo location status corresponding to the M cargo locations in the target video frame based on the cargo location area features; The target storage location status is used to indicate the occupancy status of the items in the M storage locations, and the target storage location status includes the occupancy status; If two adjacent storage locations among the M storage locations are both occupied, then the two storage locations are identified as a pair of target storage locations. In the target video frame, an image of the area to be detected containing the target storage location pair is obtained. Item detection is performed on the image of the area to be detected to obtain the item detection result corresponding to the image of the area to be detected. When the item detection result is an item overlap result, an abnormal prompt message is generated for the target storage location pair; the item overlap result is used to indicate that the target storage location pair is occupied by the same item.
2. The method according to claim 1, characterized in that, The step of acquiring a detection area image containing the target storage location pair in the target video frame, performing item detection on the detection area image, and obtaining the item detection result corresponding to the detection area image includes: Based on the vertex pixel coordinates corresponding to the target cargo location pair, the cropping region associated with the target cargo location pair is determined in the target video frame, and the pixels covered by the cropping region are determined as the area image to be detected. Obtain the object edge features in the image of the region to be detected, and determine the object edge shape in the image of the region to be detected based on the object edge features; Based on the vertex pixel coordinates corresponding to each location in the target location pair, determine the edge shape of each location in the target location pair. When the edge shape of the item overlaps with the edge shape of each storage location in the target storage location pair, the item detection result corresponding to the image of the area to be detected is determined as the item overlap result.
3. The method according to claim 1, characterized in that, The step of obtaining the target video frame from the surveillance video captured by the monitoring equipment includes: The monitoring video captured by the monitoring device is acquired, and the monitoring video is processed into frames to obtain the video frame sequence corresponding to the monitoring video. Obtain N background templates corresponding to the monitoring device, and the applicable time ranges corresponding to the N background templates respectively; the N background templates are obtained by modeling the background of the images captured by the monitoring device in different time periods, where N is a positive integer; When the shooting time of the i-th video frame falls within the applicable time range of the j-th background template, and the i-th video frame is inconsistent with the j-th background template, the i-th video frame is determined as the target video frame; the i-th video frame belongs to the video frame sequence, and the j-th background template belongs to the N background templates, where i is a positive integer less than or equal to the number of video frames in the video frame sequence, and j is a positive integer less than or equal to N.
4. The method according to claim 1, characterized in that, The step of obtaining M cropped images of cargo locations from the target video frame based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, and performing position correction on the cargo locations contained in the M cropped images of cargo locations to obtain M cargo location area images includes: In the location identification information, the vertex pixel coordinates corresponding to the kth location among the M locations are obtained. Based on the vertex pixel coordinates of the kth location, a target rectangular region is determined in the target video frame; k is a positive integer less than or equal to M. The target video frame is segmented according to the target rectangular region to obtain the kth location cropped image containing the pixels covered by the target rectangular region; the kth location cropped image belongs to the M location cropped images; Based on the vertex pixel coordinates corresponding to the kth storage location, the position transformation matrix corresponding to the cropped image of the kth storage location is determined. The position of the cropped image of the kth storage location is then corrected based on the position transformation matrix to obtain the storage location area image corresponding to the kth storage location.
5. The method according to claim 4, characterized in that, Also includes: Obtain the physical space coordinates of the kth storage location, and obtain the external and internal parameters of the monitoring device. The extrinsic parameters of the device are used to characterize the position of the monitoring device in physical space, and the intrinsic parameters of the device are determined by the optical center and the focal length of the monitoring device. Based on the physical space coordinate information, the device extrinsic parameters, and the device intrinsic parameters, determine the vertex pixel coordinates corresponding to the kth storage location; Based on the correspondence between the vertex pixel coordinates corresponding to the kth storage location and the kth storage location, the storage location identification information corresponding to the monitoring device is generated.
6. The method according to claim 1, characterized in that, The step of acquiring the cargo location area features corresponding to the M cargo location area images respectively, and determining the target cargo location status corresponding to the M cargo locations in the target video frame based on the cargo location area features, includes: Divide the k-th storage location image among the M storage location area images into L image blocks, and obtain the local region features corresponding to the L image blocks respectively; k is a positive integer less than or equal to M, and L is a positive integer; The local region features corresponding to the L image blocks are concatenated to obtain the cargo location region features corresponding to the kth cargo location region image. The features of the storage area corresponding to the kth storage area image are input into the first classifier. The storage area features of the kth storage area image are identified in the first classifier to obtain the target storage location status corresponding to the storage location in the kth storage area image.
7. The method according to claim 6, characterized in that, The step of dividing the k-th storage location image among the M storage location area images into L image blocks and obtaining the local region features corresponding to the L image blocks respectively includes: The pixels contained in the kth storage area image among the M storage area images are divided to obtain an image unit set; each image unit in the image unit set contains the same number of pixels; Based on the gradient of the pixels contained in each image unit, the gradient histogram corresponding to each image unit in the image unit set is calculated. The image units in the image unit set are combined into image blocks to obtain the L image blocks. The gradient histograms corresponding to the image units in the d-th image block are concatenated to obtain the local region features corresponding to the d-th image block. The d-th image block belongs to the L image blocks, where d is a positive integer less than or equal to L.
8. The method according to claim 6, characterized in that, The step of inputting the storage location area features corresponding to the k-th storage location area image into a first classifier, and identifying the storage location area features of the k-th storage location area image in the first classifier to obtain the target storage location status corresponding to the storage location in the k-th storage location area image includes: The features of the storage area corresponding to the kth storage area image are input into the first classifier, and the predicted distance between the classification hyperplane and the features of the storage area corresponding to the kth storage area image is determined in the first classifier. If the predicted distance is greater than the classification threshold, then the target storage location corresponding to the storage location in the kth storage location area image is determined to be in an occupied state. If the predicted distance is less than the classification threshold, then the target storage location corresponding to the storage location in the kth storage location area image is determined to be in an idle state.
9. The method according to claim 1, characterized in that, The step of acquiring the cargo location area features corresponding to the M cargo location area images respectively, and determining the target cargo location status corresponding to the M cargo locations in the target video frame based on the cargo location area features, includes: The kth storage location image among the M storage location images is input into the image recognition model, and the storage location features corresponding to the kth storage location image are obtained in the image recognition model. The second classifier in the image recognition model is used to obtain the first matching degree between the occupied attribute feature and the storage area feature corresponding to the kth storage area image, and the second matching degree between the idle attribute feature and the storage area feature corresponding to the kth storage area image. If the first matching degree is greater than the second matching degree, then the target storage location corresponding to the storage location in the kth storage location area image is determined to be in an occupied state. If the first matching degree is less than the second matching degree, then the target storage location corresponding to the storage location in the kth storage location area image is determined to be in an idle state.
10. The method according to claim 9, characterized in that, The step of inputting the kth storage location image from the M storage location area images into an image recognition model, and obtaining the storage location area features corresponding to the kth storage location area image in the image recognition model, includes: The kth storage location image among the M storage location images is input into the image recognition model. The kth storage location image is then convolved according to the convolutional layer in the image recognition model to obtain the storage location convolutional features. Based on the residual layer in the image recognition model, the cargo location convolutional features are subjected to residual convolution processing to obtain cargo location residual features. Based on the convolutional features of the cargo locations and the residual features of the cargo locations, the cargo location region features corresponding to the k-th cargo location region image are generated.
11. The method according to claim 1, characterized in that, Also includes: When the target storage location status corresponding to the kth storage location among the M storage locations is inconsistent with the historical storage location status corresponding to the kth storage location in the management system, the historical storage location status is updated to the target storage location status corresponding to the kth storage location in the management system; the historical storage location status refers to the storage location status determined based on historical video frames, the shooting time of the historical video frames is earlier than the shooting time of the target video frames, and k is a positive integer less than or equal to M.
12. The method according to claim 6, characterized in that, Also includes: Positive sample region images and negative sample region images are obtained from historical surveillance videos captured by the monitoring equipment; The positive sample region image includes occupied labels, and the negative sample region image includes idle labels; Obtain the positive sample region features corresponding to the positive sample region image, and obtain the negative sample region features corresponding to the negative sample region image; Based on the positive sample region features and the occupied label, a first hyperplane is determined in the first classifier; based on the negative sample region features and the idle label, a second hyperplane is determined in the first classifier; the first hyperplane is parallel to the second hyperplane. When the distance between the first hyperplane and the second hyperplane is optimized to the maximum value, the classification hyperplane in the first classifier is determined based on the first hyperplane and the second hyperplane. The distance between the classification hyperplane and the first hyperplane is equal to the distance between the classification hyperplane and the second hyperplane.
13. A storage location status detection device, characterized in that, include: The first acquisition module is used to acquire target video frames from the monitoring video captured by the monitoring equipment and to acquire cargo location identification information associated with the monitoring equipment. The cargo location identification information includes the vertex pixel coordinates of M cargo locations in the monitoring video, where M is a positive integer; The cargo location cropping module is used to obtain M cargo location cropping images from the target video frame based on the vertex pixel coordinates corresponding to the M cargo locations contained in the cargo location identification information, and to perform position correction on the cargo locations contained in the M cargo location cropping images to obtain M cargo location area images. The storage location detection module is used to acquire the storage location area features corresponding to the M storage location area images respectively, and determine the target storage location status corresponding to the M storage locations in the target video frame according to the storage location area features; the target storage location status is used to indicate the occupancy status of the items in the M storage locations, and the target storage location status includes the occupancy status; The location pair determination module is used to determine the two locations as a target location pair if there are two adjacent locations among the M locations, both of which are occupied. The item detection module is used to acquire an image of a region to be detected containing the target storage location pair in the target video frame, perform item detection on the image of the region to be detected, and obtain the item detection result corresponding to the image of the region to be detected. The anomaly alert module is used to generate an anomaly alert message for the target storage location pair when the item detection result is an item overlap result; the item overlap result is used to indicate that the target storage location pair is occupied by the same item.
14. A computer device, characterized in that, Including memory and processor; The memory is connected to the processor, the memory is used to store computer programs, and the processor is used to invoke the computer programs so that the computer device performs the method according to any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-12.
16. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the method described in any one of claims 1-12.
Citation Information
Patent Citations
Commodity status identification method and device, electronic device and readable storage medium
CN109446883A
Parking space state identification method and device based on video streaming
CN109817013A
Method and system for judging whether goods exist in goods allocation or not through vision
CN113313044A