Loading and unloading event detection method, device, computer equipment and storage medium

By performing object detection and optical flow processing on the event image sequence, combining optical image information and target position information, and using neural network models for loading and unloading event detection, the problem of low detection accuracy in the prior art is solved, and efficient automatic detection of loading and unloading event is achieved.

CN114596239BActive Publication Date: 2025-07-11SF TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011303391.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-19
Publication Date
2025-07-11
Estimated Expiration
2040-11-19

AI Technical Summary

Technical Problem

In the prior art, the method of judging loading and unloading events through trajectory requires complex trajectory information limitations, and is easily affected by imaging jitter and target moving route deviation, resulting in a decrease in detection accuracy.

Method used

By acquiring event image sequences, target detection and optical flow processing, combining optical image information and target position information, using a neural network model to detect loading and unloading events, fusing target position images and optical flow information images, improving detection accuracy.

Benefits of technology

Automatic loading and unloading event detection without complex trajectory information and camera jitter restrictions is realized, improving the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596239B_ABST
    Figure CN114596239B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for detecting loading and unloading events. The method includes: acquiring a collected sequence of event images; respectively performing object detection on each event image in the sequence of event images to determine the target position of the event target in each event image, and obtaining a target position image corresponding to the event image; performing optical flow processing on the sequence of event images to obtain an optical flow information image of each event image; respectively fusing each event image with the corresponding target position image and optical flow information image to obtain a fused image; and detecting loading and unloading events based on each fused image after fusion. Using this method can improve the accuracy of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, computer device, and storage medium for detecting loading and unloading events. Background Art

[0002] In the logistics industry, in order to improve the logistics operation efficiency, it is usually necessary to monitor the loading and unloading events at the transfer yard. By monitoring, the logistics status of each cargo box can be accurately understood, such as the start time and end time of loading and unloading, so as to statistically analyze the time consumed in each link of logistics to improve the logistics operation efficiency. In order to improve efficiency and reduce human consumption, the existing method has adopted a method of judging events by judging the trajectory of the target movement to replace the traditional method of manually recording events. Trajectory judgment mainly determines the event to which it belongs by judging whether the trajectory of the target movement coincides with a specific moving line.

[0003] However, the existing method of judging events by trajectory not only requires formulating complex trajectory information, but also requires restricting the camera to be fixed without jitter and restricting the moving route of the target. Once the camera jitters or the handling route deviates too much, the detection accuracy will be reduced. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for detecting loading and unloading events that can improve the detection accuracy.

[0005] A method for detecting loading and unloading events, the method comprising:

[0006] Obtaining a sequence of event images;

[0007] Performing object detection on each event image in the sequence of event images to determine the target position of the event target in each event image, and obtaining a target position image corresponding to the event image;

[0008] Performing optical flow processing on the sequence of event images to obtain an optical flow information image of each event image;

[0009] Fusing each event image with the corresponding target position image and the optical flow information image respectively to obtain a fused image;

[0010] Detecting loading and unloading events according to each fused image after fusion.

[0011] In one embodiment, the performing object detection on each event image in the sequence of event images to determine the target position of the event target in each event image, and obtaining a target position image corresponding to the event image, includes:

[0012] Input each of the event images into a trained object detection network for object detection, and output the target coordinate frames of each event target in the event image;

[0013] Use the center pixel point coordinates of the target coordinate frames of each event target in the event image as the target positions of each event target;

[0014] Mark the target positions of each event target in the event image on a blank image to obtain the target position image corresponding to the event image.

[0015] In one embodiment, the step of marking the target positions of each event target in the event image on a blank image to obtain the target position image corresponding to the event image includes:

[0016] Create a single-channel grayscale image to obtain a blank image;

[0017] Obtain the target class identifiers of each event target;

[0018] On the blank image, modify the pixel values at the same positions as the target positions of each event target to the target class identifiers corresponding to each event target to obtain the target position image corresponding to the event image.

[0019] In one embodiment, the step of performing optical flow processing on the event image sequence to obtain the optical flow information images of each event image includes:

[0020] Input the event image sequence into an optical flow algorithm interface for optical flow calculation to obtain the optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of pixels in the vertical direction;

[0021] Visualize the optical flow information to obtain an optical flow information image; the optical flow information image includes a first optical flow information image corresponding to the horizontal direction and a second optical flow information image corresponding to the vertical direction.

[0022] In one embodiment, the step of fusing each event image with the corresponding target position image and optical flow information image to obtain a fused image includes:

[0023] Perform image processing on the event image and the corresponding target position image and optical flow information image;

[0024] Stack and merge the event image and the corresponding target position image and optical flow information image after image processing in a preset order to obtain a fused image.

[0025] In one embodiment, the image processing of the event image, the corresponding target position image, and the optical flow information image includes:

[0026] Convert the event image into a grayscale image and scale it to a preset size;

[0027] Scale the target position image and the optical flow information image corresponding to the event image to the preset size respectively.

[0028] In one embodiment, the optical flow information image includes a first optical flow information image and a second optical flow information image;

[0029] Stack and merge the processed event image, the corresponding target position image, and the optical flow information image in a preset order to obtain a fused image, including:

[0030] Use the processed event image as the first layer, the processed target position image as the second layer, and the processed first optical flow information image and the second optical flow information image as the third layer and the fourth layer respectively;

[0031] Stack and merge the processed event image, the corresponding target position image, and the optical flow information image into layers in ascending order of layer numbers to obtain a fused image.

[0032] In one embodiment, the detection of loading and unloading events based on the fused fused images includes:

[0033] Call a trained neural network model, and the input channel number of the neural network model is four channels;

[0034] Input the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, loading event, or unloading event.

[0035] A loading and unloading event detection device, the device includes:

[0036] An acquisition module, configured to acquire an event image sequence;

[0037] A target detection module, configured to perform target detection on each event image in the event image sequence, determine the target position of the event target in each event image, and obtain the target position image corresponding to the event image;

[0038] An optical flow processing module, configured to perform optical flow processing on the event image sequence to obtain the optical flow information image of each event image;

[0039] A fusion module, configured to fuse each of the event images with the corresponding target position image and the optical flow information image respectively to obtain a fused image;

[0040] An event detection module, configured to detect loading and unloading events according to each of the fused images after fusion.

[0041] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the loading and unloading event detection method described in any one of the above are implemented.

[0042] A computer-readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the steps of the loading and unloading event detection method described in any one of the above are implemented.

[0043] For the above loading and unloading event detection method, device, computer device and storage medium, after obtaining an event image sequence, target detection and optical flow processing are respectively performed on each event image in the event image sequence to obtain corresponding target position images and optical flow information images; furthermore, each event image is respectively fused with the corresponding target position image and optical flow information for image fusion, and loading and unloading events are detected based on the obtained fused images. This method realizes automatic loading and unloading event detection by combining optical image information, optical flow information and target position information, thus eliminating the need to formulate complex trajectory information and being not limited by the shooting and target movement routes, thereby improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is an application environment diagram of the loading and unloading event detection method in an embodiment;

[0045] Figure 2 It is a flowchart of the loading and unloading event detection method in an embodiment;

[0046] Figure 3 It is a flowchart of the step of performing target detection on each event image in the event image sequence respectively in an embodiment, determining the target position of the event target in each event image, and obtaining the target position image corresponding to the event image;

[0047] Figure 4 It is a visualization diagram of the target rectangular box in an embodiment;

[0048] Figure 5 It is a schematic diagram of marking an event target with a blank image in an embodiment;

[0049] Figures 6a - 6b It is a schematic diagram of the optical flow information image in an embodiment;

[0050] Figure 7 Schematic flowchart of the step of fusing each event image with the corresponding target position image and optical flow information image respectively to obtain a fused image in an embodiment;

[0051] Figure 8 Schematic diagram of a fused image in an embodiment;

[0052] Figure 9 Structural block diagram of a loading and unloading event detection device in an embodiment;

[0053] Figure 10 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] The loading and unloading event detection method provided by the present application can be applied to an application environment as shown in Figure 1 , including a camera device 102, a terminal 104 and a server 106. Among them, the camera device 102 communicates with the terminal 104 through a network, and the terminal 104 communicates with the server 106 through a network. After the camera device 102 captures a sequence of event images and sends them to the terminal 104, the terminal 104 can implement the loading and unloading event detection method alone according to the sequence of event images. Alternatively, the terminal 104 can send the sequence of event images to the server 106, and the server 106 can implement the loading and unloading event detection method.

[0056] Specifically, taking the server 106 as an example, the server 106 obtains the sequence of event images captured by the camera device 102 from the terminal 104; the server 106 performs target detection on each event image in the sequence of event images to determine the target position of the event target in each event image, and obtains the target position image corresponding to the event image; the server 106 performs optical flow processing on the sequence of event images to obtain the optical flow information image of each event image; the server 106 fuses each event image with the corresponding target position image and optical flow information image respectively to obtain a fused image; the server 106 performs loading and unloading event detection according to each fused image after fusion. Among them, the camera device 102 can be but is not limited to a camera, a video camera, and various devices with a camera. The terminal 104 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server 106 can be implemented by an independent server or a server cluster composed of multiple servers.

[0057] In one embodiment, as shown inFigure 2 As shown, a method for detecting loading and unloading events is provided. Taking the server in Figure 1 as an example for illustration, the method includes the following steps:

[0058] Step S202: Obtain the collected sequence of event images.

[0059] Among them, the sequence of event images refers to a sequence including multiple consecutive event images, and each event image includes an event target corresponding to the event. The event targets included in different event images are different. Assuming that the event included in the event image is a loading and unloading event, the event targets included in the event image are not limited to transportation vehicles (trucks, airplanes, etc.), pallets, loaders, trailers, platform trucks, etc.

[0060] Specifically, corresponding camera devices are set at the event location, and the event location is continuously imaged by the camera devices to obtain a sequence of event images. Then, the camera devices send the sequence of event images to the terminal device, and the terminal device sends it to the server. It should be understood that the camera devices can also take pictures of the event location and then send a video stream obtained by shooting to the terminal. The terminal can process the video stream to obtain a sequence of event images including consecutive event images and then send it to the server. It can also directly send the video stream as a sequence of event images to the server, and the server processes the video stream by itself after receiving it to obtain the corresponding sequence of event images. In addition, for a terminal that does not communicate with the camera devices, the sequence of event images can also be obtained by manual upload or by being sent by a terminal device that communicates with the camera devices.

[0061] Step S204: Perform object detection on each event image in the sequence of event images, determine the target positions of the event targets in each event image, and obtain a target position image corresponding to the event image.

[0062] Among them, object detection refers to detecting and extracting the object in the event image to obtain the target position of the event target in the event image.

[0063] Specifically, after the server obtains the sequence of event images, it calls a trained object detection network. Each event image in the sequence of event images is input into the object detection network for object detection, and the object detection network outputs the target positions of the event targets included in the event image, obtaining an image with the target positions of the event targets marked, that is, obtaining a target position image.

[0064] The object detection network is a neural network that is pre-trained to detect event objects for a specific event. Since this embodiment is used to detect loading and unloading events, the object detection network is a neural network that detects event objects in loading and unloading events, including but not limited to models such as YOLO (You Only Look Once, unified real-time object detection) and R-CNN (Region-Convolutional Neural Networks). That is to say, in this embodiment, the event objects of the loading and unloading events are used as the network training objectives, and models such as YOLO and RCNN are trained to obtain the object detection network required in this embodiment. The training set for training the neural network annotates the positions of the event objects and the categories to which the event objects belong in the form of rectangular frames. Suppose, taking the event objects of the loading and unloading events as an example, the target categories to be annotated include all participants or entities in the loading and unloading events, usually including transportation vehicles (trucks, airplanes), cargo boxes, loaders, platform trucks, trailers, etc.

[0065] Step S206: Perform optical flow processing on the event image sequence to obtain the optical flow information images of each event image.

[0066] Among them, optical flow processing refers to the process of processing images using optical flow algorithms. It is a method that uses the changes of pixels in the time domain in the image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, so as to calculate the motion information of objects between adjacent frames.

[0067] Specifically, since the optical flow method processes images using an image sequence, after the server obtains the event image sequence, it uses the optical flow algorithm to perform optical flow processing on the event image sequence to obtain the optical flow information images of each event image. Any existing optical flow algorithm can be used, and in this embodiment, the TVL1 (Total Variation L1, total variation based on the L1 norm) optical flow algorithm is preferably used.

[0068] Step S208: Fuse each event image with the corresponding target position image and optical flow information image to obtain a fused image.

[0069] Specifically, after the server obtains the target position images and optical flow information images corresponding to each event image through object detection and optical flow processing respectively, it combines the optical image information, optical flow information, and target position information. The event image is fused with the corresponding target position image and optical flow information image. For example, the event image is merged or added to the corresponding target position image and optical flow information image to obtain the fused image after fusion.

[0070] Step S210: Detect loading and unloading events based on the fused images after fusion.

[0071] Specifically, after the server obtains the fused image, since the fused image combines optical image information, optical flow information, and target position information at the same time. Therefore, the fused image is used to detect the loading and unloading events, thereby improving the detection accuracy. The detection of loading and unloading events can also use a trained neural network. The fused image is input into the neural network for event detection. The neural network outputs the detection result, for example, the event image does not include loading and unloading events, or includes loading events / unloading events.

[0072] In the above loading and unloading event detection method, after obtaining the event image sequence, target detection and optical flow processing are respectively performed on each event image in the event image sequence to obtain the corresponding target position image and optical flow information image; and then each event image is respectively fused with the corresponding target position image and optical flow information, and the obtained fused image is used to detect the loading and unloading events. This method realizes automatic detection of loading and unloading events by combining optical image information, optical flow information, and target position information, thus eliminating the need to formulate complex trajectory information and being unaffected by the limitations of the camera and target movement routes, thereby improving the detection accuracy.

[0073] In one embodiment, as Figure 3 shown, step S204 includes the following steps:

[0074] Step S302, input each event image into a trained target detection network for target detection, and output the target coordinate frames of each event target in the event image.

[0075] Among them, the target coordinate frame is a rectangular frame composed of the coordinates of the image area where the event target is located, and the position of the event target in the event image is marked by the target rectangular frame.

[0076] Specifically, each event image is input into a trained target detection network for target detection, and the target detection network outputs the target detection frames of each event target in the event image. Refer to Figure 4 , a visualization schematic diagram of the target rectangular frame is provided. Figure 4 That is, the image after visualizing the detection result of the target detection network, and the black rectangular frame therein is the target rectangular frame of each event target.

[0077] Step S304, use the center pixel point coordinates of the target coordinate frame of each event target in the event image as the target position of each event target.

[0078] Among them, the center pixel point coordinates refer to the coordinates of the center pixel point of the target coordinate frame.

[0079] Specifically, after obtaining the target coordinate frames of each event target in the event image through target detection, since there is usually inevitable overlap in the regions between the target coordinate frames, especially when the event targets are close. For example Figure 4 the target coordinate frame of the airplane in Figure 4 completely overlaps with the target coordinate frame of the cargo box, and the target coordinate frame of the loader on the right side of the image partially overlaps with the target coordinate frame of the cargo box. Therefore, in order to avoid overlapping of the target coordinate frame regions and reduce the detection accuracy, the center pixel point coordinates of the target coordinate frames of each event target are used as the target positions of each event target.

[0080] Step S306, mark the target positions of each event target in the event image on a blank image to obtain a target position image corresponding to the event image.

[0081] Among them, the blank image refers to an image that does not include any objects. Specifically, a blank image is created, and then the corresponding positions of the determined target positions of each event target are marked in the blank image to represent each event target in the blank image.

[0082] In one embodiment, step S306 includes: creating a single-channel grayscale image to obtain a blank image; obtaining the target category identifiers of each event target; on the blank image, modifying the pixel values at the same positions as the target positions of each event target to the target category identifiers corresponding to each event target to obtain a target position image corresponding to the event image.

[0083] Specifically, first, a single-channel grayscale image is created, that is, an image with all pixel values being 0 is created to obtain a blank image. If only the event targets are marked at the corresponding positions in the blank image, it is not sufficient to distinguish the categories of each event target. Therefore, before marking on the blank image, the target category identifiers to which each event target belongs are obtained. For example, 0 represents no target, 1 represents a transportation vehicle, 2 represents a cargo box, 3 represents a loader, etc., and the specific identifiers can be set according to the actual situation. Then, on the blank image, the pixel values at the positions corresponding to the target positions of each event target are modified to the target category identifiers of the event target.

[0084] As Figure 5 shown, a schematic diagram of marking event targets on a blank image is provided. Figure 5 Take Figure 4 as an example of the schematic diagram, Figure 4 the event targets included are an airplane, two cargo boxes, and two loaders. And the target category identifier for the transportation vehicle is 1, for the cargo box is 2, and for the loader is 3. Therefore, Figure 5 the pixel values of five corresponding pixel points in Figure 5 are modified to the target category identifiers. According to Figure 4The target positions of the event targets in [description], from left to right, the pixel positions of these five modified pixel values are 3, 1, 2, 2, and 3 respectively. Thus, the result information of the target detection is completely saved in a single-channel image. It should be understood that for the sake of Figure 5 For illustration, each pixel with a value in the figure is represented by a circle, and the number in the circle represents the pixel value of that pixel point. Except for the pixels whose pixel values are modified to the target category identifier, the pixel values of other parts are all 0.

[0085] In this embodiment, by using the center point as the target position of the event target and marking each event target at the corresponding position on the blank image by modifying the pixel values, it is possible to avoid target overlap and the influence of targets unrelated to the event, thereby improving the accuracy of detection.

[0086] In one embodiment, step S206 includes: inputting the event image sequence into the optical flow algorithm interface for optical flow calculation to obtain the optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of pixels in the vertical direction; visualizing the optical flow information to obtain the optical flow information image; the optical flow information image includes the first optical flow information image corresponding to the horizontal direction and the second optical flow information image corresponding to the vertical direction.

[0087] Among them, the optical flow algorithm interface refers to an interface encapsulating the optical flow algorithm, and through this interface, the optical flow algorithm can be called for optical flow calculation processing.

[0088] Specifically, the server inputs the event image sequence into the optical flow algorithm interface, and the optical flow algorithm corresponding to the optical flow algorithm interface performs optical flow calculation processing on each event image in the event image sequence, thereby obtaining the optical flow information corresponding to each event image. Then, in order to represent the optical flow information of the event image in an image, the optical flow information is visualized to obtain the optical flow information image. However, since optical flow is a method for calculating the motion information of objects between adjacent frames. And for an image, the motion of pixels is divided into the horizontal direction and the vertical direction. Therefore, the optical flow information obtained by optical flow calculation includes two parts. One part is the U channel, which represents the motion information of pixels in the horizontal direction. The other part is the V channel, which represents the motion information of pixels in the vertical direction. So, the visualized optical flow information image includes the first optical flow information image corresponding to the horizontal direction and the second optical flow information image corresponding to the vertical direction.

[0089] As Figures 6a - 6b shown, a schematic diagram of the optical flow information image is provided. Figures 6a - 6b In [description], the uniformly gray part indicates that the object in the image is basically stationary. Figure 6a The [description] is the U channel, which represents the motion information of each region in the image in the horizontal direction, that is, the first optical flow information image. Figure 6aThe part darker than the average value indicates that the object at this position is moving to the left, and the part brighter than the average value indicates that the object at this position is moving to the right. Figure 6b is the V channel, that is, the second optical flow information image. Figure 6b The part darker than the average value indicates that the object at this position is moving upward, and the part brighter than the average value indicates that the object at this position is moving downward. In this embodiment Figures 6a - 6b is also Figure 4 the corresponding optical flow information image, refer to Figure 4 , because Figure 4 the cargo box near the aircraft cabin in [the image] is moving to the right. Therefore, Figure 6a the pixel values in the corresponding area of the U channel corresponding to the aircraft cabin door in [the image] are brighter. Also, because there is basically no movement in the vertical direction, there is no obvious pattern in the corresponding area of the U channel, only some interference information.

[0090] In one embodiment, as Figure 7 shown, step S208 includes the following steps:

[0091] Step S702, perform image processing on the event image, the corresponding target position image, and the optical flow information image.

[0092] Specifically, in order to make the event images to be fused, as well as the target position image and the optical flow information image corresponding to the event image, be fused more fittingly. Image processing is uniformly performed on the event image, the corresponding target position image, and the optical flow information image, and any existing one or more image processing methods can be used for the image processing.

[0093] In one embodiment, step S702 includes: converting the event image into a grayscale image and scaling it to a preset size; scaling the target position image and the optical flow information image corresponding to the event image to the preset size respectively.

[0094] Specifically, since the target position image and the optical flow information image are grayscale images, while the event image is a color image directly captured by the imaging device, the event image is uniformly converted into a grayscale image. Then, the event image, the target position image, and the optical flow information image, all of which are grayscale images, are uniformly scaled in size to obtain event images, target position images, and optical flow information images of the same size. For example, the width and height of the event image, the target position image, and the optical flow information image are uniformly scaled to the preset size W*H, and the preset width W and height H can be set according to specific tasks as required, which is not limited here.

[0095] Step S704, stack and merge the processed event image, the corresponding target position image, and the optical flow information image in a preset order to obtain a fused image.

[0096] Specifically, after performing image processing on the event image, the target position image, and the optical flow information image, the processed event image, the corresponding target position image, and the optical flow information image are superimposed and merged in a set preset order to obtain a fused image. The preset order can be set according to actual requirements. For example, from bottom to top, they can be the event image, the target position image, and the optical flow information image respectively.

[0097] In one embodiment, step S704 includes: taking the processed event image as the first layer, taking the processed target position image as the second layer, and taking the processed first optical flow information image and the second optical flow information image as the third layer and the fourth layer respectively; superimposing and merging the processed event image, the corresponding target position image, and the optical flow information image into layers in ascending order of the layer numbers to obtain a fused image.

[0098] Specifically, since the fused image in this embodiment is obtained by superimposing and merging, the fused image can be understood as a layer. When performing the superimposing and merging, according to the preset order, the processed event image is taken as the first layer, that is, layer 1 after the superimposing and merging. The processed target position image is taken as the second layer, that is, layer 2 after the superimposing and merging. And, since the optical flow information image includes the first optical flow information image and the second optical flow information image, they can be taken as the third layer and the fourth layer respectively, that is, layer 3 and layer 4 after the superimposing and merging. Then, according to the ascending order, the first layer, the second layer, the third layer, and the fourth layer are superimposed in turn to obtain the fused image of this embodiment. The fused image in this embodiment is as Figure 8 shown. Refer to Figure 8 , the event image is at the bottom as the first layer, followed by the target position image, the first optical flow information image, and the second optical flow information image in turn.

[0099] In this embodiment, before performing the loading and unloading detection, the event image, the target position image, and the optical flow information image are superimposed and merged to ensure that the image for performing the loading and unloading detection is an image that combines optical image information, optical flow information, and target position information at the same time, thereby improving the accuracy of the detection.

[0100] In one embodiment, step S210 includes: calling a trained neural network model, and the number of input channels of the neural network model is four channels; inputting the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, loading event, or unloading event.

[0101] Specifically, after the server obtains the fused image, it calls the trained neural network for loading and unloading event detection and classification. The fused image is input into the neural network to obtain the detection results of the event, including no event, loading event, and unloading event. No event means that there is no loading event and no unloading event in the image. Among them, the neural network in this embodiment can be any one or more of neural network models such as Resnet and EfficientNet. However, since the fused image used in this embodiment is a superposition and combination of four images, after selecting the neural network, the number of input channels of the neural network should be changed to 4 channels to adapt to the fused image. In this embodiment, event detection is performed on the fused image through the neural network, which can improve the accuracy of detection.

[0102] It should be understood that although Figure 2 , 3 , and each step in the flowchart of 7 is displayed in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2 , 3 , and at least a part of the steps in 7 may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of the steps or stages in other steps or other steps.

[0103] In one embodiment, as Figure 9 shown, a loading and unloading event detection device is provided, including: an acquisition module 902, a target detection module 904, an optical flow processing module 906, a fusion module 908, and an event detection module 910, where:

[0104] The acquisition module 902 is used to acquire an event image sequence.

[0105] The target detection module 904 is used to perform target detection on each event image in the event image sequence, determine the target position of the event target in each event image, and obtain the target position image corresponding to the event image.

[0106] The optical flow processing module 906 is used to perform optical flow processing on the event image sequence to obtain the optical flow information image of each event image.

[0107] The fusion module 908 is used to fuse each event image with the corresponding target position image and optical flow information image respectively to obtain a fused image.

[0108] An event detection module 910 for detecting loading and unloading events based on each fused image after fusion.

[0109] In one embodiment, the target detection module 904 is further configured to input each event image into a trained target detection network for target detection, and output the target coordinate frames of each event target in the event image; use the center pixel point coordinates of the target coordinate frames of each event target in the event image as the target positions of each event target; mark the target positions of each event target in the event image on a blank image to obtain a target position image corresponding to the event image.

[0110] In one embodiment, the target detection module 904 is further configured to create a single-channel grayscale image to obtain a blank image; obtain the target class identifiers of each event target; on the blank image, modify the pixel values at the same positions as the target positions of each event target to the target class identifiers corresponding to each event target to obtain a target position image corresponding to the event image.

[0111] In one embodiment, the optical flow processing module 906 is further configured to input the event image sequence into an optical flow algorithm interface for optical flow calculation to obtain the optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of pixels in the vertical direction; visualize the optical flow information to obtain an optical flow information image; the optical flow information image includes a first optical flow information image corresponding to the horizontal direction and a second optical flow information image corresponding to the vertical direction.

[0112] In one embodiment, the fusion module 908 is further configured to perform image processing on the event image, the corresponding target position image, and the optical flow information image; stack and merge the event image, the corresponding target position image, and the optical flow information image after image processing in a preset order to obtain a fused image.

[0113] In one embodiment, the fusion module 908 is further configured to convert the event image into a grayscale image and scale it to a preset size; scale the target position image and the optical flow information image corresponding to the event image to the preset size respectively.

[0114] In one embodiment, the fusion module 908 is further configured to use the event image after image processing as the first layer, the target position image after image processing as the second layer, and the first optical flow information image and the second optical flow information image after image processing as the third layer and the fourth layer respectively; stack and merge the event image, the corresponding target position image, and the optical flow information image after image processing into layers in ascending order of layer numbers to obtain a fused image.

[0115] In one embodiment, the event detection module 910 is further configured to call a trained neural network model, where the number of input channels of the neural network model is four channels; input the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, a loading event, or an unloading event.

[0116] For the specific limitations of the loading and unloading event detection device, reference can be made to the limitations of the loading and unloading event detection method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned loading and unloading event detection device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0117] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a loading and unloading event detection method.

[0118] Those skilled in the art can understand that Figure 10 the structure shown in

[0119] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0120] Obtain the collected event image sequence;

[0121] Perform target detection on each event image in the event image sequence to determine the target position of the event target in each event image, and obtain the target position image corresponding to the event image;

[0122] Perform optical flow processing on the event image sequence to obtain the optical flow information image of each event image;

[0123] Fuse each event image with the corresponding target position image and optical flow information image respectively to obtain a fused image;

[0124] Detect loading and unloading events based on each fused image after fusion.

[0125] In one embodiment, when the processor executes the computer program, the following steps are further implemented: input each event image into a trained target detection network for target detection, and output the target coordinate frames of each event target in the event image; use the center pixel point coordinates of the target coordinate frames of each event target in the event image as the target positions of each event target; mark the target positions of each event target in the event image on a blank image to obtain the target position image corresponding to the event image.

[0126] In one embodiment, when the processor executes the computer program, the following steps are further implemented: create a single-channel grayscale image to obtain a blank image; obtain the target category identifiers of each event target; on the blank image, modify the pixel values at the same positions as the target positions of each event target to the target category identifiers corresponding to each event target to obtain the target position image corresponding to the event image.

[0127] In one embodiment, when the processor executes the computer program, the following steps are further implemented: input the event image sequence into an optical flow algorithm interface for optical flow calculation to obtain the optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of pixels in the vertical direction; visualize the optical flow information to obtain the optical flow information image; the optical flow information image includes a first optical flow information image corresponding to the horizontal direction and a second optical flow information image corresponding to the vertical direction.

[0128] In one embodiment, when the processor executes the computer program, the following steps are further implemented: perform image processing on the event image and the corresponding target position image and optical flow information image; stack and merge the event image and the corresponding target position image and optical flow information image after image processing in a preset order to obtain a fused image.

[0129] In one embodiment, when the processor executes the computer program, the following steps are further implemented: convert the event image into a grayscale image and scale it to a preset size; scale the target position image and the optical flow information image corresponding to the event image to the preset size respectively.

[0130] In one embodiment, when the processor executes the computer program, the following steps are further implemented: using the event image after image processing as the first layer, the target position image after image processing as the second layer, and the first optical flow information image and the second optical flow information image after image processing as the third layer and the fourth layer respectively; stacking and merging the event image after image processing and the corresponding target position image and optical flow information image into layers in ascending order of layer numbers to obtain a fused image.

[0131] In one embodiment, when the processor executes the computer program, the following steps are further implemented: calling a trained neural network model, where the number of input channels of the neural network model is four channels; inputting the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, loading event, or unloading event.

[0132] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0133] Obtaining a collected sequence of event images;

[0134] Performing object detection on each event image in the sequence of event images respectively to determine the target position of the event target in each event image, and obtaining the target position image corresponding to the event image;

[0135] Performing optical flow processing on the sequence of event images to obtain the optical flow information image of each event image;

[0136] Fusing each event image with the corresponding target position image and optical flow information image respectively to obtain a fused image;

[0137] Performing loading and unloading event detection based on each fused image after fusion.

[0138] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: inputting each event image into a trained object detection network for object detection respectively, and outputting the target coordinate frames of each event target in the event image; using the center pixel point coordinates of the target coordinate frames of each event target in the event image as the target position of each event target; marking the target position of each event target in the event image on a blank image to obtain the target position image corresponding to the event image.

[0139] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: creating a single-channel grayscale image to obtain a blank image; obtaining the target category identifier of each event target; on the blank image, modifying the pixel value at the position same as the target position of each event target to the target category identifier corresponding to each event target to obtain the target position image corresponding to the event image.

[0140] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: inputting a sequence of event images into an optical flow algorithm interface for optical flow calculation to obtain optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of pixels in the vertical direction; visualizing the optical flow information to obtain an optical flow information image; the optical flow information image includes a first optical flow information image corresponding to the horizontal direction and a second optical flow information image corresponding to the vertical direction.

[0141] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing image processing on the event images, the corresponding target position images, and the optical flow information images; superimposing and merging the event images, the corresponding target position images, and the optical flow information images that have undergone image processing in a preset order to obtain a fused image.

[0142] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: converting the event images into grayscale images and scaling them to a preset size; scaling the corresponding target position images and optical flow information images of the event images to the preset size respectively.

[0143] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: using the event images after image processing as the first layer, the target position images after image processing as the second layer, and the first optical flow information image and the second optical flow information image after image processing as the third layer and the fourth layer respectively; superimposing and merging the event images, the corresponding target position images, and the optical flow information images that have undergone image processing into layers in ascending order of layer numbers to obtain a fused image.

[0144] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: calling a trained neural network model, the number of input channels of the neural network model being four channels; inputting the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, a loading event, or an unloading event.

[0145] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0147] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for detecting loading and unloading events, characterized in that, The method includes: Obtaining an event image sequence; the event image sequence refers to a sequence including multiple consecutive event images, and each event image includes an event target corresponding to the event; Inputting each of the event images into a trained object detection network for object detection, and outputting the target coordinate frames of each event target in the event image; taking the center pixel point coordinates of the target coordinate frames of each event target in the event image as the target positions of each event target; marking the target positions of each event target in the event image on a blank image to obtain a target position image corresponding to the event image; Performing optical flow processing on the event image sequence to obtain an optical flow information image of each event image; Performing image processing on the event image and the corresponding target position image and optical flow information image; stacking and merging the event image and the corresponding target position image and optical flow information image after image processing in a preset order to obtain a fused image; Detecting loading and unloading events based on each of the fused images after fusion.

2. The method according to claim 1, characterized in that The step of marking the target positions of each event target in the event image on a blank image to obtain a target position image corresponding to the event image includes: Creating a single-channel grayscale image to obtain a blank image; Obtaining the target category identifiers of each event target; On the blank image, modifying the pixel values at positions identical to the target positions of each event target to the target category identifiers corresponding to each event target to obtain a target position image corresponding to the event image.

3. The method according to claim 1, characterized in that, The step of performing optical flow processing on the event image sequence to obtain an optical flow information image of each event image includes: Inputting the event image sequence into an optical flow algorithm interface for optical flow calculation to obtain the optical flow information of each event image; the optical flow information includes the motion information of pixels in the horizontal direction and the motion information of the pixels in the vertical direction; Visualizing the optical flow information to obtain an optical flow information image; the optical flow information image includes a first optical flow information image corresponding to the horizontal direction and a second optical flow information image corresponding to the vertical direction.

4. The method according to claim 1, wherein The step of performing image processing on the event image and the corresponding target position image and optical flow information image includes: Converting the event image into a grayscale image and scaling it to a preset size; Scaling the target position image and the optical flow information image corresponding to the event image to the preset size respectively.

5. The method according to claim 1, characterized in that The optical flow information image includes a first optical flow information image and a second optical flow information image; The step of stacking and merging the event image and the corresponding target position image and optical flow information image after image processing in a preset order to obtain a fused image includes: Taking the event image after image processing as the first layer, taking the target position image after image processing as the second layer, and taking the first optical flow information image and the second optical flow information image after image processing as the third layer and the fourth layer respectively; Stack and merge the processed event images, the corresponding target position images, and the optical flow information images into layers in ascending order of the number of layers to obtain a fused image.

6. The method according to claim 1, wherein Perform loading and unloading event detection based on each of the fused images, including: Invoke a trained neural network model, where the number of input channels of the neural network model is four channels; Input the fused image into the neural network model with four input channels for loading and unloading event detection to obtain a detection result; the detection result includes no event, a loading event, or an unloading event.

7. A loading and unloading event detection device, characterized in that, The device includes: An acquisition module for acquiring a sequence of event images; the sequence of event images refers to a sequence including multiple consecutive event images, and each event image includes an event target corresponding to the event; A target detection module for inputting each of the event images into a trained target detection network for target detection, and outputting a target coordinate box for each event target in the event image; taking the center pixel point coordinates of the target coordinate box of each event target in the event image as the target position of each event target; marking the target positions of each event target in the event image on a blank image to obtain the target position image corresponding to the event image; An optical flow processing module for performing optical flow processing on the sequence of event images to obtain the optical flow information image of each event image; A fusion module for performing image processing on the event image, the corresponding target position image, and the optical flow information image; stacking and merging the processed event image, the corresponding target position image, and the optical flow information image in a preset order to obtain a fused image; An event detection module for performing loading and unloading event detection based on each of the fused images.

8. The device according to claim 7, characterized in that, The target detection module is further configured to create a single-channel grayscale image to obtain a blank image; acquire the target category identifier of each event target; modify the pixel value at the position identical to the target position of each event target on the blank image to the target category identifier corresponding to each event target to obtain the target position image corresponding to the event image.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multiobject fusion module for collision preparation system

    CN101837782A

  • Sudden abnormal event intelligent identification alarm device and system

    CN103839373A