Image processing method and device, computer equipment and storage medium
By calculating the frame difference of continuous image frames during industrial equipment transportation and building masks, fusing and cropping images, the problem of low image processing efficiency in the prior art is solved, and the efficiency of collision prediction is improved.
Patent Information
- Application Number
- CN202311870131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-30
- Publication Date
- 2025-07-01
AI Technical Summary
During the transportation of industrial equipment, it is difficult for the prior art to efficiently process a large number of image frames, resulting in inefficient collision prediction.
By acquiring at least two consecutive target video frame images, calculating the frame difference, building an image mask, and performing image fusion and cropping based on the mask, a cropped image is generated for collision prediction.
This method improves the efficiency of collision prediction by reducing the amount of data and computing amount of image processing and reduces the computing resource consumption during the collision prediction process of transportation equipment.
Smart Images

Figure CN120235769A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of industrial equipment detection, and particularly to an image processing method and apparatus, a computer device, and a storage medium. Background Art
[0002] At present, in the industrial field, a large number of transportation devices have been added to facilitate the transportation of goods, so as to complete the transportation of goods in the warehouse through the transportation devices. However, if the transportation device collides during transportation, the goods will be damaged, and in severe cases, there will be a safety hazard of the transportation device tipping over. In the related art, in order to reduce the safety hazard during the transportation of the transportation device, a large number of image frames of the transportation device during the entire transportation process are collected, and the large number of image frames are directly processed for collision prediction. If the resolution of each image frame is large, it will increase the burden of image processing, thereby affecting the efficiency of image processing. Therefore, how to improve the efficiency of image processing has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose an image processing method and apparatus, a computer device, and a storage medium, aiming to improve the image processing efficiency during the collision prediction process, and thereby improve the collision prediction efficiency.
[0004] To achieve the above object, a first aspect of the embodiments of this application proposes an image processing method, and the method includes:
[0005] Obtain at least two consecutive target video frame images; wherein, the target video frame images are obtained by photographing a transportation device;
[0006] Obtain the frame difference between two adjacent target video frame images to obtain an image frame difference;
[0007] Construct a mask for each image frame difference to obtain an image mask;
[0008] Perform splicing processing on the image masks to obtain a target mask;
[0009] If the target mask is greater than a preset pixel point threshold, perform fusion processing on at least two target video frame images according to the target mask to obtain a fusion image; wherein, the fusion image includes a highlighted area;
[0010] Perform cropping processing on the fusion image according to the highlighted area to obtain a cropped image;
[0011] Perform collision prediction on the transportation device according to the cropped image.
[0012] In some embodiments, the constructing a mask for each image frame difference to obtain an image mask includes:
[0013] Perform binaryzation processing on the image frame difference to obtain an image binary matrix;
[0014] Perform erosion processing on the image binary matrix according to a preset erosion matrix to obtain an image erosion matrix;
[0015] Perform dilation processing on the image erosion matrix according to a preset number of dilation times and a preset dilation matrix to obtain the image mask.
[0016] In some embodiments, if the target mask is greater than a preset pixel threshold, fusing at least two frames of the target video frame images according to the target mask to obtain a fused image, including:
[0017] If the target mask is greater than the pixel threshold, obtain the pixel points of the target video frame image to obtain image pixel points;
[0018] Perform screening processing on a preset fusion weight according to the target mask to obtain a target weight for each of the image pixel points;
[0019] Construct the fused image according to the target weight and the image pixel points.
[0020] In some embodiments, the cropping the fused image according to the highlighted area to obtain a cropped image includes:
[0021] Perform annotation processing on the highlighted area to obtain an annotated area;
[0022] Perform outward expansion processing on the annotated area according to a preset outward expansion index to obtain a selected area;
[0023] Crop the fused image according to the selected area to obtain the cropped image.
[0024] In some embodiments, the collision prediction for the transportation device according to the cropped image includes:
[0025] Perform object recognition on the cropped image to obtain an image object category;
[0026] Use the cropped image with the image object category being the transportation device as a selected image;
[0027] Perform collision prediction for the transportation device according to the selected image.
[0028] In some embodiments, the collision prediction for the transportation device according to the selected image includes:
[0029] Obtain the distance between two transportation devices according to the selected image to obtain a device distance;
[0030] Perform collision prediction on the transportation device according to the device spacing and a preset spacing range.
[0031] In some embodiments, the obtaining of at least two consecutive target video frame images includes:
[0032] Obtain target video data of the transportation device;
[0033] Perform motion detection on the target video data to obtain motion detection information;
[0034] Select a video segment from the target video data according to the motion detection information to obtain selected video segment data;
[0035] Perform frame splitting on the selected video segment data to obtain original video frame images;
[0036] Obtain at least two consecutive target video frame images from the original video frame images.
[0037] To achieve the above object, a second aspect of the embodiments of the present application proposes an image processing device, and the device includes:
[0038] An image acquisition module, configured to obtain at least two consecutive target video frame images; wherein, the target video frame images are obtained by photographing a transportation device;
[0039] A frame difference acquisition module, configured to obtain a frame difference between two adjacent target video frame images to obtain an image frame difference;
[0040] A mask construction module, configured to perform mask construction on each of the image frame differences to obtain an image mask;
[0041] A mask splicing module, configured to splice the image masks to obtain a target mask;
[0042] An image fusion module, configured to, if the target mask is greater than a preset pixel point threshold, fuse at least two target video frame images according to the target mask to obtain a fused image; wherein, the fused image includes a highlighted area;
[0043] An image cropping module, configured to crop the fused image according to the highlighted area to obtain a cropped image;
[0044] A collision prediction module, configured to perform collision prediction on the transportation device according to the cropped image.
[0045] To achieve the above object, a third aspect of the embodiments of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0046] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0047] The image processing method, device, computer device and storage medium provided by the present application analyze the amplitude of the movement of the transportation device through image masking. If the image mask of multiple video frame images is greater than a preset pixel point threshold, indicating that the movement amplitude of the transportation device is too large, then multiple video frame images are fused into a fused image, and the highlighted area in the fused image is cropped out as a cropped image, and the collision of the transportation device is predicted based on the cropped image. By judging that there is a large range of motion areas in multiple image frames, multiple video frame images are fused into a fused image, and then the motion area of the fused image is cropped out as a cropped image. By performing collision prediction on the cropped image, the computational amount of image processing in the collision prediction process is reduced, thereby improving the collision prediction efficiency. Description of the Drawings
[0048] Figure 1 is a flowchart of the image processing method provided by the embodiments of the present application;
[0049] Figure 2 is Figure 1 a flowchart of step S101 in
[0050] Figure 3 is Figure 1 a flowchart of step S103 in
[0051] Figure 4 is Figure 1 a flowchart of step S105 in
[0052] Figure 5 is a schematic diagram of the fused image in the image processing method provided by the embodiments of the present application;
[0053] Figure 6 is Figure 1 a flowchart of step S106 in
[0054] Figure 7 is a schematic diagram of the cropped image in the image processing method provided by the embodiments of the present application;
[0055] Figure 8 is Figure 1Flowchart of step S107 in
[0056] Figure 9 is Figure 8 Flowchart of step S803 in
[0057] Figure 10 Schematic structural diagram of the image processing device provided by the embodiments of the present application;
[0058] Figure 11 Schematic hardware structure diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0062] First, several terms involved in the present application are analyzed:
[0063] Mask: In image processing, it refers to a method for specifying certain regions of an image for specific operations. A mask is usually a binary image of the same size as the original image, where white pixels represent the regions that need to be operated on, and black pixels represent the regions that do not need to be operated on. In image processing, masks can be used for various purposes. For example, by creating a mask to extract the regions of interest in an image, and then performing specific operations on these regions, such as filtering, edge detection, or color conversion. Masks can also be used for image synthesis, mixing specific regions of one image with another image.
[0064] Dilation operation: It can expand the object regions in an image by expanding the boundaries of the objects outward, thereby making the objects larger. This is very useful for removing small noise or filling object holes.
[0065] Erosion operation: Opposite to dilation, it can shrink the object area in the image. By shrinking the boundaries of the object inward, sharp parts of the object can be removed or adjacent objects can be separated.
[0066] Image recognition: It is to recognize the objects in an image, that is, to label and classify each object that appears in the image. Different from image classification, the image recognition task needs to distinguish and classify each object, rather than classifying the entire image. For example, identifying multiple objects such as cats, dogs, and cars in an image. Image recognition usually refers to multi-label classification, that is, each picture may belong to multiple categories.
[0067] Image fusion: It refers to processing the image data of the same target collected from multiple source channels through image processing and computer technology, etc., to extract the beneficial information in each channel to the greatest extent, and finally synthesize high-quality images to improve the utilization rate of image information, improve the accuracy and reliability of computer interpretation, enhance the spatial resolution and spectral resolution of the original image, and facilitate monitoring.
[0068] With the development of logistics, the transportation of goods in the warehouse has changed from manual handling to intelligent transportation. In order to increase the transportation volume of goods, a large number of transportation devices have been added to complete the transportation of goods in the warehouse through the transportation devices. However, during the transportation of goods, the transportation devices may collide due to improper operation, and when the collision is serious, the devices may tip over, thus causing potential safety hazards. In the prior art, in order to predict collisions of transportation devices in advance and predict the possibility of collisions of transportation devices in advance for early warning. However, to predict collisions of transportation devices, it is necessary to collect multiple video frame images of the transportation devices during the entire transportation process and directly detect the multiple video frame images to infer the collision possibility of the transportation devices. However, directly processing multiple video frame images, if the resolution of each video frame image is large, will increase the burden of collision prediction and affect the efficiency of image processing, and thus affect the efficiency of collision prediction of transportation devices.
[0069] Based on this, the embodiments of the present application provide an image processing method and device, a computer device, and a storage medium, aiming to obtain the frame difference between two adjacent video frame images, construct an image mask for each video frame image according to the frame difference, analyze the movement amplitude of the transportation device through the image mask. If the image masks of multiple video frame images are greater than a preset pixel threshold, indicating that the movement amplitude of the transportation device is too large, then fuse the multiple video frame images into a fused image, crop out the highlighted area in the fused image as a cropped image, and predict the collision of the transportation device according to the cropped image. Therefore, cropping out the image of the moving area and then performing collision prediction reduces the quantity and size of image processing, reduces the burden of image processing, and improves the efficiency of collision prediction of transportation devices.
[0070] The image processing method and apparatus, computer device, and storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the image processing method in the embodiments of the present application will be described.
[0071] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0072] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0073] The image processing method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The image processing method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the image processing method, etc., but is not limited to the above forms.
[0074] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0075] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when an embodiment of this application needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiment of this application will be obtained.
[0076] Figure 1 is an optional flowchart of the image processing method provided by the embodiment of this application, Figure 1 The method in may include but is not limited to steps S101 to S107.
[0077] Step S101, obtain at least two consecutive target video frame images; wherein, the target video frame images are obtained by photographing a transportation device;
[0078] Step S102, obtain the frame difference between two adjacent target video frame images to obtain an image frame difference;
[0079] Step S103, construct a mask for each image frame difference to obtain an image mask;
[0080] Step S104, splice the image masks to obtain a target mask;
[0081] Step S105, if the target mask is greater than a preset pixel threshold, fuse at least two frames of target video frame images according to the target mask to obtain a fused image; wherein, the fused image includes a highlighted area;
[0082] Step S106, crop the fused image according to the highlighted area to obtain a cropped image;
[0083] Step S107, perform collision prediction on the transportation device according to the cropped image.
[0084] Steps S101 to S107 shown in the embodiments of the present application, by collecting target video frame images including a transportation device, and collecting at least two consecutive video frame images, and obtaining the frame difference between two adjacent video frame images as the image frame difference. Perform a masking process on the image frame difference to obtain an image mask, so as to judge the movement degree of the transportation device through the image mask. Combine the image masks of multiple video frame images into a target mask, so as to judge the movement amplitude of the transportation device during the process of collecting multiple video frame images according to the target mask. If the target mask is greater than the pixel threshold, it indicates that the movement amplitude of the transportation device is too large during the process of collecting video frame images, then fuse multiple video frame images into a fused image. It should be noted that when fusing multiple video frame images, fuse them according to the target mask, so there is a highlighted area in the fused image, and crop out the highlighted area in the fused image to form a cropped image. By performing collision prediction on the transportation device for the cropped image, the amount of data in the image processing process is reduced, thereby improving the efficiency of image processing, and further improving the efficiency of collision prediction of the transportation device.
[0085] In an application scenario, taking a transportation device as a cage cart as an example, the cage cart travels in the cargo hold, and the collected target video frame images are images of the entire cargo hold. If collision prediction is directly performed on the target video frame images, more computing resources are required. However, if the area including the movement area of the cage cart is cropped out and then collision prediction is performed, the computing resources for collision prediction can be reduced, thereby improving the efficiency of cage cart collision prediction.
[0086] Please refer to Figure 2 In some embodiments, step S101 may include but is not limited to steps S201 to S205:
[0087] Step S201, obtain target video data of the transportation device;
[0088] Step S202, perform motion detection on the target video data to obtain motion detection information;
[0089] Step S203, select video segments from the target video data according to the motion detection information to obtain selected video segment data;
[0090] Step S204: Perform frame division on the selected video segment data to obtain original video frame images;
[0091] Step S205: Obtain at least two consecutive target video frame images from the original video frame images.
[0092] In step S201 of some embodiments, a video recorder is set in the warehouse to capture the process of the transportation equipment transporting in the warehouse through the video recorder. Therefore, the entire warehouse needs to be photographed. Thus, the target video data is the video data that captures the entire panorama of the warehouse. It should be noted that the video screen of the target video data is the entire warehouse. If the collision prediction of the transportation equipment is directly performed on the target video data, it will consume a large amount of computing resources.
[0093] In step S202 of some embodiments, multiple video frame images are extracted from the target video data according to a preset time period. The annotation position information is obtained by annotating the area of the transportation equipment in the target video frame images, and the difference between multiple annotation position information is judged to determine whether the position of the transportation equipment has changed. If it has changed, the motion detection information is determined to indicate that the transportation equipment is in motion. It should be noted that if there is no difference in multiple annotation position information, it is determined that the motion detection information indicates that the transportation equipment in the target video data is not in motion.
[0094] In step S203 of some embodiments, if the running detection information indicates that the transportation equipment is not in motion, the selected video segment data is not extracted from the target video data, and there is no need to perform image processing on the video data, reducing the computing resources consumed by image processing during the collision prediction. It should be noted that if the motion detection information indicates that the transportation equipment is in motion, the moving segment in the target video data is found as the selected video segment data according to the motion detection information. Therefore, only the video segment of the moving process of the transportation equipment is used as the selected video segment data, and there is no need to perform collision prediction on the video data collected throughout the day, reducing the computing resources of image processing during the collision prediction.
[0095] In step S204 of some embodiments, image processing needs to process single-frame video frame images. Therefore, the selected video segment data is divided into frames to obtain original video frame images. It should be noted that frame division means splitting the target video data into frame-by-frame images as the original video frame images.
[0096] In step S205 of some embodiments, the original video frame images are not continuous. Therefore, at least two consecutive original video frame images are extracted from the original video frame images as target video frame images. It should be noted that in this embodiment, 10 consecutive target video frame images will be obtained to more accurately predict the collision situation of the transportation device. In other embodiments, the number of frames of the target video frame images can also be set according to requirements, and the number of frames of the target video frame images is not limited in this embodiment.
[0097] In steps S201 to S205 illustrated in this embodiment, by selecting the video segment data of the movement of the transportation device from the target video data, then splitting the video segment data into multiple original video frame images, and then screening out multiple consecutive target video frame images from the original video frame images. Therefore, there is no need to perform collision prediction on all video data, only the movement segment of the transportation device is obtained, reducing the computing resources consumed by image processing during the collision prediction process, and selecting consecutive video frame images to make the subsequent collision prediction more accurate.
[0098] In step S102 of some embodiments, each target video frame image is equivalent to a matrix recording the pixel values of each pixel point. By subtracting the pixel values between the corresponding pixel points of two adjacent target video frame images to obtain pixel differences, and then constructing multiple pixel differences into an image frame difference. It should be noted that if the target video frame image of the previous frame and the target video frame image of the current frame are the same picture, then it is determined that the image frame difference is all 0. If there are moving objects in the process of the target video frame image of the previous frame and the target video frame image of the current frame, then there are differences in the pixel values of the target video frame image of the previous frame and the target video frame image of the current frame, so the calculated image frame difference is not 0.
[0099] Please refer to Figure 3 , in some embodiments, step S103 may include but is not limited to steps S301 to S303:
[0100] Step S301, perform binarization processing on the image frame difference to obtain an image binarization matrix;
[0101] Step S302, perform erosion processing on the image binarization matrix according to a preset erosion matrix to obtain an image erosion matrix;
[0102] Step S303, perform dilation processing on the image erosion matrix according to a preset number of dilation times and a preset dilation matrix to obtain an image mask.
[0103] In step S301 of some embodiments, binarization is also known as image binarization, and image binarization is to set the grayscale value of the pixel points on the image to 0 or 255, making the image show an obvious black-and-white effect. In this embodiment, the image frame difference of each pixel point is compared with a preset threshold to determine whether the grayscale value of each pixel point is 0 or 255. If the image frame difference is greater than the preset threshold, the grayscale value of the corresponding pixel point is 255; if the image frame difference is less than the preset threshold, the grayscale value of the corresponding pixel point is set to 0. It should be noted that after setting the grayscale values of multiple pixel points, 255 is set to 1, and 0 remains 0 to construct an image binarization matrix composed of 0 and 1.
[0104] In step S302 of some embodiments, the most basic in morphological processing are erosion and dilation processing, and erosion and dilation are two mutually dual operations. The function of erosion is to shrink the image, while dilation is to expand the image. It should be noted that the erosion process needs to use a basic structure element of a certain shape, and the basic structure element is a matrix, so it is defined as an erosion matrix. In this embodiment, the erosion matrix is a 3*3 square structure element, and other embodiments can use basic structure elements of other sizes and shapes. In this embodiment, the image binarization matrix is subjected to an erosion operation once using the erosion matrix to obtain an image erosion matrix.
[0105] In step S303 of some embodiments, the number of dilation times is the number of times of dilation set in advance. Dilation is to make the objects in the image larger and make the objects more connected. And in this embodiment, it is found through verification that when the number of dilation times is 3, the dilation effect is optimal. Therefore, in this embodiment, the image erosion matrix is subjected to 3 dilation processes using a dilation matrix to obtain an image mask, so the image mask can accurately represent the moving area in the target video frame image. Specifically, through multiple verifications, the effect of using a 7*7 square structure element as the dilation matrix is the best. Therefore, in this embodiment, a 7*7 square structure element is used to perform 3 dilation processes on the image erosion matrix to obtain an image mask that can accurately represent the moving area.
[0106] In steps S301 to S303 illustrated in this embodiment, after performing binarization processing on the image frame difference, an erosion operation is performed on the binarized matrix using an erosion matrix, and then the eroded matrix is subjected to 3 dilations using a dilation matrix to obtain an image mask, so as to construct an image mask that can accurately represent the moving area.
[0107] In step S104 of some embodiments, after calculating the image mask of an image frame difference, if there are multiple image frame differences, multiple image masks need to be combined into a target mask. It should be noted that in this embodiment, 10 consecutive target video frame images are collected, so there will be 9 image frame differences, so 9 image masks are combined into a target mask.
[0108] Please refer toFigure 4 , in some embodiments, step S105 may include but is not limited to steps S401 to S403:
[0109] Step S401, if the target mask is greater than the pixel point threshold, obtain the pixel points of the target video frame image to obtain image pixel points;
[0110] Step S402, perform screening processing on the preset fusion weights according to the target mask to obtain the target weight of each image pixel point;
[0111] Step S403, construct a fused image according to the target weight and the image pixel points.
[0112] It should be noted that when comparing the target mask with the pixel point threshold, the pixel points with a value of 1 in the target mask are summarized as the current number of pixel points, and the current number of pixel points is compared with the pixel point threshold to determine whether the motion area of the target video frame image is large enough. If the current number of pixel points is less than the pixel point threshold, it is considered that there is no motion area in the multi-frame target video frame image, and there is no need to fuse the multi-frame target video frame images.
[0113] In step S401 of some embodiments, if the target mask is greater than the pixel point threshold, it indicates that the current number of pixel points is greater than the pixel point threshold, that is, there is a large motion area in the multi-frame target video frame image. Therefore, by obtaining the pixel points of each target video frame image as image pixel points, fusion processing is performed on the image pixel points.
[0114] In step S402 of some embodiments, the target mask is represented as a matrix, and each value of the image pixel point is recorded in the matrix. The pixel points with a value of 1 indicate that they are located in the motion area, and the pixel points with a value of 0 indicate that they are located in the non-motion area. Therefore, the fusion weights are screened according to the target mask, and different fusion weights are selected as the target weights of each image pixel point according to the value of 1 or 0, so as to distinguish the pixel points in the motion area and the non-motion area through the target weights. It should be noted that in this embodiment, the motion area needs to be brightened and the non-motion area needs to be darkened, so the target weight of the motion area is higher than that of the non-motion area.
[0115] In step S403 of some embodiments, after determining the target weight of each image pixel point, the fused image is obtained by multiplying the target weight by each image pixel point. In this embodiment, the target weight of the motion area is 1, and the target weight of the non-motion area is 0.4, so as to darken the non-motion area, so the highlighted area is the motion area.
[0116] For example, please refer to Figure 5 as shown, Figure 5 represents the fused image after fusion. By Figure 5It can be seen that the highlighted area represents the moving area, and the dark area represents the non-moving area. The moving area and the non-moving area are distinguished by brightness, so as to extract the moving area. The collision prediction of the transportation equipment through the image of the moving area is more accurate and reduces the amount of calculation.
[0117] In steps S401 to S403 shown in the present embodiment, target weights of pixels in different areas are set so that image pixels are fused into a fused image according to the target weights, and moving areas and non-moving areas are divided by different brightness levels, so that the image of the moving area can be cropped out from the fused image. There is no need to process the image of the entire warehouse, thereby reducing the computing resources for image processing.
[0118] See also Figure 6 In some embodiments, step S106 may include but is not limited to steps S601 to S603:
[0119] Step S601, marking the highlighted area to obtain a marked area;
[0120] Step S602, performing expansion processing on the marked area according to a preset expansion index to obtain a selected area;
[0121] Step S603, cropping the fused image according to the selected area to obtain a cropped image.
[0122] In step S601 of some embodiments, the highlighted area in the fused image is marked as the marked area. It should be noted that in this embodiment, the highlighted area is marked with a red border to highlight the motion area. If there are multiple highlighted areas, each highlighted area is marked with a red border.
[0123] In step S602 of some embodiments, directly cropping the fused image according to the marked area can only capture the moving transport equipment, but cannot see whether the transport equipment is about to collide with other equipment. Therefore, the marked area is expanded by an expansion index to expand the marked area to obtain a selected area, and the selected area includes the transport equipment and other nearby equipment, so cropping the cropped image from the fused image according to the selected area is beneficial for collision prediction.
[0124] In step S603 of some embodiments, the fused image is cropped according to the selected area after expansion to obtain a cropped image. It should be noted that if there are multiple selected areas, multiple cropped images will be obtained. For example, if the transport equipment is a cage truck, the cropped image is as follows: Figure 7 As shown, through Figure 7The distance between the moving cage car and other cage cars can be clearly seen, so as to make the collision prediction more accurate, and a prompt message can be output when the cage cars are about to collide, facilitating timely adjustment of the moving direction of the cage cars and reducing the potential safety hazards caused by cage car collisions.
[0125] In steps S601 to S603 shown in this embodiment, the selected area is obtained by first marking the highlighted area and then expanding it by one circle, so as to crop the cropped image from the fused image. Therefore, the cropped image not only includes the transportation equipment during the movement process, but also other equipment close to the transportation equipment, facilitating more accurate collision prediction of the transportation equipment based on the cropped image.
[0126] Please refer to Figure 8 , in some embodiments, step S107 includes but is not limited to steps S801 to S803:
[0127] Step S801, perform object recognition on the cropped image to obtain the image object category;
[0128] Step S802, use the cropped image with the image object category of the transportation equipment as the selected image;
[0129] Step S803, perform collision prediction on the transportation equipment according to the selected image.
[0130] In step S801 of some embodiments, there may be multiple cropped images, so object recognition is performed on the cropped images, that is, the object categories in the cropped images are detected to obtain the image object category. It should be noted that the image object category represents the category of the moving objects in the cropped image.
[0131] In step S802 of some embodiments, if the image object category is the transportation equipment, then the cropped image with the image object category of the transportation equipment is used as the selected image. It should be noted that if the image object category of each cropped image is not the transportation equipment, it means that the transportation equipment is not moving, but other moving objects interfere with the collision prediction of the transportation equipment, and new target video frame images need to be collected again for image cropping. For example, if the transportation equipment is a cage car and the recognized image category of the cropped image is a cage car, then the cropped image corresponding to the cage car is used as the selected image, which is equivalent to Figure 7 being used as the selected image.
[0132] Specifically, in this embodiment, an object detection network is used to perform object recognition on each cropped image, making the recognition of the image object category intelligent and saving the manpower of manual recognition.
[0133] In step S803 of some embodiments, by performing collision prediction on the selected image for the transportation device, accurate collision prediction for the transportation device can be made, and it is not necessary to perform collision prediction on all the cropped images, reducing the image processing efficiency, thereby reducing the computational amount of collision prediction and improving the efficiency of collision prediction.
[0134] In steps S801 to S803 illustrated in this embodiment, first, the object category inside each cropped image is identified, and then the cropped image with the object category of the transportation device is used as the image for collision prediction, which can not only ensure the accuracy of collision prediction for the transportation device but also reduce the computational amount of image processing during collision prediction and improve the efficiency of collision prediction.
[0135] Please refer to Figure 9 , in some embodiments, step S803 may include but is not limited to steps S901 to S902:
[0136] Step S901, obtaining the distance between two transportation devices based on the selected image to obtain the device distance;
[0137] Step S902, performing collision prediction on the transportation device according to the device distance and a preset distance range.
[0138] In step S901 of some embodiments, if multiple transportation devices are stored in the selected image, the distance between any two transportation devices is obtained as the device distance to determine whether the transportation devices in the moving process will collide through the device distance. It should be noted that if there is one transportation device in the selected image, the distance between the transportation device and other objects is judged as the device distance for judging whether the transportation device collides with other objects.
[0139] In step S902 of some embodiments, the preset distance range is the distance range that measures the likelihood of collision of the transportation device and is set after multiple verifications. If the device distance falls within the preset distance range, a collision is likely to occur. Conversely, if the device distance does not fall within the preset distance range, it is considered that the transportation device will not collide. Therefore, by comparing the device distance with the preset distance range, the collision prediction of the transportation device in the selected image is realized, making the collision prediction of the transportation device simpler.
[0140] Specifically, in this embodiment, if the equipment spacing falls within a preset spacing range, the collision prediction information is output indicating that a collision is about to occur, and an alarm message is output. The alarm message is provided to the internal management personnel of the warehouse so that the management personnel can stop the running transportation equipment that is about to collide in time to reduce the collision of the transportation equipment. If the equipment spacing does not fall within the preset spacing range, the collision prediction information is output indicating that no collision will occur, and then the video frame image at the current time is continuously collected, and the collision prediction is continued to prevent collisions of the transportation equipment in the warehouse and reduce the potential safety hazards caused by the collision of the transportation equipment, thereby improving the safety inside the warehouse.
[0141] In steps S901 to S902 shown in this embodiment, by obtaining the spacing between two transportation equipment in the selected image as the equipment spacing, it is determined whether the transportation equipment is likely to collide based on the equipment spacing and the preset spacing range, making the collision prediction of the transportation equipment easier.
[0142] As Figure 5 and Figure 7 shown, if the application scenario is a warehouse with cage car transportation, the present application embodiment continuously collects 10 target video frame images, calculates the frame difference between two adjacent target video frame images to obtain the image frame difference, then performs binary processing on each image frame difference to obtain an image binary matrix, then performs erosion processing on the image binary matrix according to a preset erosion matrix to obtain an image erosion matrix, and performs 3 times of dilation on the image erosion matrix according to a dilation matrix to obtain an image mask. Then, the image masks of multiple image frame differences are combined into a target mask, and the number of pixel points with a value of 1 in the target mask is obtained as the current number of pixel points. If the current number of pixel points is greater than 10,000 pixel points, it is proved that the range of the moving area in the 10 target video frame images is large. By obtaining the pixel points of each target video frame image as the image pixel points, the target weight value of the moving area is determined to be 1 according to the target mask, and the target weight value of the non-moving area is 0.4. The image pixel points of the non-moving area are multiplied by 0.4 to obtain a fused image. It should be noted that the non-moving area is darkened to obtain the fused image, and the fused fused image is as Figure 5 shown. Then, the highlighted area in the fused image is expanded by one circle and then cropped to obtain a cropped image including the transportation equipment and other equipment around the transportation equipment. Object recognition is performed on each cropped image to obtain the image object category, and collision prediction is performed on the cropped image with the image object category of the cage car. Therefore, after determining that there is a large range of moving areas in multiple image frames, the multiple video frame images are fused into a fused image, and then the moving area of the fused image is cropped out as a cropped image. By performing collision prediction on the cropped image, the amount of computation for image processing in the collision prediction process is reduced, thereby improving the collision prediction efficiency.
[0143] Please refer to Figure 10, embodiments of the present application further provide an image processing apparatus that can implement the above image processing method. The apparatus includes:
[0144] An image acquisition module 1001, configured to acquire at least two consecutive target video frame images; wherein, the target video frame images are obtained by photographing a transportation device;
[0145] A frame difference acquisition module 1002, configured to acquire the frame difference between two adjacent target video frame images to obtain an image frame difference;
[0146] A mask construction module 1003, configured to perform mask construction on each image frame difference to obtain an image mask;
[0147] A mask splicing module 1004, configured to splice the image masks to obtain a target mask;
[0148] An image fusion module 1005, configured to, if the target mask is greater than a preset pixel point threshold, fuse at least two target video frame images according to the target mask to obtain a fused image; wherein, the fused image includes a highlighted area;
[0149] An image cropping module 1006, configured to crop the fused image according to the highlighted area to obtain a cropped image;
[0150] A collision prediction module 1007, configured to perform collision prediction on the transportation device according to the cropped image.
[0151] The specific implementation manner of this image processing apparatus is basically the same as the specific embodiment of the above image processing method, and will not be elaborated here.
[0152] Embodiments of the present application further provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above image processing method. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0153] Please refer to Figure 11 , Figure 11 illustrates the hardware structure of a computer device in another embodiment. The computer device includes:
[0154] A processor 1101, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0155] The memory 1102 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102 and are called by the processor 1101 to execute the image processing method of the embodiments of this application;
[0156] The input / output interface 1103 is used to implement information input and output;
[0157] The communication interface 1104 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WI FI, Bluetooth, etc.);
[0158] The bus 1105 transmits information between the various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);
[0159] Among them, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 are communicatively connected to each other inside the device through the bus 1105.
[0160] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image processing method is implemented.
[0161] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0162] The image processing method, device, computer device, and storage medium provided by the embodiments of the present application analyze the movement amplitude of the transportation device through an image mask. If the image masks of multiple video frame images are greater than a preset pixel threshold, indicating that the movement amplitude of the transportation device is too large, then multiple video frame images are fused into a fused image, and the highlighted area in the fused image is cropped out as a cropped image, and a collision prediction for the transportation device is made based on the cropped image. By determining that there is a large range of movement areas in multiple image frames, multiple video frame images are fused into a fused image, and then the movement area of the fused image is cropped out as a cropped image. By making a collision prediction on the cropped image, the computational amount of image processing in the collision prediction process is reduced, thereby improving the collision prediction efficiency.
[0163] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0164] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0166] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0167] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0168] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0169] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0170] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0172] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0173] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining at least two consecutive target video frame images; wherein, the target video frame images are captured from a transportation device; Obtaining the frame difference between two adjacent target video frame images to obtain an image frame difference; Performing mask construction on each image frame difference to obtain an image mask; Performing splicing processing on the image masks to obtain a target mask; If the target mask is greater than a preset pixel point threshold, performing fusion processing on at least two target video frame images according to the target mask to obtain a fusion image; wherein, the fusion image includes a highlighted area; Performing cropping processing on the fusion image according to the highlighted area to obtain a cropped image; Performing collision prediction on the transportation device according to the cropped image.
2. The method according to claim 1, wherein The performing mask construction on each image frame difference to obtain an image mask includes: Performing binarization processing on the image frame difference to obtain an image binarization matrix; Performing erosion processing on the image binarization matrix according to a preset erosion matrix to obtain an image erosion matrix; Performing dilation processing on the image erosion matrix according to a preset number of dilation times and a preset dilation matrix to obtain the image mask.
3. The method according to claim 1, wherein The if the target mask is greater than a preset pixel point threshold, performing fusion processing on at least two target video frame images according to the target mask to obtain a fusion image includes: If the target mask is greater than the pixel point threshold, obtaining the pixel points of the target video frame images to obtain image pixel points; Performing screening processing on a preset fusion weight according to the target mask to obtain a target weight for each image pixel point; Constructing the fusion image according to the target weight and the image pixel points.
4. The method according to any one of claims 1 to 3, characterized in that, The performing cropping processing on the fusion image according to the highlighted area to obtain a cropped image includes: Performing annotation processing on the highlighted area to obtain an annotated area; Performing outward expansion processing on the annotated area according to a preset outward expansion index to obtain a selected area; Performing cropping processing on the fusion image according to the selected area to obtain the cropped image.
5. The method according to any one of claims 1 to 3, characterized in that, The performing collision prediction on the transportation device according to the cropped image includes: Performing object recognition on the cropped image to obtain an image object category; Taking the cropped image whose image object category is the transportation device as a selected image; Performing collision prediction on the transportation device according to the selected image.
6. The method according to claim 5, wherein The performing collision prediction on the transportation device according to the selected image includes: Obtaining the distance between two transportation devices according to the selected image to obtain a device distance; Performing collision prediction on the transportation device according to the device distance and a preset distance range.
7. The method according to any one of claims 1 to 3, characterized in that, The obtaining at least two consecutive target video frame images includes: Obtaining target video data of the transportation device; Performing motion detection on the target video data to obtain motion detection information; Performing video segment selection on the target video data according to the motion detection information to obtain selected video segment data; Performing frame division on the selected video segment data to obtain original video frame images; Obtain at least two consecutive target video frame images from the original video frame images.
8. An image processing apparatus, characterized in that, The device includes: An image acquisition module, configured to obtain at least two consecutive target video frame images; wherein, the target video frame images are obtained by photographing a transportation device; A frame difference acquisition module, configured to obtain the frame difference between two adjacent target video frame images to obtain an image frame difference; A mask construction module, configured to perform mask construction on each of the image frame differences to obtain an image mask; A mask splicing module, configured to splice the image masks to obtain a target mask; An image fusion module, configured to, if the target mask is greater than a preset pixel point threshold, fuse at least two target video frame images according to the target mask to obtain a fused image; wherein, the fused image includes a highlight area; An image cropping module, configured to crop the fused image according to the highlight area to obtain a cropped image; A collision prediction module, configured to perform collision prediction on the transportation device according to the cropped image.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the image processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1 to 7.