Crane anti-collision warning method, system, electronic equipment and storage medium
By preprocessing the historical video surveillance data of Kelingjiu and training a lightweight network structure model, combined with the binocular visual ranging algorithm, the problem of manual observation error operation in Kelingjiu's anti-collision is solved, and a more accurate anti-collision warning is achieved.
Patent Information
- Application Number
- CN202411309567.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The existing Keling Chime anti-collision method relies on manual observation to easily lead to misoperation and delayed reactions, and there is a great risk of collision.
By obtaining the historical video surveillance data of Kelingdan's surrounding environment for preprocessing, and importing a lightweight network structure model to train the object detection model, combining the binocular visual distance measurement algorithm to identify the target and judge the distance in real time, and perform early warning processing.
It improves the accuracy and efficiency of target recognition around Kelinghang, and achieves a more accurate anti-collision safety warning.
Smart Images

Figure CN119339315B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and in particular to an anti-collision warning method, system, electronic device, and storage medium for a crane. Background Art
[0002] Cranes are common marine lifting equipment, widely used for cargo loading and unloading operations at ports, docks, and ships. They primarily consist of a boom, slewing mechanism, balancer, crane frame, and electrical control system, driven by a hydraulic system. Crane operation involves starting the crane, loading, moving, unloading, and shutting down the crane. Their stability and safety are crucial to ensuring efficient loading and unloading operations. Currently, with the rapid development of my country's economy, cranes are increasingly used for lifting operations and are therefore increasingly important for their collision avoidance.
[0003] Currently, the existing crane collision prevention method uses video surveillance technology to record on-site footage. The operator then analyzes the footage to determine if a collision is possible. Sensors are used to detect the distance between the crane and obstacles, controlling the crane's operation and achieving collision prevention. However, manual observation can easily lead to fatigue, misoperation, or misjudgment, posing a significant risk. Furthermore, manually braking the crane when a risk is identified may fail to prevent a collision in time due to the crane's braking time, further increasing the risk of collision. Summary of the Invention
[0004] The purpose of this application is to provide a crane anti-collision warning method, system, electronic equipment and storage medium to solve the problem in related technologies that crane collision warning cannot be accurately and timely provided.
[0005] In order to achieve the above objectives, this application provides the following technical solutions:
[0006] In a first aspect, the present application provides a crane anti-collision warning method, comprising:
[0007] Historical video surveillance data of the environment around the crane is obtained, and the historical video surveillance data is preprocessed to obtain preprocessed historical video surveillance data; the preprocessed historical video surveillance data is imported into a preset lightweight network structure model for training to obtain a target detection model in the environment around the crane; real-time video surveillance data of the environment around the crane is obtained, and the real-time surveillance data is imported into the target detection model for identification and detection to obtain target information in the environment around the crane, wherein the target information includes: category information and position information of the target; the target information is imported into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; it is determined whether the real-time distance is less than a preset threshold, and if so, corresponding early warning processing is performed on the crane.
[0008] Preferably, the acquiring of historical video surveillance data of the environment around the crane and preprocessing the historical video surveillance data to obtain preprocessed historical video surveillance data includes: acquiring historical video surveillance data of the environment around the crane and performing frame extraction processing on the historical video surveillance data to obtain a historical surveillance image dataset of the environment around the crane; and performing image scale normalization and image standardization processing on each frame of the historical surveillance image dataset to obtain a processed historical surveillance image dataset of the environment around the crane.
[0009] Preferably, the preset lightweight network structure model is specifically constructed in the following steps: designing a loss function, the loss function including a positioning loss function and a similarity measurement loss function; and constructing a lightweight network structure model based on the loss function and the YOLOv8 algorithm.
[0010] Preferably, the loss function is calculated as follows:
[0011]
[0012] Among them, F loss is the loss function, L loc is the positioning loss function, L cls is the similarity metric loss function, i is the index number of the anchor point of the i-th candidate region in the image, and p i Represents the target prediction probability of the anchor point; represents the candidate region label, N loc With N cls Respectively represent the classification loss function in the positioning loss L loc With similarity metric loss L cls Normalization coefficient after functionalization, λ represents N loc With N cls The balance weight between .
[0013] As an example, the positioning loss function L loc , the calculation formula is as follows:
[0014] L loc =αL reg +βL CIoU
[0015] Among them L reg Represents the position regression loss function, L CIoU is the generalized intersection-over-union loss function, and α and β are hyperparameters used to adjust the functions of each part;
[0016] The position regression loss function L reg The calculation formula is as follows:
[0017]
[0018] Among them, B is the predicted box, B′ is the target box, IoU represents the ratio of the intersection and union of the target box and the real box, and B C Represents the smallest enclosed box in the positional relationship between the prediction box and the target box, and U represents the union of the prediction box and the target box;
[0019] The generalized intersection-over-union loss function L CIoU The calculation formula is as follows:
[0020]
[0021] Among them, a is the weight parameter; v is used to measure the consistency of the aspect ratio, w′ and h′ are the width and height of the target box, respectively, and w and h are the width and height of the prediction box, respectively;
[0022] L cls Represents the similarity measurement loss function, which is calculated as follows:
[0023]
[0024] Among them, x represents the true sample label, and x′ represents the detection output after the activation function.
[0025] Preferably, the method of importing the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment includes:
[0026] The historical video surveillance data is divided into a training set, a validation set and a test set; the training set is subjected to data enhancement processing to obtain an enhanced training set; the enhanced training set and the validation set are imported into a preset lightweight network structure model for training to obtain a target detection model in the environment around the crane.
[0027] Preferably, an anti-collision warning is provided to the crane based on the real-time distance, including: issuing a voice warning when the real-time distance is less than or equal to a first set distance; and controlling the crane to stop running when the real-time distance is less than or equal to a second set distance.
[0028] In a second aspect, the present application further provides a crane anti-collision warning system, comprising:
[0029] An acquisition module is used to acquire historical video surveillance data of the environment around the crane and preprocess the historical video surveillance data to obtain preprocessed historical video surveillance data; a training module is used to import the preprocessed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the environment around the crane; an identification and detection module is used to acquire real-time video surveillance data of the environment around the crane and import the real-time surveillance data into the target detection model for identification and detection to obtain target information in the environment around the crane, wherein the target information includes: target category information and location information; a ranging module is used to import the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; and an early warning module is used to determine whether the real-time distance is less than a preset threshold value. If so, corresponding early warning processing is performed on the crane.
[0030] In a third aspect, the present application further provides a computer electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the anti-collision warning method for a crane as described in any one of the above are implemented.
[0031] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the anti-collision warning method for a crane as described in any one of the above.
[0032] The present application provides a crane anti-collision warning method, system, electronic device and storage medium. The method comprises the following steps: obtaining historical video surveillance data of the crane's surrounding environment and preprocessing the historical video surveillance data to obtain preprocessed historical video surveillance data; importing the preprocessed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment; obtaining real-time video surveillance data of the crane's surrounding environment and importing the real-time surveillance data into the target detection model for identification and detection to obtain target information in the crane's surrounding environment, wherein the target information includes: target category information and location information; importing the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; and determining whether the real-time distance is less than a preset threshold. If so, performing corresponding warning processing on the crane. Compared with the existing technology, the present invention has the following effects: by designing a lightweight target detection model and a binocular vision ranging algorithm, the present application can improve the accuracy and efficiency of identifying targets around the crane, thereby more accurately obtaining target information and real-time distance information, thereby improving the performance of the crane's anti-collision safety warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a flow chart of a method for preventing collisions and providing early warning for a crane according to an embodiment of the present application;
[0034] Figure 2 1 is a flow chart of a binocular vision ranging algorithm in an embodiment of the present application;
[0035] Figure 3 This is a schematic diagram of the structure of the lightweight network structure model in the embodiment of the present application;
[0036] Figure 4 This is a schematic diagram of the structure of a multi-scale decoupled head in a lightweight network structure model according to an embodiment of the present application;
[0037] Figure 5 This is a structural diagram of an anti-collision warning system for a crane according to an embodiment of the present application;
[0038] Figure 6 It is a structural diagram of a computer electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0040] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. Conversely, when an element is referred to as being "directly on" another element, there is no intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.
[0041] In this application, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components or interactions between two components. Those skilled in the art will understand the specific meanings of these terms in this application based on specific circumstances.
[0042] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0043] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "the" used in one or more embodiments of the present application are also intended to include plural forms unless the context clearly indicates otherwise.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the template description herein are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0045] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when...".
[0046] Please refer to Figure 1 The embodiment of the present application provides a crane anti-collision warning method, which is applied to a shipbuilding crane and includes the following steps:
[0047] S10: Acquire historical video surveillance data of the surrounding environment of the crane, and preprocess the historical video surveillance data to obtain preprocessed historical video surveillance data.
[0048] It is understandable that due to the large working range of the crane, the crane's working scene is generally filled with multiple video surveillance devices. The historical data of these video surveillance devices is used as raw data for training and testing.
[0049] It should be noted that in order to improve the accuracy and effectiveness of raw data, historical data must first be preprocessed. Data preprocessing can include: data cleaning: including removing duplicate values, handling missing values, and handling outliers to ensure data accuracy and consistency; data conversion: converting data to adapt to specific analysis methods. This includes normalization / standardization, discretization, and data transformation; data transformation: performing mathematical transformations on data to change its distribution or reduce noise; outlier processing: detecting and handling outliers to prevent them from adversely affecting data analysis results; data smoothing: smoothing data to eliminate noise and irregular changes.
[0050] S20: Import the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the environment surrounding the crane.
[0051] Specifically, the pre-processed data monitoring data in step S10 is imported into a pre-built lightweight model for training, so as to obtain a trained target detection model in the environment surrounding the crane.
[0052] It should be noted that during training, different training strategies can be selected, such as setting a training data budget and data augmentation operations. Setting a training data budget means that when starting a new machine learning project, you must first define the goal, which will determine what type of data is needed and how many "training items" (classified data points) are needed. For example, for computer vision or image recognition projects, you may need to use manually annotated image data for training. Depending on the needs of the project, it may be necessary to continuously retrain or update the model, so it is necessary to evaluate the options for procuring data and calculate the budget. Data augmentation refers to increasing the diversity of data through technical means to improve the generalization ability of the model. This includes operations such as flipping, cropping, rotating the image, or transforming the pixel color, such as adding noise.
[0053] S30. Acquire real-time video surveillance data of the crane's surrounding environment, and import the real-time surveillance data into the target detection model for recognition and detection to obtain target information in the crane's surrounding environment, wherein the target information includes target category information and location information.
[0054] Specifically, after the target detection model is trained in step S20, real-time monitoring data from the video surveillance equipment around the crane can be collected, and the real-time monitoring data can be processed and imported into the target detection model for identification and detection, thereby obtaining target information in the environment around the crane.
[0055] It should be noted that the preprocessing process of real-time monitoring data can refer to the description of preprocessing historical monitoring data.
[0056] S40: Importing the target information into a preset binocular vision ranging algorithm model to obtain a real-time distance between the target object and the crane.
[0057] Specifically, the target information around the crane is imported into a preset binocular vision algorithm model, so that the real-time distance between the target object and the crane can be obtained.
[0058] It should be noted that the binocular visual ranging algorithm is primarily based on the similar triangle principle and the pixel scale principle. These two principles explain, from different perspectives, how to measure the distance of objects in a scene using two cameras (i.e., binocular cameras). The similar triangle principle uses two parallel cameras (Or and Ot) with a distance B between them and a focal length f for each camera lens. A point in the scene, P, is located at a distance Z from the camera. Using the similar triangle principle, an expression for Z can be derived, thereby calculating the distance between the object and the camera. The pixel scale principle uses the camera to capture the same scene at different positions, calculating the distance using the camera's movement distance and the pixel change of the object in the captured image. Specifically, the change in the pixel coordinates of an object in the images captured by the camera at positions 1 and 2, combined with the camera's field of view and lens parameters, allows the actual distance represented by each pixel to be calculated, thereby determining the distance between the object and the camera.
[0059] For example, please refer to Figure 2 , the workflow of the binocular vision ranging algorithm is as follows:
[0060] First, after building the binocular vision ranging model, complete the calibration and verification and other preparatory work;
[0061] Secondly, the left and right images acquired by the binocular camera are corrected using the binocular camera parameters, and the disparity maps of the corrected left and right images are obtained through the BM stereo matching algorithm;
[0062] Then, the target location information is input, and the distance of each point in the detection target location area is calculated using the disparity information and the average value is taken to calculate the distance D of the target;
[0063] Finally, the real-time distance between the target object and the crane in the detected image is output.
[0064] S50: Determine whether the real-time distance is less than a preset threshold; if so, perform corresponding warning processing on the crane.
[0065] Specifically, when the real-time distance between the crane and the target objects around it is less than a preset threshold, the crane can be given corresponding warning processing, such as slowing down the crane, sending a warning message, or stopping the crane operation.
[0066] In some optional embodiments, when the real-time distance is less than or equal to a first set distance, a voice warning is issued; when the real-time distance is less than or equal to a second set distance, the crane is controlled to stop operating.
[0067] For example, when the distance between the target object and the crane is less than or equal to 10 meters, a voice warning message can be sent to remind relevant personnel to take action. When the distance between the target object and the crane is less than or equal to 3 meters, the crane is directly controlled to stop running to avoid collision.
[0068] The present application provides a crane anti-collision warning method, which is applied to crane usage scenarios. The warning method obtains historical video surveillance data of the crane's surrounding environment and pre-processes the historical video surveillance data to obtain pre-processed historical video surveillance data; imports the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model for the crane's surrounding environment; obtains real-time video surveillance data of the crane's surrounding environment and imports the real-time surveillance data into the target detection model for identification and detection to obtain target information in the crane's surrounding environment, wherein the target information includes: target category information and location information; imports the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; and issues a crane anti-collision warning based on the real-time distance. Compared with the prior art, the present application has the following effects: by designing a lightweight target detection model and a binocular vision ranging algorithm, the present application can improve the accuracy and efficiency of target recognition around the crane, thereby more accurately obtaining target information and real-time distance information, thereby improving the performance of the crane anti-collision safety warning.
[0069] In some optional embodiments, step S10, obtaining historical video surveillance data of the crane's surrounding environment and preprocessing the historical video surveillance data to obtain preprocessed historical video surveillance data, includes the following steps:
[0070] S11. Acquire historical video surveillance data of the environment around the crane, and perform frame extraction processing on the historical video surveillance data to obtain a historical surveillance image dataset of the environment around the crane.
[0071] Specifically, after obtaining historical video surveillance data, the video surveillance data first needs to be frame-extracted. The purpose of this is to extract key frames from the video data. These key frames usually contain the main information of the video, which can significantly reduce the amount of data that needs to be processed, thereby improving the efficiency of image processing.
[0072] It should be noted that the frame extraction method includes uniform frame extraction and specified time frame extraction. The former extracts the picture by setting a fixed time interval, and the latter extracts the picture according to a specific time range and number of frames. In this application, there is no specific restriction on the frame extraction method.
[0073] S12. Perform image scale normalization and image standardization processing on each frame of the historical monitoring image dataset to obtain a processed historical monitoring image dataset of the crane's surrounding environment.
[0074] Specifically, after obtaining a dataset of historical surveillance images of the crane's surroundings, we need to perform image scale normalization and image standardization on each frame in the dataset. This process can achieve the following effects:
[0075] Image resizing resizes images to a uniform standard. By resizing images of varying sizes to the same size, we ensure consistency and comparability across algorithms. Furthermore, resizing helps reduce computational resource consumption, as processing images of varying sizes may require varying computational resources. Normalization ensures that algorithms have the same computational complexity when processing images of varying sizes, thereby improving processing efficiency.
[0076] Image normalization adjusts the pixel values of an image to a uniform range, such as [0, 1] or [-1, 1]. This helps improve the numerical stability and convergence speed of the algorithm. Normalization also makes the pixel values of different images comparable, allowing the algorithm to more accurately extract image features. Furthermore, normalization improves the generalization ability of the model, making it better adaptable to different data distributions and reducing the risk of overfitting.
[0077] In this embodiment, by performing frame extraction, scale normalization, and image standardization on the video image, the image processing efficiency can be improved, the consumption of computing resources can be reduced, and the model performance and stability can be enhanced.
[0078] In some optional embodiments, the preset lightweight network structure model is specifically constructed in the following steps: designing a loss function, the loss function including a positioning loss function and a similarity measurement loss function; and constructing a lightweight network structure model based on the loss function and the YOLOv8 algorithm.
[0079] Specifically, the specific structure of the lightweight network structure model is as follows Figure 3 As shown, the working principle of the lightweight network structure model is as follows:
[0080] In the first step, the monitoring stream is processed by the data enhancement module and then downsampled through a Focus layer.
[0081] In the second step, the image is passed through the data enhancement module and the Focus layer and then input into the backbone network for feature extraction to obtain three feature maps of different sizes.
[0082] It should be noted that the data enhancement module uses the mean shift filter algorithm (Mean Shift) and the contrast-constrained adaptive histogram equalization algorithm (CLAHE) to preprocess the image. These two methods can effectively reduce background noise and enhance the information of crack edges.
[0083] The backbone network consists of a reparameterization module, SC2f lightweight convolution, and SPPF module. The reparameterization module and SC2f lightweight convolution are used to perform operations such as convolution, batch normalization, and pooling. Mish function activation is applied after batch normalization to reduce the number of parameters and optimize calculations. The feature pyramid module (SPPF) performs multiple pooling on feature maps and then fuses and splices feature maps of multiple dimensions, enabling the network to extract higher-level semantic features.
[0084] In the third step, different feature maps are input into the attention mechanism module respectively, attention weights are assigned, and new feature maps are obtained.
[0085] Specifically, the attention mechanism module consists of the Efficient Channel Attention (ECA) mechanism, a cross-channel attention mechanism that does not reduce dimensionality and is beneficial for improving the learning ability of the network model. To further enhance the model's performance, the ECA input layer performs adaptive average pooling and adaptive max pooling operations, followed by an addition operation, to enhance the module's global visual information and make the model adaptable to inputs of varying sizes.
[0086] The fourth step is to input the new feature map into the multi-scale feature fusion module for feature fusion.
[0087] Specifically, the multi-scale feature fusion module first records the 20×20 feature layer output by the SPPF module as the X1 layer. The X1 feature layer is then upsampled by a factor of 2 and concatenated with the 40×40 feature layer to form a concatenated 40×40 feature layer, recorded as the X2 feature layer. Next, the X2 and X1 feature layers are upsampled by a factor of 2 and 4, respectively, and concatenated with the 80×80 feature layer to form the X3 feature layer. Finally, to speed up the network's computation, an Hi module is added to each output layer to reduce the number of feature maps: Hi(i=1,2,3…)=Conv+BN+Mish. Ultimately, the X1, X2, and X3 feature layers output head classifiers for detecting small, medium, and large objects, respectively.
[0088] In the fifth step, the information generated by the multi-scale decoupling head is processed by the small target decoupling head, the medium target decoupling head and the large target decoupling head, and then the final detection result is output after the information is generated by the multi-scale decoupling head through the corresponding post-processing steps.
[0089] It should be noted that the multi-scale decoupling head structure is as follows Figure 4 As shown, in Figure 4 In
[15] , Reg, Obj, and Cls are prediction results. Reg is used to represent the regression parameters of each feature point, and Obj and Cls are used to determine whether an object is contained and the type of the object.
[0090] In some optional embodiments, the loss function is calculated as follows:
[0091]
[0092] Among them, F loss is the loss function, L loc is the positioning loss function, L cls is the similarity metric loss function, i is the index number of the anchor point of the i-th candidate region in the image, and p i Represents the target prediction probability of the anchor point; represents the candidate region label, N loc With N cls Respectively represent the classification loss function in the positioning loss L loc With similarity metric loss L cls Normalization coefficient after functionalization, λ represents N loc With N cls The balance weight between .
[0093] In some optional embodiments, the positioning loss function L loc , the calculation formula is as follows:
[0094] L loc =αL reg +βL CIoU
[0095] Among them L reg Represents the position regression loss function, L CIoU is the generalized intersection-over-union loss function, and α and β are hyperparameters used to adjust the functions of each part;
[0096] The position regression loss function L reg The calculation formula is as follows:
[0097]
[0098] Among them, B is the predicted box, B′ is the target box, IoU represents the ratio of the intersection and union of the target box and the real box, and B C Represents the smallest enclosed box in the positional relationship between the prediction box and the target box, and U represents the union of the prediction box and the target box;
[0099] The generalized intersection-over-union loss function L CIoU The calculation formula is as follows:
[0100]
[0101] Among them, a is the weight parameter; v is used to measure the consistency of the aspect ratio, w′ and h′ are the width and height of the target box, respectively, and w and h are the width and height of the prediction box, respectively;
[0102] L cls Represents the similarity measurement loss function, which is calculated as follows:
[0103]
[0104] Among them, x represents the true sample label, and x′ represents the detection output after the activation function.
[0105] In some optional embodiments, step S20, importing the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment, includes:
[0106] S21. Divide the historical video surveillance data into a training set, a validation set, and a test set.
[0107] Exemplarily, the data in the historical video surveillance data is divided into a training set, a validation set, and a test set in a ratio of 7:2:1.
[0108] S22. Perform data enhancement processing on the training set to obtain an enhanced training set.
[0109] For example, the Mosaic and MixUp data augmentation methods can be used to enrich the training set;
[0110] It should be noted that the Mosaic data augmentation method crops four randomly selected images and then stitches them into a new image, while the MixUp data augmentation method mixes and superimposes two randomly selected images to form a new image.
[0111] S23, importing the enhanced training set and the verification set into a preset lightweight network structure model for training, so as to obtain a target detection model in the environment surrounding the crane.
[0112] Exemplarily, the lightweight network structure is preliminarily trained using a training set, a validation set, and a loss function; the model weight parameters at the end of training are saved; the saved model weight parameters are tested using a test set, and the initial accuracy and model inference speed are recorded, and a target detection model is obtained after training.
[0113] See also Figure 5The present application further provides a crane anti-collision warning system 200, comprising:
[0114] An acquisition module 201 is configured to acquire historical video surveillance data of the crane's surrounding environment and preprocess the historical video surveillance data to obtain preprocessed historical video surveillance data.
[0115] A training module 202 is configured to import the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment;
[0116] The recognition and detection module 203 is configured to obtain real-time video surveillance data of the crane's surrounding environment and import the real-time surveillance data into the target detection model for recognition and detection to obtain target information in the crane's surrounding environment, wherein the target information includes target category information and location information;
[0117] The ranging module 204 is used to import the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane;
[0118] The early warning module 205 is configured to determine whether the real-time distance is less than a preset threshold, and if so, perform corresponding early warning processing on the crane.
[0119] See also Figure 6 An embodiment of the present application further provides a computer electronic device 400, which includes a memory 403 and a processor 402. The memory 403 stores a computer program, and when the processor executes the computer program, it implements the steps of the anti-collision warning method for a crane as described above.
[0120] Specifically, the electronic device 400 includes: a transceiver 401, a bus interface and a processor 402, the processor 402 is used to obtain historical video surveillance data of the environment around the crane, and preprocess the historical video surveillance data to obtain preprocessed historical video surveillance data; import the preprocessed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the environment around the crane; obtain real-time video surveillance data in the environment around the crane, and import the real-time monitoring data into the target detection model for identification and detection to obtain target information in the environment around the crane, wherein the target information includes: target category information and location information; import the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; determine whether the real-time distance is less than a preset threshold, and if so, perform corresponding early warning processing on the crane.
[0121] In an embodiment, the electronic device 400 further includes a memory 403. Figure 6 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits such as one or more processors represented by processor 402 and memory represented by memory 403. The bus architecture may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 401 may be a plurality of components, i.e., a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium. The processor 402 is responsible for managing the bus architecture and general processing, and the memory 403 may store data used by the processor 402 when performing operations.
[0122] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for distributing Chinese herbal medicine slices as described in any one of the above are implemented.
[0123] In this embodiment, the computer-readable storage medium may be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0124] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.
[0125] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0127] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0128] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a terminal device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0129] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A crane anti-collision warning method, characterized in that: include: Acquiring historical video surveillance data of an environment surrounding the crane, and preprocessing the historical video surveillance data to obtain preprocessed historical video surveillance data; The pre-processed historical video surveillance data is imported into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment. The preset lightweight network structure model is specifically constructed in the following steps: a loss function is designed, the loss function including a positioning loss function and a similarity measurement loss function; and a lightweight network structure model is constructed based on the loss function and the YOLOv8 algorithm. Specifically, the loss function is calculated using the following formula: Among them, F loss is the loss function, L loc is the positioning loss function, which is obtained by the position regression loss function and the generalized intersection-over-union loss function. cls is the similarity metric loss function, i is the index number of the anchor point of the i-th candidate region in the image, and p i Represents the target prediction probability of the anchor point; represents the candidate region label, N loc With N cls Respectively represent the classification loss function in the positioning loss L loc With similarity metric loss L cls Normalization coefficient after functionalization, λ represents N loc With N cls The balance weight between acquiring real-time video surveillance data of an environment surrounding the crane, and importing the real-time video surveillance data into the target detection model for recognition and detection, thereby obtaining target information of the environment surrounding the crane, wherein the target information includes: target category information and location information; Importing the target information into a preset binocular vision ranging algorithm model to obtain the real-time distance between the target object and the crane; It is determined whether the real-time distance is less than a preset threshold value. If so, a corresponding early warning process is performed on the crane.
2. The crane anti-collision warning method according to claim 1, characterized in that: The acquiring of historical video surveillance data of the crane's surrounding environment and preprocessing the historical video surveillance data to obtain preprocessed historical video surveillance data includes: Acquiring historical video surveillance data of the environment surrounding the crane, and performing frame extraction processing on the historical video surveillance data to obtain a historical surveillance image dataset of the environment surrounding the crane; Image scale normalization and image standardization processing are performed on each frame of the historical monitoring image dataset to obtain a processed historical monitoring image dataset of the crane's surrounding environment.
3. The crane anti-collision warning method according to claim 1, characterized in that: The positioning loss function L loc , the calculation formula is as follows: L loc =αL reg +βL CIoU Among them L reg Represents the position regression loss function, L CIoU is the generalized intersection-over-union loss function, and α and β are hyperparameters used to adjust the functions of each part; The position regression loss function L reg The calculation formula is as follows: Among them, B is the predicted box, B' is the target box, IoU represents the ratio of the intersection and union of the target box and the real box, B C Represents the smallest enclosed box in the positional relationship between the prediction box and the target box, and U represents the union of the prediction box and the target box; The generalized intersection-over-union loss function L CIoU The calculation formula is as follows: Among them, a is the weight parameter; d represents the distance between the center points of the two boxes, which is used to penalize the center point offset; c represents the area of the minimum enclosed region, which is used to measure the proportion of the non-overlapping part; v is used to measure the consistency of the aspect ratio, w' and h' are the width and height of the target box, respectively, and w and h are the width and height of the predicted box, respectively; L cls Represents the similarity measurement loss function, which is calculated as follows: Among them, x represents the true sample label, and x' represents the detection output after the activation function.
4. The crane anti-collision warning method according to claim 1, characterized in that: The method of importing the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane surrounding environment includes: Dividing the historical video surveillance data into a training set, a validation set, and a test set; Perform data enhancement processing on the training set to obtain an enhanced training set; The enhanced training set and the validation set are imported into a preset lightweight network structure model for training, so as to obtain a target detection model in the environment surrounding the crane.
5. The crane anti-collision warning method according to claim 1, characterized in that: The determining whether the real-time distance is less than a preset threshold, and if so, performing corresponding early warning processing on the crane, includes: When the real-time distance is less than or equal to the first set distance, a voice warning is issued; When the real-time distance is less than or equal to a second set distance, the crane is controlled to stop running.
6. A crane anti-collision warning system, characterized in that: include: an acquisition module, configured to acquire historical video surveillance data of the crane's surrounding environment and preprocess the historical video surveillance data to obtain preprocessed historical video surveillance data; The training module is used to import the pre-processed historical video surveillance data into a preset lightweight network structure model for training to obtain a target detection model in the crane's surrounding environment. The preset lightweight network structure model is specifically constructed in the following steps: designing a loss function, which includes a positioning loss function and a similarity measurement loss function; and constructing a lightweight network structure model based on the loss function and the YOLOv8 algorithm. Specifically, the loss function is calculated as follows: Among them, F loss is the loss function, L loc is the positioning loss function, which is obtained by the position regression loss function and the generalized intersection-over-union loss function. cls is the similarity metric loss function, i is the index number of the anchor point of the i-th candidate region in the image, and p i Represents the target prediction probability of the anchor point; represents the candidate region label, N loc With N cls Respectively represent the classification loss function in the positioning loss L loc With similarity metric loss L cls Normalization coefficient after functionalization, λ represents N loc With N cls The balance weight between an identification and detection module for acquiring real-time video surveillance data of the crane's surroundings and importing the real-time video surveillance data into the target detection model for identification and detection, thereby obtaining target information in the crane's surroundings, wherein the target information includes target category information and location information; a distance measurement module, configured to import the target information into a preset binocular vision distance measurement algorithm model to obtain the real-time distance between the target object and the crane; The early warning module is used to determine whether the real-time distance is less than a preset threshold, and if so, perform corresponding early warning processing on the crane.
7. A computer electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the anti-collision warning method for a crane according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the anti-collision warning method for a crane according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Crane barrier monitoring and prewarning method and system based on binocular vision
CN103559703A
Anti-collision early warning monitoring method for building construction group tower cranes
WO2021046843A1