Method and system for identifying abnormal state of crane based on computer vision
Through the abnormal state recognition method of cranes based on computer vision, image processing and abnormal state recognition are used to use the DE-MultKAN-Transformer model to perform image processing and abnormal state recognition, the problems of inaccurate and low efficiency of crane construction status monitoring are solved, real-time and accurate safety warning is achieved, and safety risks are reduced.
Patent Information
- Application Number
- CN202411713373.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology cannot realize real-time and accurate monitoring of crane construction status, resulting in increased safety hazards and accident risks, low manual monitoring efficiency and inability to be carried out in harsh environments.
Using a crane abnormal state recognition method based on computer vision, the DE-MultKAN-Transformer model is used for image processing and abnormal state recognition, combined with multiplication and learning activation functions, a data set is constructed and model training is carried out, and an alarm sound is issued to alert abnormal states.
Real-time and accurate monitoring of crane construction status is achieved, the complexity of manual monitoring is reduced, safety accidents are avoided, monitoring efficiency and accuracy are improved, and operator safety is ensured.
Smart Images

Figure CN120298941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crane construction, and particularly to a method and system for identifying abnormal states of cranes based on computer vision. Background Art
[0002] With the rapid development of the modern economy, various large-scale projects emerge in an endless stream, and the demand for cranes is also increasing continuously. However, due to the long construction period, cranes need to be used for a long time, and as irreplaceable machines in engineering projects, during their construction process, due to various reasons, numerous potential safety hazards are extremely likely to occur, resulting in safety accidents such as casualties. In this regard, it is necessary to monitor the state of the crane during construction in real time, and manual monitoring cannot achieve real-time monitoring. Therefore, the significance of monitoring the abnormal states that occur during crane construction is crucial.
[0003] Since manual labor cannot monitor the construction state of the crane all-weather in real time, it cannot give early warnings to different abnormal states of the crane in a timely manner, increasing the risk of accidents during the construction process, affecting the safety of the staff, and may also have a negative impact on social economy and social life. With the continuous maturity of technologies such as the Internet of Things and artificial intelligence, the state monitoring of cranes during construction is gradually developing towards the direction of intelligence. Combining object detection in the field of computer vision with crane construction and monitoring the construction state of the crane in the monitoring images to achieve automatic acquisition, intelligent analysis, and early warning of monitoring data is an effective way to achieve real-time and accurate monitoring of the crane progress. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for identifying abnormal states of cranes based on computer vision during crane construction, which can achieve automatic, accurate, and real-time monitoring of the state of the crane during construction.
[0005] To achieve the above purpose, the technical solution of the present invention is as follows: In the first aspect, the present invention provides a method for identifying abnormal states of cranes based on computer vision, and the method includes the following steps: Step 1: Acquisition of state image data during crane construction: Obtain the video recordings during the construction process of the crane and store the video recordings in the memory. These videos contain various operating states of the crane. Then, process the video data, extract the frames containing the crane, and label these frames with the labeling software to mark the pictures of abnormal states; The abnormal states are mainly divided into several situations: First, before the crane raises the boom and starts to carry heavy objects, the outriggers are not fully extended; Second, the boom of the crane is not fully raised during operation; Third, the crane overturns, the STOP lamp lights up, and the drive shaft is deformed due to unstable outriggers; The extracted images are all PNG pictures with 512×512 pixels. A dataset is constructed in sequence, and then the pictures in the dataset are divided into a training set, a test set, and a validation set; Step 2: Construct a DE-MultKAN-Transformer model based on multiplication and learnable activation functions: The DE-MultKAN-Transformer model adopts the framework of the DETR model, including a backbone network module, a MultKAN-Transformer module, and a multiplicative Kolmogorov-Arnold prediction head module; The MultKAN-Transformer module includes a position encoding operation, a multiplicative Kolmogorov-Arnold encoder MultKAN-Encoder, a target query module, and a multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder. Between the two normalization layers of the multiplicative Kolmogorov-Arnold encoder MultKAN-Encoder, a MultKAN module and a MultKAN feed-forward network are connected in sequence. At the same time, the output of the MultKAN module and the output of the MultKAN feed-forward network are connected by a residual connection and then input into the multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder through the second normalization layer; The multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder includes three normalization layers and two multi-head attention layers. A MultKAN module is set between the first normalization layer and the second multi-head attention layer. The output of this MultKAN module serves as the input to the Q branch of the second multi-head attention layer and is connected by a residual connection with the output of the second multi-head attention layer; A MultKAN feed-forward network is set between the second normalization layer and the third normalization layer, and a residual connection is made between the input and output of this MultKAN feed-forward network; In the multiplicative Kolmogorov-Arnold prediction head MultKAN-Prediction Head module, a MultKAN feed-forward network is used on each prediction branch; Classification and output are performed in the multiplicative Kolmogorov-Arnold prediction head MultKAN-Prediction Head module, and finally the predicted output result is obtained; The MultKAN feed-forward network includes a reshaping and upscaling module, a MultKAN module, and a CBL module connected in sequence; Step 3: Use the dataset constructed in Step 1 to train the DE-MultKAN-Transformer model based on multiplication and learnable activation functions. Use the trained DE-MultKAN-Transformer model to classify the images of the crane's state during construction to be classified, and based on the classification results, obtain the state of the crane.
[0006] Further, when training the DE-MultKAN-Transformer model, the change formula for the learning rate in the t-th generation is:
[0007] where, represents the current generation number being trained, represents the decay step, represents the initial learning rate, is the decay factor and is less than 1; The loss function in the training process is divided into two steps. In the first step, first find , is a set of matching pairs with the minimum matching cost, which is expressed by the formula:
[0008] where, represents the ground truth and the prediction result The matching cost between them, N is the number of matching pairs; the ground truth is defined as:
[0009] where, represents the class label, represents the i-th predicted bounding box; is defined as:
[0010] where, is the probability that its class is and { }, is the probability of the predicted box in the matching pair ; Next is to calculate the Hungarian matching loss of all matching pairs:
[0011] where, is the probability that the class in the matching pair is and ; The loss function for the BoundingBox, and its calculation formula is:
[0012] Among them, is the GIoU loss, which is used to measure the difference between the predicted bounding box and the true bounding box, and its formula is:
[0013] Among them, represents finding the area, represents the two-dimensional sequence Batch of the bounding box and the predicted box.
[0014] Furthermore, the image to be classified is input into the trained DE-MultKAN-Transformer model for recognition, the category information of the abnormal state where the crane is located is recognized, and three different alarm sounds are emitted; In the face of the first abnormal state, that is, when the crane is working with the boom lifted, the system will emit an alarm sound for the first abnormal information, and the operator of the crane will, according to the type of the alarm sound, perform leg maintenance and fully extend the legs; when the crane is in the second abnormal state, the system will emit an alarm sound for the second abnormal information, and the operator of the crane will, according to the type of the alarm sound, perform maintenance on the boom of the crane and fully raise the boom; when the third situation occurs to the crane, an alarm sound for the third abnormal information will be emitted, and the operator of the crane will, according to the type of the alarm sound, immediately stop the crane from working and notify the maintenance personnel to conduct a comprehensive inspection and adjustment of the crane to avoid the occurrence of safety accidents.
[0015] Furthermore, the backbone network module includes a CNN module and a 1×1 convolutional module. The CNN module outputs a low-resolution feature map f with multiple channels, and its size is × × , where = 2048, = H / 32, = W / 32, where H is the height of the original input image and W is the width of the original input image; then the feature map f is input into a 1×1 convolutional module, and this convolutional module will reduce its number of channels to , and obtain a new feature map , and its size is × × .
[0016] Furthermore, the MultKAN module includes a 3×3 convolutional layer, a MultKAN structure, a depth convolutional layer, and a normalization layer. The vector obtained after processing the input feature sequence through the 3×3 convolutional layer, the MultKAN structure, and the depth convolutional layer is subjected to residual connection with the input feature sequence, and then the output of the MultKAN module is obtained after passing through the normalization layer.
[0017] Furthermore, the CBL module contains 3 convolutional blocks. The convolutional kernel sizes of the first two convolutional blocks are 3×3, while the convolutional kernel size of the third convolutional block is , and both of the first two convolutional blocks include batch normalization and Leaky_ReLU activation function operations.
[0018] In a second aspect, the present invention provides a crane abnormal state recognition system based on computer vision. The system executes the steps of the method and includes a processor, a display, and a memory.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention creatively adds a MultKAN module with multiplication and learnable activation functions to the DETR model, constructs a new DE-MultKAN-Transformer model based on learnable activation functions, which can solve the problems of inaccurate state monitoring, poor timeliness, and huge workload during traditional crane construction, replaces manual monitoring with machine monitoring, avoids potential life safety problems in manual operations, avoids the situation where workers cannot work in bad weather environments, avoids the situation where workers cannot monitor the crane construction state due to environmental factors, and avoids the situation where instruments are damaged due to improper operation.
[0020] The present invention introduces the MultKAN structure to achieve active learning of the model, reduces the complexity of crane construction state monitoring, can reduce the workload, and can accurately and efficiently monitor the state of crane construction. Then, it judges whether the crane is in a safe state according to the state, and issues a warning message according to the type of state, realizing real-time monitoring and providing an effective means for crane construction state monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic structural diagram of the DE-MultKAN-Transformer model in the present invention.
[0022] Figure 2 is a schematic structural diagram of the MultKAN-Transformer module in the present invention.
[0023] Figure 3 is a schematic structural diagram of the MultKAN structure in the present invention.
[0024] Figure 4 It is a schematic structural diagram of the MultKAN module in the present invention.
[0025] Figure 5 It is a schematic structural diagram of the MultKAN feedforward network (MultKAN-FFN) in the present invention.
[0026] Figure 6 It is a schematic flow diagram of the method of the present invention. Specific embodiments
[0027] The present invention will be further explained below in conjunction with embodiments and the accompanying drawings, but this is not used as a limitation on the protection scope of the present application.
[0028] Embodiment 1 The abnormal state recognition method of the crane based on computer vision in this embodiment has the overall method as follows: collecting the state image data during the crane construction, using the DE-MultKAN-Transformer model based on multiplication and learnable activation functions to train the images of the crane construction, identifying the state of the crane, and judging whether the operation of the crane is in a safe operation state. The method includes the following steps: Step 1: Collecting the state image data during the crane construction: A shooting machine with a fixed height is installed outside the crane construction site to shoot the crane construction process. The shooting machine can rotate 360 degrees for shooting. In this embodiment, the shooting machine uses a high-precision camera, which can capture the state of the crane working in the construction site. The shooting machine will record the crane operation process in real time and store the video in the memory. These videos contain various operation states of the crane, such as when the crane is in the state of extending the outriggers and when the crane is in the state of opening the boom for work. Then, the video data is processed to extract the frames containing the crane, and these frames are labeled with the labeling software, that is, select a picture in the labeling interface and select the format to be saved, then draw a bounding box for the target and classify its labels, and finally save it. These pictures are randomly shuffled, and the pictures in the abnormal state are marked among 144 pictures of the crane images every day, including when the crane is in the state of extending the outriggers and when the crane is in the state of opening the boom for work, as well as the possible abnormal postures of the crane, such as the crane tipping over caused by unstable outriggers, and some functions of the crane may also have errors, such as the STOP light being on and the drive shaft being deformed.
[0029] The abnormal states are mainly divided into several situations. First, before the crane raises the boom and starts to carry heavy objects, the outriggers are not fully extended. Second, the boom of the crane is not fully raised during operation. Third, the crane tips over due to unstable outriggers, the STOP light is on, and the drive shaft is deformed, etc.
[0030] Repeat this process for 3 months, and a dataset consisting of 12,960 images can be obtained in total. All the extracted images are PNG images with a size of 512×512 pixels. Then, the images are divided into a training set, a test set, and a validation set according to the ratio of 8:1:1. The number of images in each set is 10,368, 1,296, and 1,296 respectively. The images of the three abnormal states have the same proportion in the three sets.
[0031] Step 2: Build a DE-MultKAN-Transformer model based on multiplication and learnable activation functions: The DE-MultKAN-Transformer model adopts the framework of the DETR (Detection Transformer) model, and its overall structure is as Figure 1 shown, including a backbone network module, a MultKAN-Transformer module, and a MultKAN-Prediction Head module.
[0032] The backbone network module mainly includes two sub-modules, which are a convolutional neural network (CNN) module and a 1×1 convolutional module from left to right.
[0033] The input image is input into the convolutional neural network module to obtain a feature extraction map, and then input into the 1×1 convolutional module to obtain a new feature extraction map.
[0034] The MultKAN-Transformer module includes a position encoding operation, a MultKAN-Encoder, an ObjectQueries module, and a MultKAN-Decoder.
[0035] The new feature extraction map output by the backbone network is mapped into a vector, which is merged with the vector processed by the position encoding operation to form a new vector. Then, it is input into the MultKAN-Encoder. The results output by the ObjectQueries module and the MultKAN-Encoder are input into the MultKAN-Decoder together. The result output by the MultKAN-Decoder is input into the MultKAN-Prediction Head module for classification and output, and finally the predicted output result is obtained.
[0036] Step 3: Train the DE-MultKAN-Transformer model based on multiplication and learnable activation functions using the dataset constructed in Step 1: First, randomly initialize the hyperparameters (initial weights, learning rate, etc.) of the DE-MultKAN-Transformer model, and use the training set to train the constructed DE-MultKAN-Transformer model for 500 epochs. After each epoch of training, use the validation set to verify and understand the generalization ability of the deep learning model, while preventing overfitting.
[0037] The learning rate determines the speed and direction of updating the weight parameters of the model during training. Specifically, in each iteration, the DE-MultKAN-Transformer model calculates the gradient of the loss function with respect to each parameter. This gradient indicates the direction in which the parameter should be adjusted to minimize the loss, and the learning rate is the magnitude of the adjustment during the adjustment process, that is, the amount of parameter update. If the learning rate is set too large, then in each iteration, the DE-MultKAN-Transformer model parameters may miss the optimal solution, resulting in oscillation or divergence; if the learning rate is set too small, the speed at which the DE-MultKAN-Transformer model converges to the optimal solution will be very slow, and it may fall into a local minimum instead of the global optimal solution. In this embodiment, an exponential decay strategy is adopted to set the learning rate, that is, at the beginning of model training, a relatively high learning rate is first set to enable the model to quickly converge, and then the learning rate is gradually decreased according to the exponential decay strategy to enable the model to reach the optimal solution. The change formula of the learning rate in the t-th epoch is as follows:
[0038] Where, represents the current number of training epochs, represents the decay steps, represents the initial learning rate, is the decay factor, usually less than 1. As the exponent increases, will gradually decrease according to the exponential law, which helps to prevent overfitting.
[0039] Secondly, use the loss function to adjust the parameter model. The loss function of the model of the present invention is divided into two steps. The first step is to first find , is a set of matching pairs with the minimum matching cost, which is expressed by the formula:
[0040] Where, represents the true value and the prediction result The matching cost between them, N is the number of matching pairs, and the true value among them can be defined as:
[0041] Among them, is expressed as the class label, represents the i-th predicted bounding box. can be defined as:
[0042] Among them, is the probability of defining its class as and { }, is the probability of the predicted box in the matching pair .
[0043] Next is to calculate the Hungarian matching loss of all matching pairs:
[0044] Among them, is the probability that the class in the matching pair is and . is the loss function of the BoundingBox, and the formula is:
[0045] Among them, is the GIoU loss, which is used to measure the difference between the predicted bounding box and the true bounding box. The formula is as follows:
[0046] Among them, represents finding the area, represents the two-dimensional sequence Batch of the bounding box and the predicted box.
[0047] Set up a visualization chart to display the curves of accuracy and loss in real time. During the training process, after each epoch is trained, the current accuracy and time consumption will be output. By observing the change trend of the curve, analyze whether the convergence and accuracy of the DE-MultKAN-Transformer model meet the expected requirements. If the DE-MultKAN-Transformer model cannot converge or the accuracy is low after convergence, then it is necessary to adjust by adjusting the hyperparameters. When the value of the loss function remains relatively stable within multiple training cycles and no longer decreases significantly, it means that the DE-MultKAN-Transformer model has reached convergence and stop training.
[0048] Step 4: Use the trained DE-MultKAN-Transformer model to classify the images of the crane's state during construction, and based on the classification results, obtain the state of the crane: Input the image to be classified into the trained DE-MultKAN-Transformer model for recognition, and obtain information on three categories of abnormal states of the crane and emit three different alarm sounds. By classifying the output results of the model, it is possible to identify which abnormal state the crane is in. For example, whether the outriggers of the crane are extended during operation, whether the boom of the crane is opened during operation, whether the crane overturns, and the state of the STOP light being on, etc.
[0049] For example, based on the identified state of the crane, it can be analyzed whether the crane moves without fully extending its outriggers, and this state is used to issue a warning through the alarm system.
[0050] According to the types of warnings issued, perform maintenance and adjustment on the corresponding positions of the crane. In the face of the first abnormal state, that is, when the crane is working with the boom lifted, the system will emit an alarm sound for the first abnormal information. The operator of the crane will perform outrigger maintenance based on the type of the alarm sound and fully extend the outriggers; when the crane is in the second abnormal state, the system will emit an alarm sound for the second abnormal information. The operator of the crane will perform boom maintenance based on the type of the alarm sound and fully raise the boom; in the third situation of the crane, an alarm sound for the third abnormal information will be emitted. The operator of the crane will immediately stop the crane operation according to the type of the alarm sound and notify the maintenance personnel to conduct a comprehensive inspection and adjustment of the crane to avoid the occurrence of safety accidents.
[0051] In this embodiment, the structures of each part of the DE-MultKAN-Transformer model are as follows: Input an image of size C×H×W into the backbone network module, where C = 3 is the number of channels, H is the height of the image, and W is the width of the image. The convolutional neural network module will output low-resolution feature maps f of multiple channels, and its size is × × , where = 2048, = H / 32, = W / 32. Then the feature map f is input into a 1×1 convolutional module, and this convolutional module will reduce its number of channels to , and obtain a new feature map , and its size is × × .
[0052] Feature map Before entering the multiplicative Kolmogorov - Arnold encoder, an Embedding operation is first performed, which maps the input into the form of a vector. The Embedding operation is divided into two parts. The first part is the InputEmbedding operation, which maps the output feature map into a vector a. At this time, the size of vector a is × ×256. The second part is the Positional Encoding operation, which is a set of vectors b with the same dimension as the vector after the Input Embedding operation, used to provide position information. Regarding the Positional Encoding operation, its encoding rules are as follows:
[0053]
[0054] Among them, ( , ) is a certain position in the feature map, , , i represents the dimension of this position. Substituting into the two encoding formulas, a 128 - dimensional vector can be calculated, representing the positional encoding; substituting into the latter two encoding formulas, a 128 - dimensional vector can also be calculated, representing the positional encoding; concatenating the two 128 - dimensional vectors, a 256 - dimensional vector can be obtained, representing the positional encoding of ( , ). Calculating all the positional encodings, a ×256 vector is obtained, representing the positional encoding of this Batch. At this time, the dimension of the encoding is × ×256. Next, it needs to be added to vector a and then changed to a vector through the reshape operation. Then, vector a and vector b are added to get vector c, which is the key K and query Q in the multi - head attention layer of the multiplicative Kolmogorov - Arnold encoder. This multi - head attention layer can simultaneously focus on different parts of the input vector c and map the query, key, and value into multiple different linear spaces respectively. At this time, the size of vector c is 。The vector a serves as the value V of the multi-head attention layer simultaneously. For each set of inputs (query, key, value) of the vectors c and a, four steps are carried out. First, the query Q, key K, and value V are obtained, which can be expressed by the formula:
[0055]
[0056]
[0057] Among them, , and are weight matrices corresponding to different heads.
[0058] The second step is that each head calculates the similarity between the query and the key and uses it as the attention score. It can be expressed by the formula:
[0059] Among them, represents the attention score at the i-th position in the feature map, softmax is the function, represents the query vector at the i-th position in the feature map, represents the key vector at the i-th position in the feature map, represents the dimension of the key vector. Dividing by the square root of the dimension here is to stabilize the gradient and prevent the value from being too large.
[0060] The third step is to apply the attention score to the Value. Each head has its own score, so a summation is performed for each head. It can be expressed by the formula:
[0061] Among them, represents the result of the weighted summation at the i-th position, represents the value Value.
[0062] Finally, the results of all heads are concatenated and integrated into the final output vector through another linear transformation. The multi-head attention layer mechanism can capture richer dependencies. Each head processes a part of the information of the input vectors c and a, and merging the results of all heads can enhance the understanding of the global context. The finally output result is merged with the input vector a and output to the normalization layer.
[0063] The normalization layer is usually a Layer Normalization operation, that is, the result of addition is normalized. Here, an Add operation is added. The Add operation adds the output of the multi-head attention layer and the input vector a. This layer can not only normalize the vector sequence, maintain the consistency of data distribution, but also help accelerate the training of the model and improve the stability of the model. The normalization layer can also reduce the impact of internal covariate shift, making the model more sensitive to changes in the input vector. The normalized result then enters the MultKAN module, where the feature sequence is processed by a 3×3 convolutional layer, a MultKAN structure, and a depth convolutional layer. The depth convolutional layer is a convolutional kernel with the same number of channels and size as the input feature sequence. Then, the processed output result is connected with the feature sequence through a residual connection, and after passing through the normalization layer, the output of the MultKAN module is obtained.
[0064] The MultKAN structure here consists of a standard KAN layer and a multiplication layer The standard KAN layer is based on the Kolmogorov-Arnold representation theorem, that is, a multivariate continuous function can be represented as a finite combination of univariate continuous functions. The theorem can be expressed by the formula:
[0065] where, is a one-dimensional function, p is the index of the input dimension, q is the index of the output dimension, and n is the number of parameters referred to. The standard KAN layer is expressed by the formula:
[0066] where, represents a one-dimensional function, l represents the depth, belongs to the integer array , represents the input vector, and and . The l-th KAN layer has input dimensions and output dimensions, and converts the input vector to . The activation function inside the node in the KAN layer is the sum of the basic function b(x) and the spline function, which can be expressed by the formula:
[0067] where, w is used to control the overall size of the activation function. The basic function is set to:
[0068] The spline function is parameterized as a linear combination of B-splines:
[0069] where, is the B-spline function. And the whole network consists of standard KAN layers, that is
[0070] where, is the standard KAN layer. The addends and multiplicands in the l-th layer are denoted as and , respectively, where , , , , , , , . Accepts the input vector , and converts to the output , that is .
[0071] The multiplication layer consists of two parts. The multiplication part performs multiplication on pairs of child nodes, while the other part performs an identity transformation. The multiplication layer is written in Python and converts in the following way:
[0072] where, is concatenation. Therefore, the MultKAN layer can be concisely expressed by the formula: , and the entire MultKAN structure is expressed by the formula:
[0073] The activation function within the node in the MultKAN structure reduces the convolution, pooling, and other operations on sensitive data in the overall model by avoiding errors in sensitive data input, thereby further improving the accuracy of the output image of the MultKAN structure and allowing the MultKAN structure to achieve a mapping from complex multi-dimensional inputs to multi-dimensional outputs at each layer. The multiplication module therein can clearly reveal the multiplication structure in the data, enhancing the interpretability and expressive ability of the model.
[0074] The feature sequence generated by the MultKAN module is d. The feature sequence d enters the MultKAN Feed Forward Network. In the MultKAN Feed Forward Network, the reshaping and upgrading module reshapes the feature sequence d and refines it. After that, it passes through a MultKAN module, which performs further convolution operations on the feature sequence d to output the feature d1, and finally enters the CBL module. The CBL module contains 3 convolution blocks (CnovolutionBlock). Among them, the convolution kernels of the first two convolution blocks are 3×3, and the convolution kernel of the third convolution block is . The CBL module is represented by the mathematical formula:
[0075]
[0076]
[0077] Among them, d1 is the output of the MultKAN module in the MultKAN Feed Forward Network, is the output of the first convolution block in the CBL module, is the output of the second convolution block in the CBL module, and are the hidden layers of the first and second convolution blocks respectively. G is the third convolution block, is the Leaky_ReLU activation function, is the batch normalization operation, is the number of input channels, with a size of 3.
[0078] After performing a residual connection between the feature sequence d and the output G of the MultKAN-Feed Forward Network and passing through the normalization layer for normalization operation, the feature sequence A is obtained.
[0079] Next is the Object Queries module. This module predefines the number of object queries of the multiplicative Kolmogorov-Arnold decoder based on the feature sequence A output by the multiplicative Kolmogorov-Arnold encoder MultKAN-Encoder, that is, finally predicts the category of the query target and the position of the bounding box. The default number is 100. The object queries module is a learnable position encoding that is added to the input of each multi-head attention layer. The addition of object queries is because each multiplicative Kolmogorov-Arnold decoder layer needs to decode N objects in parallel, and the multiplicative Kolmogorov-Arnold decoder is invariant, so N input embeddings must be different to produce different output results. Therefore, learnable position encoding also needs to be added to the input of each multi-head attention layer. The position encoding output by the object queries enters the multiplicative Kolmogorov-Arnold decoder as a vector d2 with a dimension of 100× ×256. The vector d2 is input into the first multi-head attention layer in the multiplicative Kolmogorov-Arnold decoder, and this layer will perform the same operations as the multi-head attention layer in the multiplicative Kolmogorov-Arnold encoder on the vector d2, calculate the attention scores, perform weighted summation, and finally merge the results output by all heads. Then, the result output by the multi-head attention layer and the vector d2 are subjected to residual connection and normalized by the normalization layer. The normalized result is input into a MultKAN module. In the MultKAN module, the feature sequence is sequentially input into the Tokenization layer, 3×3 convolutional layer, MultKAN structure, depth convolutional layer, and normalization layer, and the output feature sequence is B, with a size of 100× ×256.
[0080] Next, the feature sequence B enters the lower multi-head attention layer, merges with the object queries, and is transformed into Q values, while the vector b merges with the feature sequence A and is transformed into K values, and the V values only come from the feature sequence A. Next, the multi-head attention layer will calculate the attention scores, perform weighted summation and result concatenation in sequence, and output the concatenated result to the lower normalization layer. At this time, the feature sequence B and the result output by the multi-head attention layer are output to the next normalization layer together. The normalization layer first performs residual connection on these two sequences, and then normalizes the result after residual connection. Then, the normalized result is output to the lower multiplicative Kolmogorov-Arnold feed-forward network (MultKAN-FeedForward Network). After passing through the MultKAN feed-forward network, the output feature sequence and the input of the MultKAN feed-forward network are input into a normalization layer together for residual and normalization operations, and 100 targets are output.
[0081] Next, these 100 targets will respectively enter a MultKAN feed-forward network in the MultKAN-Prediction Head module. Each MultKAN feed-forward network will independently decode the coordinates of the bounding box and the class label, generating N final prediction results, where N is a constant much larger than the number of targets in a normal image. Next, the N prediction results and the ground truth boxes are subjected to bipartite matching through the Hungarian algorithm. That is, if there are 100 targets, then 100 of the N prediction results will be able to match these 100 ground truths, and the others will all match successfully with "no.object". Then, the elements of the prediction set and the real set are made to correspond one by one to minimize the matching loss.
[0082] The present invention uses the MultKAN module in the encoder, which can not only avoid errors in the input of sensitive data, but also further improve the accuracy of the output image, and can also achieve a complex mapping from multi-dimensional input to multi-dimensional output on each layer. By adding the MultKAN module to the traditional DETR model, not only is the MultKAN structure applied, enabling the model to better capture spatially related features, but also local attention is emphasized, which helps to maintain the model's attention to details, dynamically switch between local and global information, increase the model's adaptability to input changes, contribute to enhancing the model's understanding of complex scenes, effectively fuse image features with text information (such as object class labels), enhance the model's understanding ability of multi-modal information, adjust the self-attention mechanism in the model to help the model more effectively focus on key regions in the image, improve the accuracy of model object detection, and also improve the way of feature extraction and integration to enhance the model's recognition ability for targets of different scales.
[0083] Table 1 below shows the comparison results of the effects of training different types of models using the dataset of the present invention. In the following table represents the average precision for measuring the model at different recall rate levels. represents when the IoU threshold is set to the average precision of the model detection results, reflecting the performance of the model under relatively loose matching conditions. represents when the IoU threshold is set to the average precision of the model detection results, reflecting the performance of the model under relatively strict matching conditions.
[0084]
[0085] The DE-MultKAN-Transformer model of the present invention aims to solve the problems existing in current object detection models, such as low detection accuracy, high computational complexity, non-intuitive feature fusion, strong dependence on data, poor interpretability, high hardware requirements, and many parameter requirements in crane construction. By adding an activation function with a multiplication structure and self-learning in a suitable way, the strong expressive ability and interpretability of the MultKAN structure can be combined with the global information processing ability of DETR, improving the accuracy of object detection. At the same time, its modular structure can be combined with the encoder-decoder architecture of DETR, optimizing the feature extraction and prediction processes, reducing the model training time, improving the model convergence speed and the efficiency of computing the input image, as well as enhancing the model's adaptability and accurate fitting ability, while improving the model's generalization ability and learnability, and enhancing the accuracy of crane construction status monitoring.
[0086] The present invention adopts the DE-MultKAN-Transformer model, which does not require a masking part, greatly improving the accuracy of object recognition and judgment, enabling the system to monitor the status of crane construction in real time and make corresponding early warnings and adjustments in a timely manner, ensuring the correctness and timeliness of image feature extraction, making the model more reliable and effective, improving the efficiency of the model, and enhancing the learnability of the model, being able to learn new things from data. That is, through multiplication and learnable activation functions, it can effectively learn the spatial relationship and semantic information between features, provide richer feature interaction and expression ability, provide more complex computing ability, help the model better capture key information in the image, thereby improving the detection accuracy and efficiency and enhancing the accuracy of object detection. It can accurately identify the status of the crane during operation in real time, greatly reducing the safety risks during the crane operation process and ensuring the safety of people's lives and property.
[0087] Matters not described in the present invention are applicable to the prior art.
Claims
1. A method for identifying abnormal states of a crane based on computer vision, characterized in that, The method includes the following steps: Step 1, acquisition of state image data during crane construction: Obtain the video recording during the crane construction, and store the video recording in the memory. These videos contain various operating states of the crane. Then, process the video data, extract the frames containing the crane, and label these frames with the labeling software to mark the pictures of abnormal states; The abnormal states are mainly divided into several situations: First, before the crane raises the boom and starts to carry heavy objects, the outriggers are not fully extended; Second, the boom of the crane is not fully raised during operation; Third, the crane overturns, the STOP light is on, and the drive shaft is deformed due to unstable outriggers; The extracted images are all PNG pictures with 512×512 pixels. Build a data set in sequence, and then divide the pictures in the data set into a training set, a test set, and a validation set; Step 2, construct a DE-MultKAN-Transformer model based on multiplication and learnable activation functions: The DE-MultKAN-Transformer model adopts the framework of the DETR model, including a backbone network module, a MultKAN-Transformer module, and a multiplicative Kolmogorov-Arnold prediction head module; The MultKAN-Transformer module includes a position encoding operation, a multiplicative Kolmogorov-Arnold encoder MultKAN-Encoder, an object query module, and a multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder. A MultKAN module and a MultKAN feed-forward network are sequentially connected between the two normalization layers of the multiplicative Kolmogorov-Arnold encoder MultKAN-Encoder. At the same time, the output of the MultKAN module and the output of the MultKAN feed-forward network are connected by a residual connection and then input into the multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder through the second normalization layer; The multiplicative Kolmogorov-Arnold decoder MultKAN-Decoder includes three normalization layers and two multi-head attention layers. A MultKAN module is set between the first normalization layer and the second multi-head attention layer. The output of this MultKAN module is used as the input of the Q branch of the second multi-head attention layer and is connected with the output of the second multi-head attention layer by a residual connection; A MultKAN feed-forward network is set between the second normalization layer and the third normalization layer, and a residual connection is made between the input and output of this MultKAN feed-forward network; In the multiplicative Kolmogorov-Arnold prediction head MultKAN-Prediction Head module, a MultKAN feed-forward network is used on each prediction branch; Classification and output are performed in the multiplicative Kolmogorov-Arnold prediction head MultKAN-Prediction Head module, and finally the predicted output result is obtained; The MultKAN feedforward network includes a reshaping and upgrading module, a MultKAN module, and a CBL module connected in sequence; Step 3: Use the dataset constructed in Step 1 to train the DE-MultKAN-Transformer model based on multiplication and learnable activation functions. Use the trained DE-MultKAN-Transformer model to classify the images of the crane's working state to be classified, and based on the classification results, obtain the state of the crane.
2. The method according to claim 1, wherein When the DE-MultKAN-Transformer model is trained, the change formula for the learning rate in the t-th generation is: Among them, represents the current generation number being trained, represents the decay step number, represents the initial learning rate, is the decay factor and is less than 1; The loss function of the training process is divided into two steps. In the first step, first calculate , as a set of matching pairs with the minimum matching cost, which is expressed by the formula: Among them, represents the true value and the matching cost between the prediction result where N is the number of matching pairs; the true value among them is defined as: Among them, is represented as a class label, represents the i-th predicted bounding box; Defined as: Among them, is the probability of defining its category as and { }, is the probability of the predicted box in the matching pair ; Next is to calculate the Hungarian matching loss for all matching pairs: Among them, is the matching pair in which the category is with a probability of ; is the loss function of the Bounding Box, and the calculation formula is: Among them, is the GIoU loss, which is used to measure the difference between the predicted bounding box and the ground truth bounding box, and its formula is: Among them, represents calculating the area, represents a two-dimensional sequence Batch of the bounding box and the prediction box.
3. The method according to claim 1, wherein Input the image to be classified into the trained DE-MultKAN-Transformer model for recognition, recognize the category information of the abnormal state of the crane, and emit three different alarm sounds; In the face of the first abnormal state, that is, when the crane is working with the boom lifted, the system will emit an alarm sound for the first abnormal information. The operator of the crane will, according to the type of the alarm sound, perform leg maintenance and fully extend the legs; when the crane is in the second abnormal state, the system will emit an alarm sound for the second abnormal information. The operator of the crane will, according to the type of the alarm sound, perform boom maintenance and fully raise the boom; when the crane is in the third situation, an alarm sound for the third abnormal information will be emitted. The operator of the crane will, according to the type of the alarm sound, immediately stop the crane and notify the maintenance personnel to conduct a comprehensive inspection and adjustment of the crane to avoid the occurrence of safety accidents.
4. The method according to claim 1, characterized in that The backbone network module includes a CNN module and a 1×1 convolutional module. The CNN module outputs low-resolution feature maps f of multiple channels, with a size of × × , where = 2048, = H / 32, = W / 32, where H is the height of the original input image and W is the width of the original input image. Then, the feature map f is input into a 1×1 convolutional module, which reduces its number of channels to , and obtains a new feature map , with a size of × × .
5. The method according to claim 1, characterized in that The MultKAN module includes a 3×3 convolutional layer, a MultKAN structure, a depth convolutional layer, and a normalization layer. The vector obtained after the input feature sequence is processed by the 3×3 convolutional layer, the MultKAN structure, and the depth convolutional layer is connected in residual with the input feature sequence, and then the output of the MultKAN module is obtained after passing through the normalization layer.
6. The method according to claim 1, characterized in that, The CBL module contains 3 convolutional blocks. The convolutional kernels of the first two convolutional blocks are 3×3 in size, while the convolutional kernel of the third convolutional block is , and both of the first two convolutional blocks include batch normalization and Leaky_ReLU activation function operations.
7. A crane abnormal state recognition system based on computer vision, characterized in that, The system executes the steps of the method according to any one of claims 1-6, and includes a processor, a display, and a memory.