Building constructor operation violation behavior monitoring method based on artificial intelligence
By using the combined technology of composite perturbation generation adversarial network, adaptive fluctuation optimization neural network, manifold mapping learning self-coded neural network and meta-learning modulation limit learning machine algorithm in the construction site monitoring, the problems of insufficient data and weak model generalization capabilities in complex environments in the construction site are solved, efficient feature extraction, dimensionality reduction and classification are achieved, and monitoring accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202411899474.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
AI Technical Summary
In the complex environment of monitoring the construction site, the prior art faces the problems of insufficient data and weak model generalization capabilities, and is inefficient and consumes high computing resources when processing large-scale or complex data.
The image data is expanded by a generative adversarial network algorithm based on composite perturbation, feature extraction is performed through neural networks based on adaptive fluctuation optimization, feature dimensionality reduction is performed using an auto-coded neural network based on manifold mapping learning, and classification training is performed through an adaptive modulation limit learning machine algorithm based on meta-learning.
It improves the diversity of data and the generalization ability of the model, optimizes the oscillation and local minimum value problems in the traditional gradient descent method, ensures the data retention of key information after dimensionality reduction, and improves the accuracy and efficiency of data reconstruction.
Smart Images

Figure CN119992442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence model training, and in particular to a method for monitoring construction workers' illegal operations based on artificial intelligence. Background Art
[0002] With the development of technology, especially the advancement of artificial intelligence and machine learning technology, it has become feasible to use these technologies to conduct automated safety monitoring at construction sites. By collecting video data through cameras installed at the construction site and automatically analyzing this data using intelligent algorithms, the safety behavior of construction workers can be monitored in real time, such as whether they are wearing safety helmets and seat belts, thereby improving the efficiency and effectiveness of safety management.
[0003] For example, the Chinese invention patent with application number CN202010201329.X proposes a method and system for identifying the wearing of a safety helmet. Through the method provided by the present invention, the image can be scaled, and feature point detection can be performed on the basis of the scale division to determine the final detection frame. Therefore, through this method, the existing construction site video stream can be fully utilized to realize the wearing identification of the safety helmet. There is no need to install other monitoring equipment, which saves costs. In addition, the algorithm can ensure the stability, robustness and high precision of the results, improve the accuracy, and the recognition efficiency is also improved.
[0004] The Chinese invention patent with application number CN202010447288.2 proposes a construction worker identification and helmet wearing detection method and system, including the following steps: collecting construction worker identity information and worker images monitored by real-name system entrance; adding labels to images in the public data set SHWD and normalizing them; building a system architecture based on YOLO V3 and Light cnn-29 neural network algorithms; training YOLO V3 on the public data set SHWD, inputting the labeled images, detecting helmets and face parts; training Light CNN-29; testing the accuracy and recall rate of recognition under various construction environments, and correcting them; using the corrected neural network model, authenticate the identity of personnel entering the construction site, and detect helmets for personnel entering the construction site. The present invention uses a deep learning algorithm to authenticate the identity of personnel entering the construction site, detect helmets for personnel entering the construction site, prevent non-project personnel and those who do not wear helmets from entering the construction site, and effectively improve the safety management capabilities of the construction site.
[0005] The Chinese invention patent with application number CN202311825131.9 proposes a safety warning method and medium based on smart helmets, which are applied to smart helmets, including: the smart helmet is connected to the management platform and bound through the communication module; the question voice is recorded through the audio module, and the management platform uses artificial intelligence application to receive and return the answer voice to achieve interaction; the virtual image is kept consistent with the real inspection equipment through the MR head display device to achieve interaction; the camera and the management platform are used to make video calls and return the video; the gas detection module is used to detect the environment around the smart helmet and issue an early warning; the Beidou positioning module is used to determine the location of the wearer and issue an early warning; the acceleration detection module and the infrared detection module are used to detect whether the wearer falls or takes off the hat and issue an early warning; the optical detection module and the infrared detection module are used to detect the wearer's blood oxygen level and body temperature, and issue an early warning. The present invention is more intelligent and ensures the safety of the wearer.
[0006] The prior art has the following deficiencies:
[0007] 1. When faced with complex or changing construction scenarios, traditional monitoring methods are prone to poor generalization capabilities due to insufficient data or overly simple models.
[0008] 2. Existing technologies often suffer from low efficiency and high consumption of computing resources when processing large-scale or complex data, especially in the feature extraction and dimensionality reduction stages.
[0009] 3. Traditional neural networks often encounter oscillation and local optimal problems during training, which affects the training effect and model stability. Summary of the invention
[0010] In order to overcome the defects existing in the prior art, the present invention provides a method for monitoring the illegal behavior of construction workers based on artificial intelligence, which improves the diversity of data, enhances the generalization ability of the model, optimizes the common oscillation and local minimum problems in the traditional gradient descent method, makes the parameter updating process smoother and more adaptive, ensures that the data after dimensionality reduction can still retain key information, and improves the accuracy and efficiency of data reconstruction.
[0011] In order to achieve the above object, the present invention provides a method for monitoring construction workers' illegal operations based on artificial intelligence, comprising the following steps:
[0012] Collect image data of the construction site and mark different types of work violations in the image data;
[0013] The collected image data is expanded using a generative adversarial network algorithm based on composite perturbations;
[0014] Constructing a feature extraction model, inputting the expanded image data into the feature extraction model and performing feature extraction training by using a neural network feature extraction algorithm based on adaptive fluctuation optimization to obtain image data after feature extraction;
[0015] Construct a feature dimensionality reduction model, input the image data after feature extraction into the feature dimensionality reduction model, and perform feature dimensionality reduction training by using an autoencoder neural network based on manifold mapping learning to obtain the image data after feature dimensionality reduction;
[0016] A classifier model is constructed, the image data after feature dimension reduction is input into the classifier model, and classification training is performed by using an adaptive modulation extreme learning machine algorithm based on meta-learning to obtain a category result consistent with the labeled category of the work violation;
[0017] The trained feature extraction model, feature dimension reduction model and classifier model are used to process the images of the real-time construction site to monitor the work violations of construction workers and identify the corresponding categories.
[0018] Preferably, different categories of work violations are marked by manual marking, and the marked categories include: not wearing a safety helmet, not using a safety belt, and normal wearing.
[0019] Preferably, the composite perturbation-based generative adversarial network algorithm includes a generator and a discriminator, wherein the generator includes a combined weighted convolutional layer.
[0020] Preferably, when the collected image data is expanded using a composite perturbation-based generative adversarial network algorithm, the expansion process includes:
[0021] Initialize and set the network parameters of the generator and discriminator;
[0022] The generator receives random noise and outputs a synthetic image. The discriminator determines whether the input image data is real or synthetic, and calculates the loss function that minimizes the generator and the discriminator. The calculation method of the loss function is expressed as:
[0023]
[0024] Where bx G represents the generator, bx D represents the discriminator, bx~p data Indicates sampling from the real data distribution, bz~p z represents sampling from a random noise distribution, represents the expectation of the real image, represents the expectation of random noise;
[0025] In the generator, the performance of the generator is enhanced by adjusting the weights of the combined weighted convolution layer to reduce the difference between the generated image and the real image. The weight adjustment method of the combined weighted convolution layer is expressed as:
[0026]
[0027] In the formula, represents the weight at iteration t, η vk represents the learning rate, Represents the loss function About weight α bx The gradient of
[0028] The discriminator provides feedback after each training, and adjusts the network parameters of the generator according to the feedback of the discriminator. The adjustment method is expressed as:
[0029]
[0030] In the formula, λ gs represents the adjustment factor, represents the loss function of the discriminator, represents the gradient of the loss function with respect to the generator network parameters;
[0031] Repeat the above steps until the preset stop iteration condition is met.
[0032] Preferably, when the performance of the generator is enhanced by adjusting the weights of the combined weighted convolutional layer in the generator, the loss function is set The gradient relative to the output of one of the convolutional layers in the combined weighted convolutional layer is g out , and the input of the convolutional layer is bx in , then the gradient of the weight The calculation method is expressed as:
[0033]
[0034] Where ⊙ represents the Hadamard product.
[0035] Preferably, the training process of the neural network feature extraction algorithm based on adaptive fluctuation optimization includes:
[0036] Define the neural network structure and initialize the parameters of the neural network. The parameters of the neural network include weights and biases, and the initialization setting method is expressed as:
[0037] θ0={w0,b0}
[0038] In the formula, θ0 represents the initial parameters of the neural network, w0 represents the initial weight, and b0 represents the initial bias;
[0039] Use the parameters to perform forward propagation on the input to get the output value, compare the output value with the target value to calculate the loss, and the loss calculation method is expressed as:
[0040] Y = frs(X;θ)
[0041] L=Loss(Y,T)
[0042] Where Y represents the output value, T represents the target value, L represents the loss, frs() represents the network structure, and the output layer of the network structure uses the preset Softmax classifier to obtain the predicted label, and Loss(Y,T) represents the cross entropy loss;
[0043] The gradient of each parameter is calculated based on the loss function, and the calculation method is expressed as:
[0044]
[0045] In the formula represents the gradient of each parameter, represents the gradient with respect to the weight, represents the gradient with respect to the bias;
[0046] According to the direction and size of the gradient, a wave source is generated at the corresponding parameter position, and the intensity of each wave source is updated. The update method is expressed as:
[0047]
[0048] In the formula represents the intensity of the updated wave source, α ck represents the learning rate, β bp represents the attenuation coefficient, γ bp Indicates the amplification factor;
[0049] The fluctuations of all wave sources are superimposed in the parameter space to update each parameter. The parameter update method is expressed as:
[0050]
[0051] Where k represents the wave number, ω represents the angular frequency, Re() represents the ReLU activation function, and τ i,l represents the fluctuation threshold;
[0052] According to the loss reduction of this iteration, the attenuation coefficient and amplification coefficient of each wave source are dynamically adjusted. The adjustment method is expressed as:
[0053]
[0054] In the formula, Respectively represent the attenuation coefficient and amplification coefficient of the previous iteration, λ bp ,μ bp They represent the attenuation coefficient adjustment factor and the amplification coefficient adjustment factor respectively, and ΔL represents the change in loss;
[0055] Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
[0056] Preferably, when initializing the parameters of the neural network, the position, intensity and initial phase of the wave source are initialized and set. The method of initializing and setting the wave source is expressed as:
[0057] S i (x i ,a i ,φ i ),i=1,2,…,n
[0058] In the formula, x i Indicates the initial position of the parameter, a i represents the initial value of the randomly set strength, φ i represents the initial value of the randomly set initial phase, S i represents the i-th wave source, and n represents the total number of wave sources.
[0059] Preferably, the training process of the autoencoder neural network algorithm based on manifold mapping learning includes:
[0060] An autoencoder neural network model is constructed, wherein the structure of the autoencoder neural network model includes an encoder for mapping high-dimensional input data to a low-dimensional feature space and a decoder for reconstructing original input data from the low-dimensional feature space, and the network parameters of the encoder are initialized and set, wherein the network parameters of the encoder include weights and biases, and the initialization setting method is expressed as:
[0061]
[0062] b i =0
[0063] In the formula, w i represents the weight of the i-th layer, σ ch represents standard deviation;
[0064] During the encoding process, the input data is mapped to a low-dimensional feature space by the encoder, and during the decoding process, the original data is reconstructed by the decoder;
[0065] In the encoded low-dimensional feature space, Riemann factors are used for popular learning;
[0066] The Riemann factor based on popular learning is combined with the reconstruction error to calculate the loss function, and the calculation method of the loss function is expressed as:
[0067]
[0068] In the formula, λ cm represents the regularization coefficient, represents the reconstruction error;
[0069] Based on the loss function, the network parameters are adjusted through the back-propagation algorithm to optimize the performance of the encoder and decoder;
[0070] Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
[0071] Preferably, the training process of the meta-learning-based adaptive modulation extreme learning machine algorithm includes:
[0072] Build a basic extreme learning machine model and initialize the model parameters;
[0073] A subset of metadata is extracted from the training data, and in the meta-learning modulation stage, the model is self-adjusted according to the metadata to adapt to different data characteristics, and the adjustment method is expressed as:
[0074]
[0075] In the formula, qx m Represents metadata, represents the adjusted weight, represents the adjusted bias, η q represents the learning rate; L() represents the loss function calculated based on metadata, represents the gradient of the loss function with respect to the weight, Represents the gradient of the loss function with respect to the bias;
[0076] According to the suggestions of the meta-learning modulator, the weights and biases of the model are dynamically adjusted to optimize the responsiveness and accuracy of the model. The adjustment method is expressed as:
[0077]
[0078] In the formula, represents the final adjusted weight, represents the final adjusted bias, α q represents the adjustment coefficient, D q Represents a dynamically adjusted matrix;
[0079] During the training process, based on the dynamic hierarchical gradient sparsification strategy of Caiyong, the training efficiency and prediction accuracy of the model are improved by reducing the weights of unimportant connections. The sparsification method is expressed as:
[0080]
[0081] In the formula, S q Represents the weight after applying sparseness, and mask() represents the weight according to the weight The value and threshold θ q Function to generate mask;
[0082] Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
[0083] Preferably, when using the trained feature extraction model, feature dimension reduction model and classifier model to process the real-time construction site image to monitor the construction workers' work violations and identify the corresponding categories, the specific steps include:
[0084] The data of the current image is converted into a representative feature vector through a feature extraction model;
[0085] The feature vector is further processed by the trained feature dimensionality reduction model to reduce the dimensionality while retaining information useful for classification decisions;
[0086] The processed feature vector is input into the trained classifier model for monitoring and identification and output of the corresponding category.
[0087] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0088] 1) A combined weighted convolution layer is used in the generator to automatically adjust the weight of the convolution kernel, optimize feature extraction according to the characteristics of the input data, improve the diversity and realism of the generated images, and generate data in complex environments that are closer to the actual construction site, thereby increasing the diversity of the data and enhancing the generalization ability of the model.
[0089] 2) Simulate the fluctuation phenomenon and dynamically update the weights and biases in the network through the characteristics of periodic changes and energy transfer, optimize the common oscillation and local minimum problems in the traditional gradient descent method, and make the parameter update process smoother and more adaptive.
[0090] 3) Try to reconstruct the input through the decoding process, use the Riemann factor to optimize the representation of the feature space, strengthen the understanding of the intrinsic structure of the data, ensure that the data after dimensionality reduction can still retain key information, and improve the accuracy and efficiency of data reconstruction.
[0091] 4) Adjust model parameters through metadata supervision to cope with data diversity, reduce the risk of overfitting, improve the generalization ability of the model on different data sets, and enhance the stability and adaptability of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0093] Figure 1 It is a training flow chart of a neural network feature extraction algorithm based on adaptive fluctuation optimization in an embodiment of the present invention.
[0094] Figure 2 It is a training flow chart of the autoencoder neural network algorithm based on manifold mapping learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0095] The specific embodiments of the present invention are further described below in conjunction with the accompanying drawings. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0096] See also Figure 1 and Figure 2 As shown, an embodiment of the present invention provides a method for monitoring illegal behaviors of construction workers based on artificial intelligence, comprising the following steps:
[0097] 1) Data collection and annotation
[0098] Collect image data of the construction site and mark different types of work violations in the image data;
[0099] 2) Data expansion
[0100] The collected image data is expanded using a composite perturbation-based generative adversarial network algorithm to generate more diverse and realistic image data;
[0101] 3) Feature extraction model training
[0102] Constructing a feature extraction model, inputting the expanded image data into the feature extraction model and performing feature extraction training by using a neural network feature extraction algorithm based on adaptive fluctuation optimization to obtain image data after feature extraction;
[0103] 4) Feature dimensionality reduction model training
[0104] Construct a feature dimensionality reduction model, input the image data after feature extraction into the feature dimensionality reduction model, and perform feature dimensionality reduction training by using an autoencoder neural network based on manifold mapping learning to obtain the image data after feature dimensionality reduction;
[0105] 5) Classifier model training
[0106] A classifier model is constructed, the image data after feature dimension reduction is input into the classifier model, and classification training is performed by using an adaptive modulation extreme learning machine algorithm based on meta-learning to obtain a category result consistent with the labeled category of the work violation;
[0107] 6) Monitoring and identification of illegal behaviors
[0108] The trained feature extraction model, feature dimension reduction model and classifier model are used to process the images of the real-time construction site to monitor the work violations of construction workers and identify the corresponding categories.
[0109] Furthermore, in this embodiment, manual labeling is used to label different categories of work violations, and the labeled categories include: not wearing a safety helmet, not using a safety belt, and normal wearing. The image data is stored in JPEG format, and each pixel contains the values of the three RGB color channels.
[0110] Furthermore, the composite perturbation-based generative adversarial network algorithm includes the generator (bx G ) and the discriminator (bx D ), wherein the generator includes a combined weighted convolution layer. Preferably, when the collected image data is expanded using a composite perturbation-based generative adversarial network algorithm, the expansion process includes:
[0111] a. Initialize and set the network parameters of the generator and the discriminator. It should be noted that the network parameters in this embodiment include the weights and biases of the convolutional layer and the weight assignment coefficients in the combined weighted convolutional layer. The initialization method follows the normal distribution and can be expressed as:
[0112]
[0113] In the formula, θ G The network parameters representing the weights, θ D represents the biased network parameters, represents a normal distribution with a mean of 0 and a standard deviation of 0.02.
[0114] b. In the adversarial training stage, the generator receives random noise and outputs a synthetic image. The discriminator determines whether the input image data is real or synthetic, and calculates the loss function that minimizes the generator and the discriminator. The calculation method of the loss function is expressed as:
[0115]
[0116] Where bx Grepresents the generator, bx D represents the discriminator, bx~p data Indicates sampling from the real data distribution, bz~p z represents sampling from a random noise distribution, represents the expectation of the real image, represents the expectation of random noise.
[0117] c. In the generator, the performance of the generator is enhanced by adjusting the weights of the combined weighted convolution layer to reduce the difference between the generated image and the real image. The weight adjustment method of the combined weighted convolution layer is expressed as:
[0118]
[0119] In the formula, represents the weight at iteration t, η vk represents the learning rate, Represents the loss function About weight α bx The gradient of vk Set to 0.02.
[0120] It should be noted that, in this embodiment, the gradient of the loss function relative to the output of one of the convolutional layers of the combined weighted convolutional layer is assumed to be g out , the input of this convolutional layer is bx in , then the gradient of the weight The calculation method can be expressed as:
[0121]
[0122] Wherein, ⊙ represents the Hadamard product (i.e., the Hadamard product, which is a type of matrix operation).
[0123] d. Provide feedback through the discriminator after each training, and adjust the network parameters of the generator according to the feedback of the discriminator. The adjustment method is expressed as:
[0124]
[0125] In the formula, λ gs represents the adjustment factor, represents the loss function of the discriminator, represents the gradient of the loss function with respect to the generator network parameters, preferably, λ gs Set to 0.01.
[0126] e. Repeat the above steps until a preset stop iteration condition is met. It should be noted that the preset stop iteration condition in this embodiment is to reach a preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0127] See also Figure 1 As shown in the figure, the neural network feature extraction algorithm based on adaptive fluctuation optimization is inspired by the fluctuation phenomenon. The fluctuation presents the characteristics of periodic change and energy transfer, and can transmit and transform energy forms in different media. It iteratively updates the weights and biases in the network by simulating the propagation and energy distribution characteristics of the fluctuation. Each parameter update is affected not only by the current gradient, but also by the cumulative influence of the previous parameter update fluctuations, which is similar to the fluctuation influence caused by the superposition of wave sources. This makes the parameter update process smoother and more adaptive, and effectively avoids the oscillation and local minimum problems in the traditional gradient descent method.
[0128] Preferably, the training process of the neural network feature extraction algorithm based on adaptive fluctuation optimization includes:
[0129] a. Define the neural network structure and initialize the parameters of the neural network. The parameters of the neural network include weights and biases, and the initialization setting method is expressed as:
[0130] Y0={w0,b0}
[0131] In the formula, θ0 represents the initial parameters of the neural network, w0 represents the initial weight, and b0 represents the initial bias.
[0132] It should be noted that, in this embodiment, the number of hidden layers of the neural network is 2, the first hidden layer includes 100 neurons, the second hidden layer includes 50 neurons, the activation function is the ReLU activation function, and when initializing the parameters of the neural network, the position, intensity and initial phase of the wave source are initialized, and the method of initializing the wave source is expressed as:
[0133] S i (x i ,a i ,φ i ),i=1,2,…,n
[0134] In the formula, x i Indicates the initial position of the parameter, a i represents the initial value of the randomly set strength, φ i represents the initial value of the randomly set initial phase, S i represents the i-th wave source, and n represents the total number of wave sources.
[0135] b. Use the parameters to perform forward propagation on the input to get the output value, compare the output value with the target value to calculate the loss. The way to calculate the loss is expressed as:
[0136] Y = frs(X;θ)
[0137] L=Loss(Y,T)
[0138] In the formula, Y represents the output value, T represents the target value, L represents the loss, frs() represents the network structure, and the output layer of the network structure uses the preset Softmax classifier to obtain the predicted label, and Loss(Y,T) represents the cross entropy loss.
[0139] c. Calculate the gradient of each parameter based on the loss function. The calculation method is expressed as:
[0140]
[0141] In the formula represents the gradient of each parameter, represents the gradient with respect to the weight, In this embodiment, the gradient is calculated by using the chain method, where the gradient calculation method for the weight w is expressed as:
[0142]
[0143] In the formula, y k represents the output of Softmax, z k Represents the linear output before the Softmax layer.
[0144] Further, The calculation method can be expressed as:
[0145]
[0146] and, The calculation method can be expressed as:
[0147]
[0148] and, The calculation method can be expressed as:
[0149]
[0150] d. According to the direction and size of the gradient, a wave source is generated at the corresponding parameter position, and the intensity of each wave source is updated. The update method is expressed as:
[0151]
[0152] In the formula represents the intensity of the updated wave source, α ck represents the learning rate, β bp represents the attenuation coefficient, γ bp represents the amplification factor, preferably, α ck Set to 0.01.
[0153] e. Superimpose the fluctuations of all wave sources in the parameter space to update each parameter. The update amount of each parameter is determined by the superposition result of the fluctuations of all wave sources at this point. The parameter update method is expressed as:
[0154]
[0155] Where k represents the wave number, ω represents the angular frequency, Re() represents the ReLU activation function, and τ i,l Indicates the fluctuation threshold.
[0156] Furthermore, the fluctuation threshold τ of each wave source i,l The layer sensitivity index is adjusted dynamically, and the adjustment method can be expressed as:
[0157]
[0158] In the formula, τ0 represents the initial fluctuation threshold, S l Represents layer sensitivity index, κ pq represents the adjustment coefficient. Preferably, τ0 is set to 0.1, κ pq Set to 3.
[0159] Furthermore, the layer sensitivity index S of the lth layer l The norm calculation based on the gradient output of this layer is used to measure the contribution of this layer to the overall network error. The calculation method can be expressed as:
[0160]
[0161] In the formula, is the norm of the parameter gradient of the lth layer, Lc is the number of network layers
[0162] f. According to the loss reduction of this iteration, dynamically adjust the attenuation coefficient and amplification coefficient of each wave source. The adjustment method is expressed as:
[0163]
[0164] In the formula, are the attenuation coefficient and amplification coefficient of the previous iteration, λ bp ,μ bp are the attenuation coefficient adjustment factor and the amplification coefficient adjustment factor, respectively, and ΔL is the change in loss. Preferably, the parameter λ bp and μbp Set to 0.1 and 0.05 respectively.
[0165] g. Repeat the above steps until the preset stop iteration condition is met to complete the model training. In this embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0166] See also Figure 2 As shown, when the image data after feature extraction is input into the feature dimensionality reduction model and the image data is trained for feature dimensionality reduction by adopting the autoencoder neural network based on manifold mapping learning, the autoencoder neural network not only learns the intrinsic structure of the data in the encoding stage, but also attempts to reconstruct the input through the decoding process to ensure that the features after dimensionality reduction can still retain key information. The autoencoder neural network model structure includes an encoder for mapping high-dimensional input data to a low-dimensional feature space and a decoder for reconstructing the original input data from the low-dimensional feature space. By utilizing the Riemann factor based on popular learning to optimize the representation of the feature space, and by utilizing the geometric properties on the Riemann manifold, the local structure of the data can be better understood and more accurate data reconstruction can be achieved.
[0167] Preferably, the training process of the autoencoder neural network algorithm based on manifold mapping learning includes:
[0168] a. Construct an autoencoding neural network model and initialize the network parameters of the encoder, where the network parameters of the encoder include weights and biases. In this embodiment, Gaussian distribution is used to initialize the parameters of the encoder, and the initialization setting method is expressed as:
[0169]
[0170] b i =0
[0171] In the formula, w i represents the weight of the i-th layer, σ ch represents standard deviation;
[0172] b. In the encoding process, the input data is mapped to a low-dimensional feature space by the encoder, and in the decoding process, the original data is reconstructed by the decoder. Specifically, the way in which the data x is mapped to the low-dimensional feature space z by the autoencoder can be expressed as:
[0173] z=Sig(W e x+b e )
[0174] And, the decoder attempts to reconstruct the original input can be expressed as:
[0175]
[0176] Where Sig() is the activation function; W e and b e are the weight and bias of the encoder respectively; W d and b d are the decoder weights and biases.
[0177] Furthermore, in this embodiment, the weight W e By presetting it in the way of principal component analysis, the calculation method can be expressed as:
[0178]
[0179] Where V k is the covariance matrix X T The matrix consisting of the first k eigenvectors obtained from X, Σ k is the corresponding diagonal matrix of eigenvalues.
[0180] c. In the encoded low-dimensional feature space, Riemann factors are used for popular learning, where the way in which the Riemann factors constrain the encoded feature z can be expressed as:
[0181]
[0182] In the formula, μ cy represents the preset manifold center, σ R is a parameter that controls the compactness of the manifold. Preferably, μ cy Set to the mean vector of the low-dimensional space, σ R Set to 0.1.
[0183] d. The Riemann factor based on popular learning is combined with the reconstruction error to calculate the loss function, where the reconstruction error ensures that the data can be effectively restored, the popular constraint strengthens the geometric structure of the feature space, and the loss function is calculated as:
[0184]
[0185] In the formula, λ cm represents the regularization coefficient, represents the reconstruction error.
[0186] Furthermore, the reconstruction error The calculation method can be expressed as:
[0187]
[0188] In the formula, z i represents the value of z in the i-th dimension, μ i Represents the value of the preset manifold center in the i-th dimension.
[0189] e. Based on the loss function, adjust the network parameters through the back propagation algorithm to optimize the performance of the encoder and decoder. Preferably, in this embodiment, the weight W e For example, the calculation method of its gradient can be expressed as:
[0190]
[0191] Furthermore, the gradient It is calculated by the error back propagation method and can be expressed as:
[0192]
[0193] and, Through the decoder weight W d The transposed representation of can be expressed as:
[0194]
[0195] Furthermore, since z is W e The result obtained by the activation function after the linear combination of and x is The calculation method can be expressed as:
[0196]
[0197] Where Sig′() is the derivative of the Sigmoid activation function, diag(v) represents the diagonal matrix with vector v as the diagonal element, and x T is the transpose of x.
[0198] It should be noted that in this embodiment, the gradient descent method is used to update the parameters, and the updating method can be expressed as:
[0199] In the formula, α rv represents the learning rate, t represents the number of iterations, represents the updated weight parameter, represents the weight parameter before updating. Preferably, α rv Set to 0.05.
[0200] f. Repeat the above steps until the preset stop iteration condition is met and the model training is completed. In this embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations.
[0201]
[0202] Optionally, the preset maximum number of iterations is set to 1000.
[0203] The present invention adopts an adaptive modulation extreme learning machine algorithm based on meta-learning as a classifier. The adaptive modulation extreme learning machine algorithm based on meta-learning uses the basic extreme learning machine model to provide fast learning capabilities, and uses the meta-learning modulation mechanism to adjust and optimize model parameters through supervised metadata to reduce the risk of overfitting and improve the generalization ability of the algorithm on diversified data.
[0204] Among them, the training process of the adaptive modulation extreme learning machine algorithm based on meta-learning includes:
[0205] a. Construct a basic extreme learning machine model and initialize the parameters of the model. Preferably, a random strategy is used in this embodiment to initialize the parameters, and the initialization method is expressed as:
[0206] qx ω =σ q (Γ q )
[0207] qx n =β q (Γ q )
[0208] In the formula, qx ω Represents the initialization value of the network weight, qx b represents the initialization value of the bias, σ q and β q represents the initialization function, Γ q represents the parameter generation distribution used to generate weights and biases, preferably, a standard normal distribution is selected.
[0209] b. Extract a subset of metadata from the training data, and in the meta-learning modulation stage, adjust the model according to the metadata to adapt to different data characteristics, and the adjustment method is expressed as:
[0210]
[0211] In the formula, qx m Represents metadata, represents the adjusted weight, represents the adjusted bias, η q represents the learning rate; L() represents the loss function calculated based on metadata, represents the gradient of the loss function with respect to the weight, Represents the gradient of the loss function with respect to the bias.
[0212] Furthermore, the loss function L() is applied to the weight qx ω The gradient calculation method can be expressed as:
[0213]
[0214] In the formula, Represents the loss function relative to the model output qx out The gradient of The Jacobian matrix representing the model output with respect to the weights.
[0215] c. According to the suggestions of the meta-learning modulator, the weights and biases of the model are dynamically adjusted to optimize the responsiveness and accuracy of the model. Specifically, in this process, the adjustment of weights and biases is not a simple gradient descent, but includes the evaluation of the importance of weights of each layer and the corresponding adjustment to ensure that key features are strengthened. The adjustment of model parameters is further refined, which can be expressed as:
[0216]
[0217] In the formula, represents the final adjusted weight, represents the final adjusted bias, α q represents the adjustment coefficient, D q Represents a dynamically adjusted matrix.
[0218] Furthermore, the matrix D is dynamically adjusted q The calculation method can be expressed as:
[0219] D q =diag(softmax(∥qx ω ∥))
[0220] In the formula, ∥qx ω ∥The modulus of the weight vector is calculated, and the softmax() function is used to convert the modulus of the weight vector into a probability distribution.
[0221] d. During the training process, based on the dynamic hierarchical gradient sparsification strategy of Caiyong, the training efficiency and prediction accuracy of the model are improved by reducing the weights of unimportant connections. The sparsification method is expressed as:
[0222]
[0223] In the formula, S q Represents the weight after applying sparseness, and mask() represents the weight according to the weight The value and threshold θ q Function to generate the mask.
[0224] Preferably, in this embodiment, the mask function mask() is used to implement dynamic layered gradient sparsification, which can be expressed as:
[0225]
[0226] e. Repeat the above steps until the preset stop iteration condition is met to complete the model training. In this embodiment, the preset stop iteration condition is to reach a preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0227] When using the trained feature extraction model, feature dimension reduction model and classifier model to process the real-time construction site images to monitor the construction workers' work violations and identify the corresponding categories, the specific steps include:
[0228] The data of the current image is converted into a representative feature vector through a feature extraction model;
[0229] The feature vector is further processed by the trained feature dimensionality reduction model to reduce the dimensionality while retaining information useful for classification decisions;
[0230] The processed feature vector is input into the trained classifier model for monitoring and identification and output of corresponding categories. The identified categories include: not wearing a helmet, not using a seat belt, and normal wearing, a total of 3 categories.
[0231] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.
Claims
1. A method for monitoring construction workers' illegal behaviors based on artificial intelligence, characterized in that: The following steps are involved: Collect image data of the construction site and mark different types of work violations in the image data; The collected image data is expanded using a composite perturbation-based generative adversarial network algorithm; Constructing a feature extraction model, inputting the expanded image data into the feature extraction model and performing feature extraction training by using a neural network feature extraction algorithm based on adaptive fluctuation optimization to obtain image data after feature extraction; Construct a feature dimensionality reduction model, input the image data after feature extraction into the feature dimensionality reduction model, and perform feature dimensionality reduction training by using an autoencoder neural network based on manifold mapping learning to obtain the image data after feature dimensionality reduction; A classifier model is constructed, the image data after feature dimension reduction is input into the classifier model, and classification training is performed by using an adaptive modulation extreme learning machine algorithm based on meta-learning to obtain a category result consistent with the labeled category of the work violation; The trained feature extraction model, feature dimension reduction model and classifier model are used to process the images of the real-time construction site to monitor the work violations of construction workers and identify the corresponding categories.
2. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1 is characterized in that: Manual labeling is used to mark different types of work violations, and the marked categories include: not wearing a safety helmet, not using a safety belt, and normal wearing.
3. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1 is characterized in that: When a composite perturbation-based generative adversarial network algorithm is used to expand collected image data to generate more diverse and realistic image data, the composite perturbation-based generative adversarial network algorithm includes a generator and a discriminator, wherein the generator includes a combined weighted convolutional layer.
4. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 3 is characterized in that: When the collected image data is expanded using the composite perturbation-based generative adversarial network algorithm, the expansion process includes: Initialize and set the network parameters of the generator and discriminator; The generator receives random noise and outputs a synthetic image. The discriminator determines whether the input image data is real or synthetic, and calculates the loss function that minimizes the generator and the discriminator. The calculation method of the loss function is expressed as: Where bx G represents the generator, bx D represents the discriminator, bx~p data Indicates sampling from the real data distribution, bz~p z represents sampling from a random noise distribution, represents the expectation of the real image, represents the expectation of random noise; In the generator, the performance of the generator is enhanced by adjusting the weights of the combined weighted convolution layer to reduce the difference between the generated image and the real image. The weight adjustment method of the combined weighted convolution layer is expressed as: In the formula, represents the weight at iteration t, η vk represents the learning rate, Represents the loss function About weight α bx The gradient of The discriminator provides feedback after each training, and adjusts the network parameters of the generator according to the feedback of the discriminator. The adjustment method is expressed as: In the formula, λ gs represents the adjustment factor, represents the loss function of the discriminator, represents the gradient of the loss function with respect to the generator network parameters; Repeat the above steps until the preset stop iteration condition is met.
5. The method for monitoring illegal behaviors of construction workers based on artificial intelligence as claimed in claim 4, wherein when the performance of the generator is enhanced by adjusting the weights of the combined weighted convolutional layer in the generator, the loss function is set The gradient of the output of one of the convolutional layers relative to the combined weighted convolutional layer is g out , and the input of the convolutional layer is bx in , then the gradient of the weight The calculation method is expressed as: Where ⊙ represents the Hadamard product.
6. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1 is characterized in that: The training process of the neural network feature extraction algorithm based on adaptive fluctuation optimization includes: Define the neural network structure and initialize the parameters of the neural network. The parameters of the neural network include weights and biases, and the initialization setting method is expressed as: θ0={w0,b0} In the formula, θ0 represents the initial parameters of the neural network, w0 represents the initial weight, and b0 represents the initial bias; Use the parameters to perform forward propagation on the input to get the output value, compare the output value with the target value to calculate the loss, and the loss calculation method is expressed as: Y = frs(X;θ) L=Loss(Y,T) Where Y represents the output value, T represents the target value, L represents the loss, frs() represents the network structure, and the output layer of the network structure uses the preset Softmax classifier to obtain the predicted label, and Loss(Y,T) represents the cross entropy loss; The gradient of each parameter is calculated based on the loss function, and the calculation method is expressed as: In the formula represents the gradient of each parameter, represents the gradient with respect to the weight, represents the gradient with respect to the bias; According to the direction and size of the gradient, a wave source is generated at the corresponding parameter position, and the intensity of each wave source is updated. The update method is expressed as: In the formula represents the intensity of the updated wave source, α ck represents the learning rate, β bp represents the attenuation coefficient, γ bp Indicates the amplification factor; The fluctuations of all wave sources are superimposed in the parameter space to update each parameter. The parameter update method is expressed as: Where k represents the wave number, ω represents the angular frequency, Re() represents the ReLU activation function, and τ i,l represents the fluctuation threshold; According to the loss reduction of this iteration, the attenuation coefficient and amplification coefficient of each wave source are dynamically adjusted. The adjustment method is expressed as: In the formula, Respectively represent the attenuation coefficient and amplification coefficient of the previous iteration, λ bp ,μ bp They represent the attenuation coefficient adjustment factor and the amplification coefficient adjustment factor respectively, and ΔL represents the change in loss; Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
7. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 6 is characterized in that: When initializing the parameters of the neural network, the position, intensity and initial phase of the wave source are initialized. The method of initializing the wave source is expressed as: S i (x i ,a i ,f i ),i=1,2,…,n In the formula, x i Indicates the initial position of the parameter, a i represents the initial value of the randomly set strength, φ i represents the initial value of the randomly set initial phase, S i represents the i-th wave source, and n represents the total number of wave sources.
8. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1 is characterized in that: The training process of the autoencoder neural network algorithm based on manifold mapping learning includes: An autoencoder neural network model is constructed, wherein the structure of the autoencoder neural network model includes an encoder for mapping high-dimensional input data to a low-dimensional feature space and a decoder for reconstructing original input data from the low-dimensional feature space, and the network parameters of the encoder are initialized and set, wherein the network parameters of the encoder include weights and biases, and the initialization setting method is expressed as: b i =0 In the formula, w i represents the weight of the i-th layer, σ ch represents standard deviation; During the encoding process, the input data is mapped to a low-dimensional feature space by the encoder, and during the decoding process, the original data is reconstructed by the decoder; In the encoded low-dimensional feature space, Riemann factors are used for popular learning; The Riemann factor based on popular learning is combined with the reconstruction error to calculate the loss function, and the calculation method of the loss function is expressed as: In the formula, λ cm represents the regularization coefficient, represents the reconstruction error; Based on the loss function, the network parameters are adjusted through the back-propagation algorithm to optimize the performance of the encoder and decoder; Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
9. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1, characterized in that: The training process of the meta-learning-based adaptive modulation extreme learning machine algorithm includes: Build a basic extreme learning machine model and initialize the model parameters; A subset of metadata is extracted from the training data, and in the meta-learning modulation stage, the model is self-adjusted according to the metadata to adapt to different data characteristics, and the adjustment method is expressed as: In the formula, qx m Represents metadata, represents the adjusted weight, represents the adjusted bias, η q represents the learning rate; L() represents the loss function calculated based on metadata, represents the gradient of the loss function with respect to the weight, Represents the gradient of the loss function with respect to the bias; According to the suggestions of the meta-learning modulator, the weights and biases of the model are dynamically adjusted to optimize the responsiveness and accuracy of the model. The adjustment method is expressed as: In the formula, represents the final adjusted weight, represents the final adjusted bias, α q represents the adjustment coefficient, D q Represents a dynamically adjusted matrix; During the training process, based on the dynamic hierarchical gradient sparsification strategy of Caiyong, the training efficiency and prediction accuracy of the model are improved by reducing the weights of unimportant connections. The sparsification method is expressed as: In the formula, S q Represents the weight after applying sparseness, and mask() represents the weight according to the weight The value and threshold θ q Function to generate mask; Repeat the above steps until the preset stop iteration condition is met and the model training is completed.
10. The method for monitoring construction workers' illegal behaviors based on artificial intelligence as claimed in claim 1, characterized in that: When using the trained feature extraction model, feature dimension reduction model and classifier model to process the real-time construction site images to monitor the construction workers' work violations and identify the corresponding categories, the specific steps include: The data of the current image is converted into a representative feature vector through a feature extraction model; The feature vector is further processed by the trained feature dimensionality reduction model to reduce the dimensionality while retaining information useful for classification decisions; The processed feature vector is input into the trained classifier model for monitoring and identification and output of the corresponding category.
Citation Information
Patent Citations
A method and system for identifying helmet wearing
CN111401276B
Construction worker identity recognition and safety helmet wearing detection method and system
CN111598040A
Safety early warning method based on intelligent safety helmet and medium
CN117994930A