A method and system for identifying prefabricated building components based on computer vision

By collecting image data of construction scenes at construction sites, building a model training set, and dynamically adjusting the loss function and model parameters using non-European geometric distance metrics, the problem of low component recognition accuracy in complex environments on construction sites is solved, and the recognition effect of high precision and rapid convergence is achieved.

CN119992346BActive Publication Date: 2025-06-24CHINA STATE CONSTR HAILONG TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510468078.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art is difficult to achieve stable component recognition in complex environments of construction sites, which can easily lead to problems such as instability in training, degradation of recognition accuracy and insufficient generalization capabilities.

Method used

The pre-deployed image acquisition device collects construction scenes at construction sites, builds a model training set, and dynamically adjusts the loss function using non-European geometric distance metrics to dynamically adjust the model parameters to improve recognition accuracy.

Benefits of technology

In complex scenarios such as lighting changes, occlusion and local similarity, high recognition accuracy is achieved, and the training process converges faster by dynamically adjusting the model parameters to ensure higher accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992346B_ABST
    Figure CN119992346B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image data processing, and particularly relates to a method and system for identifying prefabricated building components based on computer vision, including: collecting construction scenes at different time periods on a construction site to construct a model training set; training a pre-constructed component recognition model according to each training image in the model training set and the corresponding component type label of each training image; during the training process, every time the component recognition model obtains a corresponding prediction output based on the training image, the loss function and model parameters are dynamically adjusted in real time based on the non-Euclidean geometric distance metric between the training image and the prediction output; when the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result. The beneficial effect is that through the dynamic optimization of the loss function and model parameters, the training process converges faster while ensuring the recognition accuracy of the component recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular, to a method and system for identifying prefabricated building components based on computer vision. Background Art

[0002] With the development of building industrialization and intelligence, prefabricated buildings have been widely promoted globally due to their advantages such as high construction efficiency, controllable quality, and resource conservation. However, due to a large number of standardized and non-standardized components in prefabricated building designs, such as steel structures, precast concrete slabs, connectors, etc., how to efficiently and accurately identify and manage these components has become a key issue in building construction and quality control.

[0003] Currently, the identification of prefabricated components mainly relies on manual inspection and management, that is, construction workers identify and manage components by visual judgment or based on manual records. However, this method has many deficiencies, including low efficiency, large errors, high costs, and in large construction sites, the manual identification method is difficult to meet the requirements of precise positioning, quality tracking, construction progress management, etc. Therefore, using computer vision technology to achieve automated and intelligent component identification is of great significance for improving the construction management level of prefabricated buildings.

[0004] However, the application of existing computer vision recognition technology in building component identification still faces many challenges. First, the environment of construction sites is complex, and factors such as light changes, noise interference, and perspective distortion will affect the image quality, making it difficult for existing methods to extract stable features. Second, there are various types of prefabricated components, and the shapes of some components are similar, such as different types of steel components or precast concrete slabs. Traditional classification methods based on convolutional neural networks are prone to misclassification when distinguishing these similar components. In addition, components may be partially occluded, missing, or damaged during the construction process, further increasing the difficulty of identification. The traditional method of using fixed momentum coefficient, fixed learning rate, and simple cross-entropy loss in the training process of convolutional neural networks is difficult to cope with the complex construction site environment, easily leading to problems such as unstable training, decreased recognition accuracy, and insufficient generalization ability. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method and system for identifying prefabricated building components based on computer vision, which solves the technical problems of low efficiency, large errors, and high costs of the manual detection method, and is difficult to meet the requirements of precise positioning, quality tracking, construction progress management, etc. in large construction sites, as well as the technical problems that existing computer vision technology is difficult to cope with the complex construction site environment, easily leading to unstable training, decreased recognition accuracy, and insufficient generalization ability.

[0007] (2) Technical solution

[0008] To achieve the above object, the main technical solutions adopted by the present invention include:

[0009] In a first aspect, an embodiment of the present invention provides a method for identifying prefabricated building components based on computer vision, including:

[0010] S11. Collect construction site construction scenes at different times through a pre-deployed image acquisition device to construct a model training set;

[0011] The model training set includes a preset number of training images, and component type labels corresponding to the prefabricated components in each training image;

[0012] Each of the training images includes prefabricated components in the construction scene;

[0013] S12. Train a pre-constructed component recognition model according to each training image in the model training set and the component type label corresponding to each training image;

[0014] During the training process, every time the component recognition model obtains a corresponding prediction output according to the training image, the loss function is dynamically adjusted based on the non-Euclidean geometric distance metric between the training image and the prediction output, and the model parameters are dynamically adjusted according to the dynamically adjusted loss function;

[0015] When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized, and a component recognition result is obtained.

[0016] Optionally, the S12 includes:

[0017] Perform multi-dimensional enhancement processing on all training images in the model training set according to a pre-set feature enhancement algorithm;

[0018] Input the model training set after multi-dimensional feature enhancement into the component recognition model for model training. During the training process, every time the component recognition model obtains a corresponding prediction output according to the training image, a corresponding cross-entropy loss value is obtained according to the training image and the prediction output, and a pre-set cross-entropy loss function;

[0019] And, dynamically adjust the corresponding cross-entropy loss value according to the non-Euclidean geometric distance metric between the training image and the prediction output to obtain a corresponding true loss value;

[0020] And, dynamically adjust the model parameters according to the true loss value;

[0021] When the end condition is reached, the trained component recognition model is obtained to detect the component image to be recognized, and the component recognition result is obtained.

[0022] Optionally, in S12, according to the non-Euclidean geometric distance metric between the training image and the prediction output, the corresponding cross-entropy loss value is dynamically adjusted to obtain the corresponding true loss value, including:

[0023] Based on the non-Euclidean geometric distance metric between the training image and the prediction output, and a preset formula 1, the corresponding cross-entropy loss value is dynamically adjusted; the formula 1 is:

[0024] ;

[0025] where L is the true loss value, L cross is the cross-entropy loss value, λ is the regularization coefficient, is the non-Euclidean geometric distance metric corresponding to the i-th training image, y i is the component category label of the i-th training image, x i is the i-th training image, is the prediction output of the component recognition model for the i-th training image, n is the number of training images input to the component recognition model, is the noise factor;

[0026] The expression of the non-Euclidean geometric distance metric is:

[0027] ;

[0028] where f s (•) is the high-dimensional depth feature representation extracted from the training image by the component recognition model, is the depth feature class center of the component category to which the training image belongs, is the depth feature class variance of the component category to which the training image belongs, is the nearest neighbor of the training image in the component category to which it belongs, G m is the m-th Gaussian scale kernel, M is the number of scales, is the convolution operation, ||•|| F is the Frobenius norm, ||•||2 is the L2 norm, is the weight coefficient of the m-th scale, and are both extremely small positive numbers.

[0029] Optionally, the model parameters include a momentum coefficient;

[0030] Then in S12, according to the true loss value, the model parameters are dynamically adjusted, including:

[0031] According to the true loss value of the current training step, obtain the gradient of the true loss value of the current training step with respect to the weights, and save the gradient of the true loss value of the current training step with respect to the weights;

[0032] According to the gradient of the true loss value of the current training step with respect to the weights, and the gradients of the true loss values of all previous training steps with respect to the weights pre-saved, obtain the gradient change factor corresponding to the current training step;

[0033] According to the gradient change factor corresponding to the current training step, and the preset initial momentum coefficient and momentum adjustment parameter, dynamically adjust the momentum coefficient corresponding to the current training step.

[0034] Optionally, the model parameters further include convolutional kernel weights;

[0035] Then, S12, when dynamically adjusting the model parameters according to the true loss value, further includes:

[0036] According to the true loss value of the current training step, obtain the second derivative of the true loss value of the current training step with respect to the weights;

[0037] According to the second derivative of the true loss value of the current training step with respect to the weights and the gradient of the true loss value of the current training step with respect to the weights, obtain the second derivative correction term corresponding to the current training step;

[0038] According to the mean value of information entropy and the gradient variance at each position in the feature map output by the model convolutional layer when the training image of the current training step is input, obtain the training complexity coefficient corresponding to the current training step;

[0039] According to the gradient of the true loss value with respect to the weights corresponding to the previous training step, the gradient of the true loss value with respect to the weights corresponding to the current training step, the second derivative correction term corresponding to the current training step, the complexity coefficient corresponding to the current training step, the convolutional kernel weights corresponding to the current training step, and the preset learning rate, obtain the convolutional kernel weights corresponding to the next training step.

[0040] Optionally, the model parameters further include the learning rate;

[0041] S12, when dynamically adjusting the model parameters according to the true loss value, further includes:

[0042] In the current training step, before updating and correcting the convolutional kernel weights of the component recognition model, dynamically adjust the learning rate according to the number of features corresponding to the training image in the current training step; the expression of the learning rate is:

[0043] ;

[0044] Wherein, is the initial learning rate, is the learning rate adjustment factor, t is the current training step, int(t) is the index value of the first t training steps, T is the maximum number of training times, n fea is the number of features in the training image, is the gradient of the k-th feature corresponding to the true loss value at the t-th training step, is the gradient of the o-th feature corresponding to the true loss value at the t-th training step, is the adjustment factor of the k-th feature, and the expression of the adjustment factor of the feature is:

[0045] ;

[0046] Among them, is the variance of the gradient of the k-th feature corresponding to the true loss value in the past h training steps of the component recognition model, is the mean of the gradient of the k-th feature corresponding to the true loss value in the past h training steps of the component recognition model, is the sensitivity coefficient, is a very small positive number.

[0047] Optionally, in S12, according to a pre-set feature enhancement algorithm, multi-dimensional enhancement processing is performed on all training images received in the model training set, including:

[0048] Performing stretching / scaling processing on the received training images according to a pre-set local adaptive contrast enhancement algorithm;

[0049] According to a pre-set multi-dimensional data enhancement algorithm, performing random rotation, random mirroring, random cropping, and affine transformation processing on the training images after stretching / scaling processing.

[0050] Optionally, S12 further includes:

[0051] Before inputting the model training set into the component recognition model for model training, initializing the weights and biases of the convolutional kernels of the component recognition model so that the weights and biases of the convolutional kernels are initialized to follow a normal distribution with a mean of 0 and a variance of 0.01.

[0052] Optionally, after S12, it further includes:

[0053] S13. Collecting the original image of the target area through a pre-set image acquisition device, where the original image includes the component to be recognized;

[0054] Inputting the original image into a pre-set component recognition model to obtain the corresponding component recognition result;

[0055] The component recognition model includes a recognition model obtained by training a convolutional neural network model with a pre-set model training set;

[0056] The model training set includes at least one training image; any one of the training images includes at least one component and a pre-annotated component category label for each component;

[0057] The component recognition model is used to extract features from the received original image through a pre-set convolutional layer to obtain a corresponding feature image;

[0058] and perform a non-linear mapping on the feature image through a pre-set non-linear activation function ReLU;

[0059] and perform pooling processing on the feature image after non-linear mapping through a pre-set pooling layer, and input the feature image output by the pooling layer into a pre-set fully connected layer to output the probability distribution of the component categories in the original image, and the maximum probability is the corresponding component recognition result.

[0060] In a second aspect, an embodiment of the present invention provides a building prefabricated component recognition system based on computer vision, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the above-mentioned building prefabricated component recognition method based on computer vision.

[0061] (III) Beneficial effects

[0062] The beneficial effect of the present invention is that for the building prefabricated component recognition method based on computer vision of the present invention, the loss function is dynamically optimized through non-Euclidean geometric distance measurement, so that the model can still maintain high recognition accuracy in complex scenarios such as light changes, occlusion, and local similarity. The dynamic adjustment of the model parameters enables the training process to converge faster while ensuring higher accuracy. Description of the drawings

[0063] Figure 1 It is a flowchart of a building prefabricated component recognition method provided by Embodiment 1 of the present invention;

[0064] Figure 2 It is a schematic flowchart of a building prefabricated component recognition method provided by Embodiment 2 of the present invention. Detailed implementation manners

[0065] In order to better explain the present invention for easy understanding, the present invention will be described in detail below with reference to the drawings through specific implementation manners.

[0066] An object recognition method for building prefabricated components based on computer vision proposed in an embodiment of the present invention dynamically optimizes the loss function through non-Euclidean geometric distance measurement, enabling the model to maintain high recognition accuracy in complex scenarios such as light changes, occlusion, and local similarity. The dynamic adjustment of the model parameters enables the training process to converge faster while ensuring higher accuracy.

[0067] To better understand the above technical solution, the exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0068] Embodiment 1

[0069] This embodiment provides an object recognition method for building prefabricated components based on computer vision, as Figure 1 shown, including:

[0070] S11. Collect construction scenes at different times on a construction site through a pre-deployed image acquisition device to construct a model training set;

[0071] The model training set includes a preset number of training images and the component type labels corresponding to the prefabricated components in each training image;

[0072] Each of the training images includes prefabricated components in the construction scene;

[0073] S12. Train a pre-constructed component recognition model according to each training image in the model training set and the component type label corresponding to each training image;

[0074] During the training process, every time the component recognition model obtains a corresponding prediction output according to the training image, the loss function is dynamically adjusted in real time based on the non-Euclidean geometric distance measurement between the training image and the prediction output, and the model parameters are dynamically adjusted in real time according to the dynamically adjusted loss function;

[0075] When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

[0076] An object recognition method for building prefabricated components based on computer vision provided in this embodiment dynamically optimizes the loss function through non-Euclidean geometric distance measurement, enabling the model to maintain high recognition accuracy in complex scenarios such as light changes, occlusion, and local similarity. The dynamic adjustment of the model parameters enables the training process to converge faster while ensuring higher accuracy.

[0077] Example 2

[0078] This embodiment provides a method for identifying prefabricated building components based on computer vision. As Figure 2 shown, this method is used for identifying prefabricated building components, and the method includes:

[0079] S1. The acquisition source, method, and storage format of training data;

[0080] The source of the model training set in this embodiment is the actual construction scene of a construction site, including but not limited to various prefabricated components such as steel structures and precast concrete slabs;

[0081] The acquisition method uses a high-resolution digital camera system (image acquisition device), which is fixed at different angles to cover the full range of views of the components, ensuring that the collected training image data can comprehensively reflect the component characteristics and environmental information;

[0082] The data storage format of the training images is JPEG format, and the resolution of each image is not less than 1920×1080 pixels.

[0083] S2. Data preprocessing;

[0084] Due to the complex environment of the construction site, the training images may be affected by factors such as light changes, noise interference, and perspective distortion. Gaussian filtering and bilateral filtering are used to remove random noise and improve the image quality.

[0085] S3. Data annotation;

[0086] Data annotation is performed on the collected training images, and the annotation results are in the form of component type labels of the prefabricated components in the training images; the annotation type labels include:

[0087] Type 1: Steel structure components (such as H-shaped steel beams, steel columns, steel connection plates);

[0088] Type 2: Precast concrete components (such as precast floor slabs, precast wall panels, precast beams, precast columns);

[0089] Type 3: Connectors (such as bolts, welds, embedded parts);

[0090] Type 4: Auxiliary equipment (support frames, construction tools);

[0091] ……

[0092] Type n: …….

[0093] A model training set is constructed with the training images and the corresponding component type labels of the training images.

[0094] S4. Training of the building prefabricated component recognition model;

[0095] S401. Perform multi-dimensional enhancement processing on the training images in the model training set;

[0096] For the training images, uneven illumination, noise, and partial missing of components are likely to occur during the actual acquisition process. These image characteristics complicate the subsequent feature extraction and classification steps.

[0097] In the prior art, usually only simple normalization processing is performed, which cannot fully solve the interference caused by illumination, noise, and geometric deformation, and easily leads to unstable feature extraction effects.

[0098] In this embodiment, by performing multi-dimensional enhancement on the images, it is ensured that the subsequent convolution operations of the convolutional neural network are more adaptable when extracting the key regions and texture features of building components; that is, according to the pre-set feature enhancement algorithm, multi-dimensional enhancement processing is performed on all the training images received in the model training set; the feature enhancement algorithm is expressed as:

[0099] ;

[0100] Among them, is the image after multi-dimensional feature enhancement processing, representing the result after processing in terms of illumination, noise, and geometric transformation, etc.; is the received training image, representing the component image collected from the construction site or prefabrication scene; is the normalization function, representing stretching or scaling the image grayscale value or color range; is the data enhancement function, representing operations such as randomly rotating and mirroring the image;

[0101] The normalization function stretches or scales the grayscale or color distribution of the training image through local adaptive contrast enhancement, effectively reducing the influence of uneven illumination and suppressing image noise at the same time; the calculation method of the local adaptive contrast enhancement algorithm is expressed as:

[0102] ;

[0103] Among them, I(a, b) is the pixel value of the training image at the position (a, b), is the local mean of the pixels within the window of size c×c centered at the position (a, b), is the local standard deviation of the pixels within the window of size c×c centered at the position (a, b), is a very small positive number to prevent division by zero; (a, b) is the coordinate value, and c is the window size. Preferably, is 0.0001, and c is set to 3.

[0104] The data augmentation function combines random rotation, mirroring, random cropping, and affine transformation to comprehensively achieve multi-dimensional data augmentation, enhancing the network's robustness to geometric deformations. The calculation method of the multi-dimensional data augmentation algorithm is expressed as:

[0105] ;

[0106] where A aff (•) is the random affine transformation operation, C crop (•) is the random cropping operation, M flip (•) is the random mirroring operation, is the random rotation operation, is the rotation angle.

[0107] S402. Initialize the parameters of the convolutional neural network.

[0108] By selecting the normal distribution to initialize the weights and biases of the convolutional kernels of the component recognition model (usually a convolutional neural network model), the problem of gradient explosion or disappearance during the early training of building component images (training images) is avoided, ensuring stable subsequent gradients. Specifically, the initialization of the weights and biases of the convolutional kernels follows a normal distribution with a mean of 0 and a variance of 0.01.

[0109] S403. Perform forward propagation calculation of the component recognition model (convolutional neural network model).

[0110] By inputting the training images processed by multi-dimensional augmentation into the convolutional neural network (component recognition model), first perform convolutional layer operations to extract local features of the images; then perform non-linear mapping on the features through the non-linear activation function ReLU; then reduce the size of the feature map through the pooling layer to reduce network parameters; finally, expand the feature map output by the pooling layer into a one-dimensional vector and send it to the fully connected layer to output the probability distribution corresponding to the training images, and select the one with the maximum probability as the corresponding predicted output according to the maximum value principle, which is the component type recognition result of the corresponding training image.

[0111] S404. Calculate the momentum coefficient of the convolutional neural network (component recognition model).

[0112] During the training process of the convolutional neural network in the model training set, due to the complex details, many blocks, and large local differences of the components.

[0113] Conventional convolutional neural networks are prone to being insensitive to the gradient fluctuations caused by local details of building components, resulting in the loss of some detailed information of building components.

[0114] In this embodiment, by setting a dynamic momentum coefficient, adaptive adjustment is performed according to the gradient change at each training step, and a gradient change factor is used to quantify the gradient accumulation difference, so as to find a balance between high-speed convergence and local stability, overcoming the difficulty that the conventional technology is insensitive to local gradient changes. The calculation method of the momentum coefficient is expressed as:

[0115] ;

[0116] where u t is the momentum coefficient at the t-th training step, representing the accumulation of the previous step gradient during update, u0 is the initial momentum coefficient, representing the starting setting of the momentum; α is the momentum adjustment parameter, representing the amplitude of the dynamic change of the momentum with training; β t is the gradient change factor at the t-th training step, representing the influence of the gradient difference between this training step and the historical training steps; t is a positive integer. Preferably, u0 is set to 0.1 and α is set to 2.

[0117] The gradient change factor dynamically adjusts the momentum coefficient based on the gradient change of the current training step. The expression of the gradient change factor is:

[0118] ;

[0119] where γ is the gradient adjustment constant, representing the amplification or reduction of the gradient difference accumulation; is the gradient of the loss function with respect to the weight at the j-th training step, int(t - 1) is the index of the previous t - 1 training steps, is the gradient of the loss function with respect to the weight at the j - 1-th training step. i is a positive integer. Preferably, γ is set to 0.5.

[0120] S405. Update the weight parameters of the convolutional neural network (component recognition model).

[0121] Still in the training of the convolutional neural network for building prefabricated component images, due to the complex details, many blocks and large local differences of the components. Conventional convolutional neural networks usually do not consider the historical changes of the gradient, or only use the simple accumulation of the current and historical gradients, and often it is difficult to balance local accuracy and global convergence speed when facing high-dimensional features and local significant differences.

[0122] This embodiment combines the momentum coefficient with second-order information, and uses the correction of the second-order information for weight update to realize the correction of the second derivative of the loss during the training process of the convolutional neural network, so as to better process the component features with large local differences. That is, according to the dynamically adjusted momentum coefficient and second-order information, update and correct the convolutional kernel weights of the component recognition model; the second-order information is the second derivative information of the loss function with respect to the model parameters; the expression for updating and correcting the convolutional kernel weights of the component recognition model is:

[0123] ;

[0124] wherein, W t+1 is the convolutional kernel weight at the (t + 1)-th training step, and W t is the convolutional kernel weight at the t-th training step, is the learning rate at the t-th iteration, is the training complexity coefficient, is the second derivative correction term, obtained through second-order information, is the gradient of the weight corresponding to the loss function at the t-th training step; is the gradient of the weight corresponding to the loss function at the (t - 1)-th training step, and the expression of the second derivative correction term is:

[0125] ;

[0126] wherein, is the second derivative of the weight corresponding to the loss function at the t-th training step.

[0127] The determination of the training complexity coefficient is based on the local feature complexity and occlusion degree shown by the current training image in the convolutional neural network.

[0128] Specifically, by analyzing the feature map after the input image passes through the convolutional layer, calculate the local information entropy and the variance of the gradient amplitude of each feature channel; by statistically calculating the mean information entropy and the gradient variance of each position of the feature map, and normalizing and weighting them, obtain ; wherein, E t is the mean information entropy of each position in the feature map at the t-th training step, and V t is the gradient variance of each position in the feature map at the t-th training step, is a very small positive number.

[0129] The training complexity coefficient dynamically adjusts the strength of the second derivative correction term, so as to enhance the second-order correction on complex training images (such as local missing regions), thereby improving the robustness of the model to occlusion.

[0130] S406. Calculate the classification result (predicted output).

[0131] Based on step S403, after operations such as convolution, planning, pooling, and fully connected layer, obtain the predicted probability distribution corresponding to each category, and then determine the classification result through the principle of maximizing the probability value, that is, the category with the highest probability is used as the final predicted output.

[0132] S407. Calculate the loss function of the convolutional neural network.

[0133] In the images of prefabricated building components, samples that are extremely similar locally can make it difficult for the network to distinguish during classification. It is necessary to make the loss function pay more attention to difficult-to-separate samples to improve the overall classification accuracy. Conventional convolutional neural networks usually use the cross-entropy loss function, which only cares about the global error and does not distinguish between local or difficult-to-separate samples, resulting in insufficient attention allocation in the network.

[0134] In this embodiment, a non-Euclidean geometric interval metric correction term is combined in the cross-entropy loss, so that difficult-to-separate samples obtain a higher penalty weight in the formula. At the same time, a noise factor is used to suppress the environmental noise during the acquisition process, thus taking into account both complex features and noise interference. That is, during the training process, every time the component recognition model obtains the corresponding prediction output according to the training image, the loss function is adjusted in real time and dynamically based on the non-Euclidean geometric distance metric between the training image and the prediction output, and the model parameters are adjusted in real time and dynamically according to the loss function after the dynamic adjustment; the expression (Formula 1) of the loss function is:

[0135] ;

[0136] where L is the loss function, representing the total loss of the model; L cross is the cross-entropy function; λ is the regularization coefficient, representing the weight of the non-Euclidean geometric metric; is the non-Euclidean geometric distance metric of the i-th training image; y i is the component category label of the i-th training image, and x i is the i-th training image, representing the i-th training image input into the model; is the prediction output of the component recognition model for the i-th training image, n is the number of training images input into the component recognition model, is the noise factor, representing the suppression strength for noisy samples or incomplete images. Preferably, λ is set to 0.3, is set to 0.01.

[0137] The non-Euclidean geometric distance metric is realized by combining the multi-scale local feature differences of the component images and the semantic distance between categories, and the expression of the non-Euclidean geometric distance metric is:

[0138] ;

[0139] where f s (•) is the high-dimensional deep feature representation extracted from the training image by the component recognition model, is the deep feature class center of the component category to which the training image belongs, calculated by the mean of the features of the same category samples; is the depth feature class variance of the component category to which the training image belongs, reflecting the degree of dispersion of features within the category; is the nearest neighbor of the training image in the component category to which it belongs (measured by the K-nearest neighbor algorithm), representing the boundary fuzzy points between categories; G m is the m-th Gaussian scale kernel, used to extract the differences in image features at different scales; M is the number of scales, preferably set to 3, capturing the differences in component features at different granularities from small scales to large scales; is the convolution operation, ||•|| F is the Frobenius norm, representing the scale sensitivity of the feature difference matrix; ||•||2 is the L2 norm, is the weight coefficient of the m-th scale, used to dynamically balance the contribution degrees of differences at different scales, preset by humans. For example, when M is set to 3, the weight coefficient of the first scale is set to 0.1, the weight coefficient of the second scale is set to 0.5, and the weight coefficient of the third scale is set to 0.4; and are both extremely small positive numbers, used to prevent division by zero. Preferably, and are both set to 0.0001.

[0140] That is, through the expression of the above loss function, in each training step, the predicted output corresponding to each training step and the true loss value corresponding to the training image can be obtained to dynamically adjust the model parameters.

[0141] S408. Dynamically adjust the learning rate of the convolutional neural network (component recognition model).

[0142] In the training of the convolutional neural network for prefabricated building component images, due to the complex details, numerous blocks, and large local differences of components, as the depth of the convolutional neural network and the data scale increase, if the learning rate is not dynamically adjusted, it may converge too slowly in the initial stage or oscillate too much in the later stage. Conventional methods of fixing the learning rate or simple piecewise decay cannot adaptively adjust according to the differences in the importance of different features in component images, resulting in a decline in training stability.

[0143] In this embodiment, by dynamically monitoring the importance of the proportion of the magnitudes of each feature gradient, the learning rate of the network is adaptively adjusted, enabling the network to converge quickly in the initial stage and focus on adjusting the key gradients, while steadily reducing the update amplitude in the later stage. Specifically, dynamic trade-offs are implemented through the contribution degree of the feature gradient to the learning rate to achieve efficient and stable convergence. That is, in the training step, before updating and correcting the convolutional kernel weights of the component recognition model, the learning rate is dynamically adjusted according to the number of features corresponding to the training image in this training step; the expression of the learning rate is:

[0144] ;

[0145] Among them, is the learning rate at the t-th training step, representing the gradient update amplitude of the current update iteration; is the initial learning rate; is the learning rate adjustment factor, representing the decay rate of the learning rate as the number of training steps increases; int(t) is the index value of the first t training steps, T is the maximum number of training times, and n fea is the number of features in the training image, is the gradient of the k-th feature corresponding to the loss function at the t-th training step, is the gradient of the o-th feature corresponding to the loss function at the t-th training step, is the adjustment factor of the k-th feature, representing the influence of the gradient corresponding to this feature on the learning rate; preferably, is set to 0.01, is set to 0.3.

[0146] The adjustment factor of the feature is determined by measuring the contribution degree of the feature to the network prediction accuracy at different training stages. The expression of the adjustment factor of the feature is:

[0147] ;

[0148] Among them, is the variance of the gradient of the k-th feature corresponding to the loss function in the past h training steps of the component recognition model, representing the strength of the gradient fluctuation; is the mean value of the gradient of the k-th feature corresponding to the loss function in the past h training steps of the component recognition model; is the sensitivity coefficient, used to adjust the influence intensity of the fluctuating gradient on the learning rate adjustment; is a very small positive number to prevent division by zero.

[0149] S409. Judgment on the end of the training of the convolutional neural network (component recognition model);

[0150] When the difference between the loss function values of the current training step and the previous training step is less than the preset threshold, the iteration is stopped, indicating that the convolutional neural network reaches the convergence state and the training of the convolutional neural network model is completed.

[0151] S5. Identification of building prefabricated components;

[0152] After the model training is completed, the trained convolutional neural network is deployed to the actual application scenario of the construction site for identification through real-time acquisition or pre-stored component images.

[0153] First, the input image needs to go through a preprocessing process consistent with the training stage, including Gaussian filtering, bilateral filtering for denoising, and local adaptive contrast enhancement, to eliminate environmental interference and improve image quality;

[0154] Further, the enhanced image is input into the trained convolutional neural network, and multi-scale depth features are extracted through multiple convolutional, activation, and pooling operations, and the probability distribution of component categories is generated using the fully connected layer;

[0155] Further, the model determines the category to which the target component belongs (such as steel structure, precast concrete, connection piece, or auxiliary facility) according to the maximum probability, and dynamically optimizes the classification boundary in combination with the non-Euclidean geometric distance metric correction term to ensure accurate discrimination of locally similar or occluded samples;

[0156] Further, the recognition result is output through a visualization interface, annotating the component category and confidence level, and synchronizing to the construction management platform to assist on-site personnel in component positioning, quality inspection, and progress tracking.

[0157] For example: At the construction site of a large-scale prefabricated building, the computer vision-based method for identifying prefabricated building components proposed in the present invention is adopted. High-resolution cameras are used to collect images of steel structures, precast concrete components, connection pieces, etc. at different construction stages, and real-time recognition and classification are carried out through the trained convolutional neural network. First, the collected images are denoised by Gaussian filtering and the local adaptive contrast is enhanced to reduce environmental interference. Then, through multiple convolutional, pooling, activation, etc. processes, multi-scale features of the components are extracted, and combined with the classification algorithm optimized by non-Euclidean geometry metrics, accurate recognition of different types of components is realized, and the component category, confidence level, and position information are displayed in real time on the visualization interface. Finally, the recognition system is linked with the construction management platform to realize automatic classification of components, progress tracking, and quality inspection, reducing manual operations and improving the intelligent level of construction management.

[0158] The method for identifying prefabricated building components based on computer vision provided in this embodiment, through multi-dimensional data augmentation and non-Euclidean geometric loss function optimization, enables the model to still maintain high recognition accuracy in complex scenarios such as light changes, occlusion, and local similarity. The dynamic momentum coefficient and second-order information optimize the weight update to effectively balance local feature extraction and global convergence speed, making the training process more stable and reducing the possibility of training failure. The adoption of dynamic learning rate adjustment and non-Euclidean geometric loss optimization enables the model to better distinguish different types of components, and can accurately recognize even in the case of similar structures or partial missing. By adaptively adjusting the momentum coefficient, the training process converges faster while maintaining a high classification accuracy. Through the noise suppression strategy and feature gradient adjustment of the learning rate, the network can still accurately classify in high-noise and complex backgrounds.

[0159] Furthermore, the training method of the component recognition model provided in this embodiment combines various enhancement strategies such as local adaptive contrast enhancement, random rotation, mirroring, cropping, and affine transformation to improve the network's adaptability to complex environments such as illumination changes, noise interference, and geometric deformations, solves the limitations of traditional normalization preprocessing only, and enhances the stability of feature extraction. By adaptively adjusting the momentum coefficient through the gradient change factor, the network can better balance local detail capture and global convergence speed during training, improve convergence stability, and solve the problem that traditional fixed momentum coefficients are insensitive to local gradient changes. The second derivative correction term is used to adjust the weight update, making the network more robust in recognizing components with local occlusion and rich details, and solving the problem that conventional methods only using historical gradient accumulation make it difficult to balance local accuracy and global convergence for high-dimensional features. Based on the cross-entropy loss, a non-Euclidean geometric distance metric is added to enhance the penalty for difficult-to-separate samples, improve the classification accuracy, and combine with the noise factor to reduce the impact of environmental noise and enhance the robustness under the complex background of construction sites. The learning rate is dynamically adjusted according to the contribution degree of feature gradients, enabling the network to converge quickly in the early stage and update smoothly in the later stage, and solving the problem of unstable training caused by traditional fixed learning rates or simple piecewise decay methods. The component recognition model provided in this embodiment can identify building prefabricated components in real time, visually output category labels and confidence levels, and synchronize them to the construction management platform to assist construction management, solving the problems of traditional manual annotation dependence and low recognition efficiency.

[0160] Embodiment 3

[0161] This embodiment provides a computer vision-based building prefabricated component recognition system, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the computer vision-based building prefabricated component recognition method described in Embodiment 1 or Embodiment 2.

[0162] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0163] In the present invention, unless otherwise clearly specified or limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral body; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium; it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0164] In the present invention, unless otherwise clearly specified or limited, when a first feature is "on" or "under" a second feature, it may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, when a first feature is "above", "over" and "on top of" a second feature, it may be that the first feature is directly above or obliquely above the second feature, or simply means that the first feature has a higher horizontal height than the second feature. When a first feature is "under", "beneath" and "underneath" a second feature, it may be that the first feature is directly below or obliquely below the second feature, or simply means that the first feature has a lower horizontal height than the second feature.

[0165] In the description of this specification, the descriptions of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0166] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for identifying building prefabricated components based on computer vision, characterized in that: include: S11, constructing a model training set by collecting construction scenes at different time periods of the construction site through a pre-deployed image acquisition device; The model training set includes a preset number of training images and component type labels corresponding to the assembled components in each training image; Each of the training images includes a prefabricated component in a construction scene; S12, training a pre-built component recognition model according to each training image in the model training set and the component type label corresponding to each training image; During the training process, each time the component recognition model obtains a corresponding predicted output according to a training image, a corresponding cross entropy loss value is obtained according to the training image and the predicted output, and a pre-set cross entropy loss function. Based on the non-Euclidean distance metric between the training image and the predicted output, and a pre-set formula 1, the corresponding cross entropy loss value is dynamically adjusted to obtain a corresponding true loss value; and dynamically adjusting model parameters according to the true loss value; When the end condition is reached, the trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result; The formula 1 is: ; Among them, L is the true loss value, L cross is the cross entropy loss value, λ is the regularization coefficient, is the non-Euclidean distance metric corresponding to the i-th training image, y i is the component category label of the i-th training image, x i is the i-th training image, is the predicted output of the component recognition model for the i-th training image, n is the number of training images input to the component recognition model, is the noise factor; The expression of the non-Euclidean distance metric is: ; Among them, f s (•) is the high-dimensional deep feature representation extracted by the component recognition model for the training image, is the deep feature class center of the component category to which the training image belongs, is the deep feature class variance of the component category to which the training image belongs, is the nearest neighbor of the training image in the component category to which it belongs, G m is the mth Gaussian scaling kernel, M is the number of scales, is the convolution operation, ||•|| F is the Frobenius norm, ||•||2 is the L2 norm, is the weight coefficient of the mth scale, and are all very small positive numbers.

2. The computer vision-based building assembly component recognition method according to claim 1, characterized in that: The step S12, training the pre-built component recognition model according to each training image in the model training set and the component type label corresponding to each training image, includes: According to the preset feature enhancement algorithm, all training images in the model training set are subjected to multi-dimensional enhancement processing; The model training set with multi-dimensional features enhanced is input into the component recognition model for model training.

3. The method for identifying building prefabricated components based on computer vision according to claim 1, characterized in that: The model parameters include momentum coefficient; Then the S12, dynamically adjusting the model parameters according to the true loss value, includes: According to the true loss value of the current training step, obtain the gradient of the true loss value of the current training step with respect to the weight, and save the gradient of the true loss value of the current training step with respect to the weight; According to the gradient of the real loss value of the current training step with respect to the weight, and the pre-saved gradient of the real loss value of all training steps before the current training step with respect to the weight, the gradient change factor corresponding to the current training step is obtained; According to the gradient change factor corresponding to the current training step, as well as the preset initial momentum coefficient and momentum adjustment parameter, the momentum coefficient corresponding to the current training step is dynamically adjusted.

4. The method for identifying building prefabricated components based on computer vision according to claim 3, characterized in that: The model parameters also include convolution kernel weights; Then the S12, dynamically adjusting the model parameters according to the true loss value, also includes: According to the true loss value of the current training step, obtain the second-order derivative of the true loss value of the current training step with respect to the weight; According to the second-order derivative of the true loss value of the current training step with respect to the weight and the gradient of the true loss value of the current training step with respect to the weight, the second-order derivative correction term corresponding to the current training step is obtained; According to the information entropy mean and gradient variance of each position in the feature map output by the convolution layer of the training image input model in the current training step, the training complexity coefficient corresponding to the current training step is obtained; According to the gradient of the true loss value with respect to the weight corresponding to the previous training step, the gradient of the true loss value with respect to the weight corresponding to the current training step, the second-order derivative correction term corresponding to the current training step, the complexity coefficient corresponding to the current training step and the convolution kernel weight corresponding to the current training step, as well as the preset learning rate, the convolution kernel weight corresponding to the next training step is obtained.

5. The method for identifying building prefabricated components based on computer vision according to claim 4, characterized in that: The model parameters also include learning rate; The S12, dynamically adjusting the model parameters according to the true loss value, further includes: In the current training step, before updating and correcting the convolution kernel weights of the component recognition model, the learning rate is dynamically adjusted according to the number of features corresponding to the training image in the current training step; the expression of the learning rate is: ; in, is the initial learning rate, is the learning rate adjustment factor, t is the current training step, int(t) is the index value of the previous t training steps, T is the maximum number of training times, n fea is the number of features in the training image, is the gradient of the kth feature corresponding to the true loss value of the tth training step, is the gradient of the o-th feature corresponding to the true loss value of the t-th training step, is the adjustment factor of the kth feature, and the expression of the feature adjustment factor is: ; in, is the variance of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the mean value of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the sensitivity coefficient, is a very small positive number.

6. The computer vision-based building assembly component recognition method according to claim 2, characterized in that: S12, performing multi-dimensional enhancement processing on all training images received in the model training set according to a preset feature enhancement algorithm, including: Perform stretching / scaling processing on the received training image according to a preset local adaptive contrast enhancement algorithm; According to the preset multi-dimensional data enhancement algorithm, the stretched / scaled training images are randomly rotated, randomly mirrored, randomly cropped, and affine transformed.

7. The method for identifying building prefabricated components based on computer vision according to claim 1, characterized in that: The S12 further includes: Before inputting the model training set into the component recognition model for model training, the weights and biases of the convolution kernel of the component recognition model are initialized so that the weights and biases of the convolution kernel are initialized to obey a normal distribution with a mean of 0 and a variance of 0.

01.

8. The method for identifying building prefabricated components based on computer vision according to claim 1, characterized in that: The S12 also includes: S13, collecting an original image of the target area by a preset image collection device, wherein the original image includes a component to be identified; Inputting the original image into a preset component recognition model to obtain a corresponding component recognition result; The component recognition model includes a recognition model obtained by training a convolutional neural network model through a preset model training set; The model training set includes at least one training image; any of the training images includes at least one component and a component category label pre-labeled for each component; The component recognition model is used to extract features from the received original image through a pre-set convolution layer to obtain a corresponding feature image; And, nonlinear mapping is performed on the feature image through a preset nonlinear activation function ReLU; Furthermore, the feature image after nonlinear mapping is pooled through a preset pooling layer, and the feature image output by the pooling layer is input into a preset fully connected layer to output the probability distribution of the component category in the original image, and the maximum probability value is the corresponding component recognition result.

9. A computer vision-based building assembly component recognition system, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the computer vision-based building prefabricated component recognition method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method, device and equipment for detecting component information in image and readable storage medium

    CN111476351A

  • Person re-identification method and apparatus based on deep learning network, device, and medium

    WO2023272994A1