Building fabricated component identification method and system based on computer vision

By collecting construction scene images at construction sites and dynamically adjusting the loss function and model parameters using non-European geometric distance metrics, the problems of low efficiency, large error and instability in the recognition of building prefabricated components are solved, and high recognition accuracy and rapid convergence training are achieved in complex scenarios.

CN119992346AActive Publication Date: 2025-05-13CHINA STATE CONSTR HAILONG TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510468078.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency, large error and high cost in the identification of prefabricated components of building, and in complex construction site environments, computer vision technology is difficult to maintain high recognition accuracy and training stability.

Method used

The computer vision-based building prefabricated component recognition method is used to construct a model training set by pre-acquisition of construction scene images of construction sites, dynamically adjust the loss function using non-European geometric distance metrics, and dynamically adjust the model parameters to improve recognition accuracy and training stability.

Benefits of technology

In complex scenarios such as lighting changes, occlusions and local similarities, high recognition accuracy is achieved, and the training process converges faster by dynamically adjusting the model parameters, improving the recognition accuracy and training stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992346A_ABST
    Figure CN119992346A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, in particular to a building fabricated component identification method and system based on computer vision, and the method comprises the steps: collecting construction scenes of a construction site at different time periods, and constructing a model training set; according to each training image in the model training set and the component type label corresponding to each training image, training a pre-constructed component identification model; in the training process, when the component identification model obtains the corresponding prediction output according to the training image, the loss function and the model parameters are dynamically adjusted in real time based on the non-Euclidean geometric distance measurement between the training image and the prediction output; and when an end condition is met, obtaining a trained component identification model to detect the to-be-identified component image, and obtaining a component identification result. The method has the beneficial effects that through dynamic optimization of the loss function and the model parameters, the training process is enabled to converge more quickly, and meanwhile, the recognition precision of the component recognition model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular to a method and system for identifying building prefabricated components based on computer vision. Background Art

[0002] With the development of building industrialization and intelligence, prefabricated buildings have been widely promoted around the world due to their advantages of high construction efficiency, controllable quality, and resource conservation. However, since prefabricated buildings are designed with a large number of standardized and non-standardized components, such as steel structures, precast concrete panels, connectors, etc., how to efficiently and accurately identify and manage these components has become a key issue in building construction and quality control.

[0003] At present, the identification of prefabricated components mainly relies on manual inspection and management, that is, construction workers identify and manage components through visual judgment or based on manual records. However, this method has many shortcomings, including low efficiency, large errors, high costs, and in large construction sites, manual identification methods are difficult to meet the needs of accurate positioning, quality tracking, construction progress management, etc. Therefore, using computer vision technology to achieve automated and intelligent component identification is of great significance to improving the level of construction management of prefabricated buildings.

[0004] However, the application of existing computer vision recognition technology in the recognition of building components still faces many challenges. First, the environment of the construction site is complex. Factors such as lighting changes, noise interference, and perspective distortion will affect the image quality, making it difficult for existing methods to extract stable features. Secondly, there are many types of prefabricated components, and some components are similar in shape, such as different types of steel components or precast concrete panels. Traditional classification methods based on convolutional neural networks are prone to misclassification when distinguishing these similar components. In addition, components may be partially occluded, missing, or damaged during the construction process, further increasing the difficulty of recognition. The traditional convolutional neural network training process uses a fixed momentum coefficient, a fixed learning rate, and a simple cross-entropy loss method that is difficult to cope with the complex construction site environment, and is prone to problems such as unstable training, reduced recognition accuracy, and insufficient generalization ability. Summary of the invention

[0005] 1. Technical issues to be resolved In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method and system for identifying prefabricated building components based on computer vision, which solves the technical problems that manual detection methods are inefficient, have large errors, are costly, and are difficult to meet the needs of precise positioning, quality tracking, construction progress management, etc. at large construction sites, as well as the technical problems that existing computer vision technology is difficult to cope with complex construction site environments, easily leading to unstable training, reduced recognition accuracy, and insufficient generalization ability.

[0006] (II) Technical solution In order to achieve the above object, the main technical solutions adopted by the present invention include: In a first aspect, an embodiment of the present invention provides a method for identifying assembled building components based on computer vision, comprising: S11, constructing a model training set by collecting construction scenes at different time periods of the construction site through a pre-deployed image acquisition device; The model training set includes a preset number of training images and component type labels corresponding to the assembled components in each training image; Each of the training images includes a prefabricated component in a construction scene; S12, training a pre-built component recognition model according to each training image in the model training set and the component type label corresponding to each training image; During the training process, each time the component recognition model obtains a corresponding prediction output according to a training image, the loss function is dynamically adjusted based on a non-Euclidean distance metric between the training image and the prediction output, and the model parameters are dynamically adjusted according to the dynamically adjusted loss function; When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

[0007] Optionally, the S12 includes: According to the preset feature enhancement algorithm, all training images in the model training set are subjected to multi-dimensional enhancement processing; The model training set after multi-dimensional feature enhancement is input into the component recognition model for model training. During the training process, each time the component recognition model obtains a corresponding prediction output according to a training image, a corresponding cross entropy loss value is obtained according to the training image and the prediction output, as well as a pre-set cross entropy loss function; And, dynamically adjusting the corresponding cross entropy loss value according to the non-Euclidean distance metric between the training image and the predicted output to obtain the corresponding true loss value; And, dynamically adjusting the model parameters according to the true loss value; When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

[0008] Optionally, the S12, dynamically adjusting the corresponding cross entropy loss value according to the non-Euclidean distance metric between the training image and the predicted output to obtain the corresponding true loss value, includes: Based on the non-Euclidean distance metric between the training image and the predicted output, and a preset formula 1, the corresponding cross entropy loss value is dynamically adjusted; the formula 1 is: ; Among them, L is the true loss value, L cross is the cross entropy loss value, λ is the regularization coefficient, is the non-Euclidean distance metric corresponding to the i-th training image, y i is the component category label of the i-th training image, x i is the i-th training image, is the predicted output of the component recognition model for the i-th training image, n is the number of training images input to the component recognition model, is the noise factor; The expression of the non-Euclidean distance metric is: ; Among them, f s (•) is the high-dimensional deep feature representation extracted by the component recognition model for the training image, is the deep feature class center of the component category to which the training image belongs, is the deep feature class variance of the component category to which the training image belongs, is the nearest neighbor of the training image in the component category to which it belongs, G m is the mth Gaussian scale kernel, M is the number of scales, is the convolution operation, ||•|| F is the Frobenius norm, ||•||2 is the L2 norm, is the weight coefficient of the mth scale, and are all very small positive numbers.

[0009] Optionally, the model parameters include momentum coefficients; Then the S12, dynamically adjusting the model parameters according to the true loss value, includes: According to the true loss value of the current training step, obtain the gradient of the true loss value of the current training step with respect to the weight, and save the gradient of the true loss value of the current training step with respect to the weight; According to the gradient of the real loss value of the current training step with respect to the weight, and the pre-saved gradient of the real loss value of all training steps before the current training step with respect to the weight, the gradient change factor corresponding to the current training step is obtained; According to the gradient change factor corresponding to the current training step, as well as the preset initial momentum coefficient and momentum adjustment parameter, the momentum coefficient corresponding to the current training step is dynamically adjusted.

[0010] Optionally, the model parameters also include convolution kernel weights; Then the S12, dynamically adjusting the model parameters according to the true loss value, also includes: According to the true loss value of the current training step, obtain the second-order derivative of the true loss value of the current training step with respect to the weight; According to the second-order derivative of the true loss value of the current training step with respect to the weight and the gradient of the true loss value of the current training step with respect to the weight, the second-order derivative correction term corresponding to the current training step is obtained; According to the information entropy mean and gradient variance of each position in the feature map output by the convolution layer of the training image input model in the current training step, the training complexity coefficient corresponding to the current training step is obtained; According to the gradient of the true loss value with respect to the weight corresponding to the previous training step, the gradient of the true loss value with respect to the weight corresponding to the current training step, the second-order derivative correction term corresponding to the current training step, the complexity coefficient corresponding to the current training step and the convolution kernel weight corresponding to the current training step, as well as the preset learning rate, the convolution kernel weight corresponding to the next training step is obtained.

[0011] Optionally, the model parameters also include a learning rate; The S12, dynamically adjusting the model parameters according to the true loss value, further includes: In the current training step, before updating and correcting the convolution kernel weights of the component recognition model, the learning rate is dynamically adjusted according to the number of features corresponding to the training image in the current training step; the expression of the learning rate is: ; in, is the initial learning rate, is the learning rate adjustment factor, t is the current training step, int(t) is the index value of the previous t training steps, T is the maximum number of training times, n fea is the number of features in the training image, is the gradient of the kth feature corresponding to the true loss value of the tth training step, is the gradient of the o-th feature corresponding to the true loss value of the t-th training step, is the adjustment factor of the kth feature, and the expression of the feature adjustment factor is: ; in, is the variance of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the mean value of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the sensitivity coefficient, is a very small positive number.

[0012] Optionally, the step S12, performing multi-dimensional enhancement processing on all training images received in the model training set according to a preset feature enhancement algorithm, includes: Perform stretching / scaling processing on the received training image according to a preset local adaptive contrast enhancement algorithm; According to the preset multi-dimensional data enhancement algorithm, the stretched / scaled training images are randomly rotated, randomly mirrored, randomly cropped, and affine transformed.

[0013] Optionally, the S12 further includes: Before inputting the model training set into the component recognition model for model training, the weights and biases of the convolution kernel of the component recognition model are initialized so that the weights and biases of the convolution kernel are initialized to obey a normal distribution with a mean of 0 and a variance of 0.01.

[0014] Optionally, after S12, the step further includes: S13, collecting an original image of the target area by a preset image collection device, wherein the original image includes a component to be identified; Inputting the original image into a preset component recognition model to obtain a corresponding component recognition result; The component recognition model includes a recognition model obtained by training a convolutional neural network model through a preset model training set; The model training set includes at least one training image; any of the training images includes at least one component and a component category label pre-labeled for each component; The component recognition model is used to extract features from the received original image through a pre-set convolution layer to obtain a corresponding feature image; And, nonlinear mapping is performed on the feature image through a preset nonlinear activation function ReLU; Furthermore, the feature image after nonlinear mapping is pooled through a preset pooling layer, and the feature image output by the pooling layer is input into a preset fully connected layer to output the probability distribution of the component category in the original image, and the maximum probability value is the corresponding component recognition result.

[0015] In a second aspect, an embodiment of the present invention provides a computer vision-based building prefabricated component recognition system, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-mentioned computer vision-based building prefabricated component recognition method.

[0016] (III) Beneficial effects The beneficial effects of the present invention are as follows: a computer vision-based building prefabricated component recognition method of the present invention dynamically optimizes the loss function through non-Euclidean distance measurement, so that the model can still maintain high recognition accuracy in complex scenes such as lighting changes, occlusion, and local similarity, and the dynamic adjustment of model parameters makes the training process converge faster while ensuring higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a method for identifying assembled building components based on computer vision provided in Example 1 of the present invention; Figure 2 A schematic flow chart of a method for identifying assembled building components based on computer vision provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0018] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.

[0019] The embodiment of the present invention proposes a method for identifying prefabricated building components based on computer vision. The loss function is dynamically optimized through non-Euclidean distance measurement, so that the model can still maintain high recognition accuracy in complex scenarios such as lighting changes, occlusion, and local similarity. The dynamic adjustment of model parameters makes the training process converge faster while ensuring higher accuracy.

[0020] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0021] Example 1

[0022] This embodiment provides a method for identifying building prefabricated components based on computer vision. Figure 1 As shown, including: S11, constructing a model training set by collecting construction scenes at different time periods of the construction site through a pre-deployed image acquisition device; The model training set includes a preset number of training images and component type labels corresponding to the assembled components in each training image; Each of the training images includes a prefabricated component in a construction scene; S12, training the pre-built component recognition model according to each training image in the model training set and the component type label corresponding to each training image; During the training process, each time the component recognition model obtains a corresponding prediction output according to a training image, the loss function is dynamically adjusted in real time based on the non-Euclidean distance measurement between the training image and the prediction output, and the model parameters are dynamically adjusted in real time according to the dynamically adjusted loss function; When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

[0023] This embodiment provides a method for identifying prefabricated building components based on computer vision. The loss function is dynamically optimized through non-Euclidean distance measurement, so that the model can maintain high recognition accuracy in complex scenarios such as lighting changes, occlusion, and local similarity. The dynamic adjustment of model parameters enables the training process to converge faster while ensuring higher accuracy.

[0024] Example 2

[0025] This embodiment provides a method for identifying building prefabricated components based on computer vision. Figure 2 As shown, the method is used for identifying building prefabricated components, and the method includes: S1. Source, method and storage format of training data collection; The model training set of this embodiment is derived from actual construction scenes at construction sites, including but not limited to various assembled components such as steel structures and precast concrete panels; The acquisition method uses a high-resolution digital camera system (image acquisition device) fixed at different angles to cover the full range of components, ensuring that the acquired training image data can fully reflect the component characteristics and environmental information; The data storage format of the training images is JPEG, and the resolution of each image is not less than 1920×1080 pixels.

[0026] S2, data preprocessing; Due to the complex environment of construction sites, training images may be affected by lighting changes, noise interference, perspective distortion, etc. Gaussian filtering and bilateral filtering are used to remove random noise and improve image quality.

[0027] S3, data annotation; The collected training images are annotated with data, and the annotated results are in the form of component type labels of the assembled components in the training images; the annotated type labels include: Type 1: Steel structural components (such as H-shaped steel beams, steel columns, steel connecting plates); Type 2: Precast concrete components (such as precast floor slabs, precast wall panels, precast beams, precast columns); Type 3: Connectors (such as bolts, welds, embedded parts); Type 4: Auxiliary equipment (support frame, construction tools); … Type n:…….

[0028] A model training set is constructed using training images and component type labels corresponding to the training images.

[0029] S4, building assembly component recognition model training; S401, performing multi-dimensional enhancement processing on the training images in the model training set; For training images, uneven lighting, noise and partial missing components are prone to occur during the actual acquisition process. These image characteristics complicate the subsequent feature extraction and classification steps.

[0030] In the prior art, usually only a simple normalization process is performed, which cannot fully solve the interference caused by illumination, noise and geometric deformation, and easily leads to unstable feature extraction effect.

[0031] This embodiment performs multi-dimensional enhancement on the image to ensure that the convolution operation of the subsequent convolutional neural network is more adaptable when extracting key areas and texture features of building components; that is, according to a preset feature enhancement algorithm, multi-dimensional enhancement processing is performed on all training images received in the model training set; the feature enhancement algorithm is expressed as: ; in, It is an image processed by multi-dimensional feature enhancement, representing the results after processing in terms of illumination, noise, and geometric transformation; The received training images represent component images collected from the construction site or prefabrication scene; It is a normalization function, which represents the stretching or scaling of the image grayscale value or color range; is a data enhancement function, which represents random rotation, mirroring and other operations on the image; The normalization function stretches or scales the grayscale or color distribution of the training image by means of local adaptive contrast enhancement, effectively reducing the impact of uneven illumination and suppressing image noise. The calculation method of the local adaptive contrast enhancement algorithm is expressed as: ; Where I(a,b) is the pixel value of the training image at position (a,b), is the local mean of pixels in a window of size c×c centered at position (a,b), is the local standard deviation of pixels in a window of size c×c centered at position (a,b), is a very small positive number to prevent division by zero; (a, b) is the coordinate value, and c is the window size. Preferably, is 0.0001, and c is set to 3.

[0032] The data enhancement function combines random rotation, mirroring, random cropping and affine transformation to comprehensively implement multi-dimensional data enhancement and enhance the network's robustness to geometric deformation. The calculation method of the multi-dimensional data enhancement algorithm is expressed as: ; Among them, A aff (•) is a random affine transformation operation, C crop (•) is a random cropping operation, M flip (•) is a random mirror operation, is a random rotation operation, is the rotation angle.

[0033] S402: Initialize the parameters of the convolutional neural network.

[0034] By selecting the normal distribution to initialize the weights and bias of the convolution kernel of the component recognition model (usually a convolutional neural network model), the problem of gradient explosion or disappearance during the early training of the building component image (training image) can be avoided, and the subsequent gradient can be guaranteed to be stable. Specifically, the weights and bias of the convolution kernel are initialized to obey the normal distribution with a mean of 0 and a variance of 0.01.

[0035] S403, performing forward propagation calculation of the component recognition model (convolutional neural network model).

[0036] By inputting the training image that has undergone multi-dimensional enhancement processing into the convolutional neural network (component recognition model), the convolution layer operation is first performed to extract the local features of the image; then the features are nonlinearly mapped through the nonlinear activation function ReLU; the size of the feature map is reduced through the pooling layer to reduce the network parameters; finally, the feature map output by the pooling layer is expanded into a one-dimensional vector and sent to the fully connected layer to output the probability distribution corresponding to the training image. The one with the largest probability based on the maximum principle is taken as the corresponding prediction output, which is the component type recognition result corresponding to the training image.

[0037] S404, calculating the momentum coefficient of the convolutional neural network (component identification model).

[0038] During the convolutional neural network training process of the model training set, due to the complex component details, many blocks and large local differences.

[0039] Conventional convolutional neural networks are easily insensitive to gradient fluctuations caused by local details of building components, and lose some detailed information of building components.

[0040] In this embodiment, a dynamic momentum coefficient is set to perform adaptive adjustment according to the gradient change in each training step, and the gradient change factor is used to quantify the gradient accumulation difference, so as to find a balance between high-speed convergence and local stability, and overcome the difficulty that conventional technologies are insensitive to local gradient changes. The calculation method of the momentum coefficient is expressed as: ; Among them, u t is the momentum coefficient of the t-th training step, which represents the accumulation of the gradient of the previous step during the update; u0 is the initial momentum coefficient, which represents the initial setting of the momentum; α is the momentum adjustment parameter, which represents the amplitude of the dynamic change of momentum with training; β t is the gradient change factor of the tth training step, which characterizes the influence of the gradient difference between the current training step and the historical training step; t is a positive integer. Preferably, u0 is set to 0.1 and α is set to 2.

[0041] The gradient change factor dynamically adjusts the momentum coefficient based on the gradient change of the current training step. The expression of the gradient change factor is: ; Among them, γ is the gradient adjustment constant, which represents the amplification or reduction of the accumulated gradient difference; is the gradient of the loss function of the jth training step with respect to the weight, int(t-1) is the index of the previous t-1 training steps, is the gradient of the loss function of the j-1th training step with respect to the weight. i is a positive integer. Preferably, γ is set to 0.5.

[0042] S405: Update the weight parameters of the convolutional neural network (component recognition model).

[0043] Still in the training of convolutional neural networks for images of building prefabricated components, due to the complex details of the components, many blocks, and large local differences, conventional convolutional neural networks usually do not consider the historical changes of gradients, or only use the simple accumulation of current and historical gradients. It is often difficult to balance local accuracy and global convergence speed when facing high-dimensional features and local significant differences.

[0044] This embodiment combines the momentum coefficient with the second-order information, and uses the second-order information to correct the weight update, so as to correct the second-order derivative of the loss during the convolutional neural network training process, so as to better handle the component features with large local differences. That is, according to the dynamically adjusted momentum coefficient and second-order information, the convolution kernel weights of the component recognition model are updated and corrected; the second-order information is the second-order derivative information of the loss function with respect to the model parameters; the expression for updating and correcting the convolution kernel weights of the component recognition model is: ; Among them, W t+1 is the convolution kernel weight of the t+1th training step, W t is the convolution kernel weight of the tth training step, is the learning rate of the tth iteration, is the training complexity coefficient, is the second-order derivative correction term, obtained through the second-order information, is the gradient of the weight corresponding to the loss function of the tth training step; is the gradient of the weight corresponding to the loss function of the t-1th training step, and the expression of the second-order derivative correction term is: ; in, is the second-order derivative of the weight corresponding to the loss function of the tth training step.

[0045] The training complexity coefficient is determined based on the local feature complexity and occlusion level of the current training image in the convolutional neural network.

[0046] Specifically, by analyzing the feature map of the input image after the convolution layer, the local information entropy and the variance of the gradient amplitude of each feature channel are calculated; by statistically analyzing the mean information entropy and gradient variance of each position in the feature map, and normalizing and weighting them, we get ; Among them, E t is the mean information entropy of each position in the feature map of the tth training step, V t is the gradient variance of each position in the feature map of the tth training step, is a very small positive number.

[0047] The training complexity coefficient dynamically adjusts the strength of the second-order derivative correction term, so that the second-order correction is enhanced on complex training images (such as locally missing areas), thereby improving the model's robustness to occlusion.

[0048] S406: Calculate the classification result (prediction output).

[0049] Based on step S403, after convolution, planning, pooling and full connection layer operations, the prediction probability distribution corresponding to each category is obtained, and then the classification result is determined by the probability value maximization principle, that is, the category with the largest probability is used as the final prediction output.

[0050] S407, calculating the loss function of the convolutional neural network.

[0051] Samples that are extremely similar locally in images of building prefabricated components will make it difficult for the network to distinguish them during classification. The loss function needs to pay more attention to the difficult samples to improve the overall classification accuracy. Conventional convolutional neural networks usually use the cross-entropy loss function, which only cares about the global error and does not differentiate between local or difficult samples, resulting in insufficient network attention allocation.

[0052] This embodiment combines the non-Euclidean interval metric correction term in the cross entropy loss, so that the difficult-to-separate samples get a higher penalty weight in the formula, and at the same time adopts a noise factor to suppress the environmental noise in the acquisition process, thereby taking into account both complex features and noise interference. That is, during the training process, every time the component recognition model obtains the corresponding predicted output based on the training image, the loss function is dynamically adjusted in real time based on the non-Euclidean distance metric between the training image and the predicted output, and the model parameters are dynamically adjusted in real time based on the dynamically adjusted loss function; the expression of the loss function (Formula 1) is: ; Among them, L is the loss function, which represents the total loss of the model; L cross is the cross entropy function; λ is the regularization coefficient, which represents the weight of non-Euclidean geometry metrics; is the non-Euclidean distance metric of the i-th training image; y i is the component category label of the i-th training image, x i is the i-th training image, representing the i-th training image input into the model; is the predicted output of the component recognition model for the i-th training image, n is the number of training images input to the component recognition model, is the noise factor, which represents the suppression strength of noisy samples or incomplete images. Preferably, λ is set to 0.3. Set to 0.01.

[0053] The non-Euclidean distance metric is implemented by combining the multi-scale local feature differences of component images and the semantic distance between categories through non-linear manifold embedding. The expression of the non-Euclidean distance metric is: ; Among them, f s(•) is the high-dimensional deep feature representation extracted by the component recognition model for the training image, is the center of the deep feature class of the component category to which the training image belongs, which is calculated by the mean of the features of samples of the same category; is the deep feature class variance of the component category to which the training image belongs, reflecting the degree of discreteness of the features within the category; is the nearest neighbor of the training image in the component category (measured by the K nearest neighbor algorithm), representing the boundary fuzzy point between categories; G m is the mth Gaussian scale kernel, which is used to extract the differences of image features at different scales; M is the number of scales, which is preferably set to 3 to capture the differences in component features of different granularities from small scale to large scale; is the convolution operation, ||•|| F is the Frobenius norm, which represents the scale sensitivity of the feature difference matrix; ||•||2 is the L2 norm, is the weight coefficient of the mth scale, which is used to dynamically balance the difference contribution at different scales. It is preset manually. For example, when M is set to 3, the weight coefficient of the first scale is set to 0.1, the weight coefficient of the second scale is set to 0.5, and the weight coefficient of the third scale is set to 0.4; and are all very small positive numbers to prevent division by zero. Preferably, and Both are set to 0.0001.

[0054] That is, through the expression of the above loss function, the predicted output corresponding to each training step and the true loss value corresponding to the training image can be obtained in each training step to dynamically adjust the model parameters.

[0055] S408. Dynamically adjust the learning rate of the convolutional neural network (component recognition model).

[0056] In the training of convolutional neural networks for images of architectural prefabricated components, due to the complex details of the components, the large number of blocks, and the large local differences, as the depth of the convolutional neural network and the growth of the data scale, if the learning rate is not dynamically adjusted, it may converge too slowly in the early stage or oscillate too much in the later stage. Conventional methods fix the learning rate or simply decay in segments, which cannot adaptively adjust the differences in the importance of different features in the component images, resulting in reduced training stability.

[0057] This embodiment dynamically monitors the importance of the proportion of each feature gradient amplitude and adaptively adjusts the learning rate of the network, so that the network converges quickly in the early stage and focuses on adjusting the key gradient, while steadily reducing the update amplitude in the later stage. Specifically, the contribution of the feature gradient to the learning rate is dynamically weighed to achieve efficient and stable convergence. That is, in the training step, before updating and correcting the convolution kernel weights of the component recognition model, the learning rate is dynamically adjusted according to the number of features corresponding to the training image in the training step; the expression of the learning rate is: ; in, is the learning rate of the tth training step, representing the gradient update amplitude of the current update iteration; is the initial learning rate; is the learning rate adjustment factor, which represents the decay rate of the learning rate as the number of training steps increases; int(t) is the index value of the previous t training steps, T is the maximum number of training steps, n fea is the number of features in the training image, is the gradient of the k-th feature corresponding to the loss function of the t-th training step, is the gradient of the o-th feature corresponding to the loss function of the t-th training step, is the adjustment factor of the kth feature, representing the influence of the gradient corresponding to the feature on the learning rate; preferably, Set to 0.01, Set to 0.3.

[0058] The adjustment factor of the feature is determined by measuring the contribution of the feature to the network prediction accuracy at different training stages. The expression of the adjustment factor of the feature is: ; in, is the variance of the gradient of the kth feature corresponding to the loss function of the component recognition model in the past h training steps, representing the strength of the gradient fluctuation; is the mean value of the gradient of the kth feature corresponding to the loss function of the component recognition model in the past h training steps; is the sensitivity coefficient, which is used to adjust the intensity of the impact of the fluctuation gradient on the learning rate adjustment; A very small positive number to prevent division by zero.

[0059] S409, judging the completion of convolutional neural network (component recognition model) training; When the difference between the loss function value of this training step and the previous training step is less than the preset threshold, the iteration is stopped, indicating that the convolutional neural network has reached a convergence state and the convolutional neural network model training is completed.

[0060] S5. Identification of building prefabricated components; After the model training is completed, the trained convolutional neural network is deployed to the actual application scenario of the construction site for recognition through real-time or pre-stored component images.

[0061] First, the input image needs to go through the same preprocessing process as the training stage, including Gaussian filtering, bilateral filtering denoising, and local adaptive contrast enhancement to eliminate environmental interference and improve image quality; Furthermore, the enhanced image is input into the trained convolutional neural network, multi-scale deep features are extracted through multi-layer convolution, activation and pooling operations, and the probability distribution of component categories is generated using a fully connected layer; Furthermore, the model determines the category of the target component (such as steel structure, precast concrete, connector or auxiliary facility) based on the maximum probability, and dynamically optimizes the classification boundary by combining the non-Euclidean distance metric correction term to ensure accurate distinction of locally similar or occluded samples; Furthermore, the recognition results are output through a visual interface, with component categories and confidence levels marked, and synchronized to the construction management platform to assist on-site personnel in component positioning, quality inspection, and progress tracking.

[0062] For example: At a large-scale prefabricated building construction site, the computer vision-based building prefabricated component recognition method proposed in the present invention is used to collect images of steel structures, precast concrete components, connectors, etc. at different construction stages using a high-resolution camera, and real-time recognition and classification are performed through a trained convolutional neural network. First, the collected images are subjected to Gaussian filtering denoising and local adaptive contrast enhancement to reduce environmental interference. Then, after multi-layer convolution, pooling, activation and other processing, the multi-scale features of the components are extracted, and the classification algorithm optimized by non-Euclidean geometry is combined to achieve accurate recognition of components of different categories, and the component category, confidence and location information are displayed in real time on the visual interface. Finally, the recognition system is linked with the construction management platform to realize automatic classification of components, progress tracking, and quality inspection, reducing manual operations and improving the intelligent level of construction management.

[0063] The present embodiment provides a method for identifying prefabricated building components based on computer vision, which uses multi-dimensional data enhancement and non-Euclidean loss function optimization to enable the model to maintain high recognition accuracy in complex scenes such as lighting changes, occlusion, and local similarity. The dynamic momentum coefficient and second-order information optimization weight update effectively balance local feature extraction and global convergence speed, making the training process more stable and reducing the possibility of training failure. By using dynamic learning rate adjustment and non-Euclidean loss optimization, the model can better distinguish between components of different categories, and can accurately identify them even in the case of similar structures or partial missing. By adaptively adjusting the momentum coefficient, the training process converges faster while maintaining a high classification accuracy. The learning rate is adjusted through noise suppression strategies and feature gradients, so that the network can still accurately classify under high noise and complex backgrounds.

[0064] Furthermore, the training method of the component recognition model provided in this embodiment combines a variety of enhancement strategies such as local adaptive contrast enhancement, random rotation, mirroring, cropping, and affine transformation to improve the adaptability of the network to complex environments such as illumination changes, noise interference, and geometric deformation, solves the limitations of traditional normalization preprocessing, and enhances the stability of feature extraction. The momentum coefficient is adaptively adjusted by the gradient change factor, so that the network can better balance the capture of local details and the global convergence speed during the training process, improve the convergence stability, and solve the problem that the traditional fixed momentum coefficient is insensitive to local gradient changes. The weight update is adjusted by using the second-order derivative correction term, so that the network is more robust when identifying components with local occlusion and rich details, and solves the problem that the conventional only uses historical gradient accumulation, resulting in high-dimensional features that are difficult to take into account local accuracy and global convergence. On the basis of cross entropy loss, a non-Euclidean distance metric is added to enhance the penalty for difficult samples, improve classification accuracy, combine noise factors to reduce the impact of environmental noise, and enhance robustness against complex backgrounds of construction sites. By dynamically adjusting the learning rate based on the contribution of feature gradients, the network converges quickly in the early stage and updates smoothly in the later stage, solving the problem of unstable training caused by traditional fixed learning rate or simple segmented attenuation method. The component recognition model provided in this embodiment can recognize building prefabricated components in real time, and visualize the output category labels and confidence levels, synchronize them to the construction management platform, and assist in construction management, solving the problem of traditional reliance on manual labeling and low recognition efficiency.

[0065] Example 3 This embodiment provides a computer vision-based building prefabricated component recognition system, including a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement a computer vision-based building prefabricated component recognition method described in Example 1 or Example 2.

[0066] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0067] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0068] In the present invention, unless otherwise clearly specified and limited, when a first feature is “on” or “below” a second feature, it may be that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Moreover, when a first feature is “above”, “above” or “above” a second feature, it may be that the first feature is directly above or obliquely above the second feature, or it may simply mean that the first feature is higher in level than the second feature. When a first feature is “below”, “below” or “below” a second feature, it may be that the first feature is directly below or obliquely below the second feature, or it may simply mean that the first feature is lower in level than the second feature.

[0069] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0070] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may alter, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for identifying building prefabricated components based on computer vision, characterized in that: include: S11, constructing a model training set by collecting construction scenes at different time periods of the construction site through a pre-deployed image acquisition device; The model training set includes a preset number of training images and component type labels corresponding to the assembled components in each training image; Each of the training images includes a prefabricated component in a construction scene; S12, training a pre-built component recognition model according to each training image in the model training set and the component type label corresponding to each training image; During the training process, each time the component recognition model obtains a corresponding prediction output according to a training image, the loss function is dynamically adjusted based on a non-Euclidean distance metric between the training image and the prediction output, and the model parameters are dynamically adjusted according to the dynamically adjusted loss function; When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

2. The computer vision-based building assembly component recognition method according to claim 1, characterized in that: The S12 includes: According to the preset feature enhancement algorithm, all training images in the model training set are subjected to multi-dimensional enhancement processing; The model training set after multi-dimensional feature enhancement is input into the component recognition model for model training. During the training process, each time the component recognition model obtains a corresponding prediction output according to a training image, a corresponding cross entropy loss value is obtained according to the training image and the prediction output, as well as a pre-set cross entropy loss function; And, dynamically adjusting the corresponding cross entropy loss value according to the non-Euclidean distance metric between the training image and the predicted output to obtain the corresponding true loss value; And, dynamically adjusting the model parameters according to the true loss value; When the end condition is reached, a trained component recognition model is obtained to detect the component image to be recognized and obtain the component recognition result.

3. The computer vision-based building assembly component recognition method according to claim 2, characterized in that: The step S12, dynamically adjusting the corresponding cross entropy loss value according to the non-Euclidean distance metric between the training image and the predicted output to obtain the corresponding true loss value, includes: Based on the non-Euclidean distance metric between the training image and the predicted output, and a preset formula 1, the corresponding cross entropy loss value is dynamically adjusted; the formula 1 is: ; Among them, L is the true loss value, L cross is the cross entropy loss value, λ is the regularization coefficient, is the non-Euclidean distance metric corresponding to the i-th training image, y i is the component category label of the i-th training image, x i is the i-th training image, is the predicted output of the component recognition model for the i-th training image, n is the number of training images input to the component recognition model, is the noise factor; The expression of the non-Euclidean distance metric is: ; Among them, f s (•) is the high-dimensional deep feature representation extracted by the component recognition model for the training image, is the deep feature class center of the component category to which the training image belongs, is the deep feature class variance of the component category to which the training image belongs, is the nearest neighbor of the training image in the component category to which it belongs, G m is the mth Gaussian scale kernel, M is the number of scales, is the convolution operation, ||•|| F is the Frobenius norm, ||•||2 is the L2 norm, is the weight coefficient of the mth scale, and are all very small positive numbers.

4. The computer vision-based building assembly component recognition method according to claim 2, characterized in that: The model parameters include momentum coefficient; Then the S12, dynamically adjusting the model parameters according to the true loss value, includes: According to the true loss value of the current training step, obtain the gradient of the true loss value of the current training step with respect to the weight, and save the gradient of the true loss value of the current training step with respect to the weight; According to the gradient of the real loss value of the current training step with respect to the weight, and the pre-saved gradient of the real loss value of all training steps before the current training step with respect to the weight, the gradient change factor corresponding to the current training step is obtained; According to the gradient change factor corresponding to the current training step, as well as the preset initial momentum coefficient and momentum adjustment parameter, the momentum coefficient corresponding to the current training step is dynamically adjusted.

5. The method for identifying building prefabricated components based on computer vision according to claim 4, characterized in that: The model parameters also include convolution kernel weights; Then the S12, dynamically adjusting the model parameters according to the true loss value, also includes: According to the true loss value of the current training step, obtain the second-order derivative of the true loss value of the current training step with respect to the weight; According to the second-order derivative of the true loss value of the current training step with respect to the weight and the gradient of the true loss value of the current training step with respect to the weight, the second-order derivative correction term corresponding to the current training step is obtained; According to the information entropy mean and gradient variance of each position in the feature map output by the convolution layer of the training image input model in the current training step, the training complexity coefficient corresponding to the current training step is obtained; According to the gradient of the true loss value with respect to the weight corresponding to the previous training step, the gradient of the true loss value with respect to the weight corresponding to the current training step, the second-order derivative correction term corresponding to the current training step, the complexity coefficient corresponding to the current training step and the convolution kernel weight corresponding to the current training step, as well as the preset learning rate, the convolution kernel weight corresponding to the next training step is obtained.

6. The computer vision-based building assembly component recognition method according to claim 5, characterized in that: The model parameters also include learning rate; The S12, dynamically adjusting the model parameters according to the true loss value, further includes: In the current training step, before updating and correcting the convolution kernel weights of the component recognition model, the learning rate is dynamically adjusted according to the number of features corresponding to the training image in the current training step; the expression of the learning rate is: ; in, is the initial learning rate, is the learning rate adjustment factor, t is the current training step, int(t) is the index value of the previous t training steps, T is the maximum number of training times, n fea is the number of features in the training image, is the gradient of the kth feature corresponding to the true loss value of the tth training step, is the gradient of the o-th feature corresponding to the true loss value of the t-th training step, is the adjustment factor of the kth feature, and the expression of the feature adjustment factor is: ; in, is the variance of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the mean value of the gradient of the kth feature corresponding to the true loss value of the component recognition model in the past h training steps, is the sensitivity coefficient, is a very small positive number.

7. The computer vision-based building assembly component recognition method according to claim 2, characterized in that: S12, performing multi-dimensional enhancement processing on all training images received in the model training set according to a preset feature enhancement algorithm, including: Perform stretching / scaling processing on the received training image according to a preset local adaptive contrast enhancement algorithm; According to the preset multi-dimensional data enhancement algorithm, the stretched / scaled training images are randomly rotated, randomly mirrored, randomly cropped, and affine transformed.

8. The method for identifying building prefabricated components based on computer vision according to claim 1, characterized in that: The S12 further includes: Before inputting the model training set into the component recognition model for model training, the weights and biases of the convolution kernel of the component recognition model are initialized so that the weights and biases of the convolution kernel are initialized to obey a normal distribution with a mean of 0 and a variance of 0.

01.

9. The method for identifying building prefabricated components based on computer vision according to claim 1, characterized in that: The S12 also includes: S13, collecting an original image of the target area by a preset image collection device, wherein the original image includes a component to be identified; Inputting the original image into a preset component recognition model to obtain a corresponding component recognition result; The component recognition model includes a recognition model obtained by training a convolutional neural network model through a preset model training set; The model training set includes at least one training image; any of the training images includes at least one component and a component category label pre-labeled for each component; The component recognition model is used to extract features from the received original image through a pre-set convolution layer to obtain a corresponding feature image; And, nonlinear mapping is performed on the feature image through a preset nonlinear activation function ReLU; Furthermore, the feature image after nonlinear mapping is pooled through a preset pooling layer, and the feature image output by the pooling layer is input into a preset fully connected layer to output the probability distribution of the component category in the original image, and the maximum probability value is the corresponding component recognition result.

10. A computer vision-based building assembly component recognition system, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the computer vision-based building prefabricated component recognition method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • A method of building component extraction based on a Faster-RCNN model

    CN109002841A

  • Method, device and equipment for detecting component information in image and readable storage medium

    CN111476351A

  • Person re-identification method and apparatus based on deep learning network, device, and medium

    WO2023272994A1