A machine vision-based drilling sample defect intelligent identification method
Patent Information
- Application Number
- CN202610911604.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0005]为解决现有技术静态合并及缺乏动态终止机制导致土样缺陷识别准确率低的技术问题,本发明提出了一种基于机器视觉的钻探取样土样缺陷智能识别方法
本发明通过利用双分支卷积神经网络提取图像中缺陷基元的分割掩膜、特征向量及置信度,构建融合特征相似度、结构连续性以及置信度增益的合并得分计算机制,并根据初始置信度的谐波均值分配加权项权重,配合基于概率分布的采样策略与基于信息熵的客观终止条件,避免了传统生硬合并造成的误判与遗漏,确保了复合区域生成的科学性,通过计算特征空间内的马氏距离进行缺陷类别判定,排除了特征变量间相关性的干扰,提升了钻探取样土样缺陷智能识别与定位的准确率。
Smart Images

Figure CN122453828B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology. More specifically, this invention relates to a machine vision-based intelligent identification method for defects in drilled soil samples. Background Technology
[0002] In the fields of geological exploration and geotechnical engineering, soil core samples obtained through drilling are crucial for assessing the mechanical properties of strata and engineering geological conditions. For a long time, the identification of defects in soil core samples has relied heavily on manual visual observation and measurement records by geologists. This traditional method is not only extremely inefficient and time-consuming, but also highly susceptible to the subjective experience and fatigue of the inspectors, leading to inconsistent evaluation standards. Furthermore, soil samples are easily subjected to mechanical disturbance during drilling, resulting in secondary fractures or deformations; genuine natural defects and man-made damage caused by drilling are often intricately intertwined. Faced with the ever-increasing volume of modern engineering exploration data, the traditional subjective judgment model relying on manual methods has revealed serious shortcomings in terms of accuracy, efficiency, and consistency, and is far from meeting the urgent need for high-precision processing of massive amounts of data in current geotechnical engineering exploration.
[0003] Currently, feature extraction from core images obtained through drilling is performed using deep learning networks, and semantic segmentation techniques can be used to initially define the contours of defects or rock blocks. Chinese patent application CN117291871A discloses a deep learning-based method for core image crack identification and crack dip angle calculation. This method utilizes a semantic segmentation network to classify and label crack regions, core blocks, and fracture zones in horizontal core images. Then, it employs either a direct method for skeletalization or an indirect method to extract edge lines and generate crack skeleton lines. Based on these techniques, automated preliminary identification of core cracks can be achieved, and the interference of cores containing fracture zones on crack identification can be reduced to some extent.
[0004] However, the aforementioned technical solutions can improve the automation and processing efficiency of core sample defect identification to some extent. However, existing technologies have significant shortcomings in their identification mechanisms when dealing with complex, fragmented, and broken real soil core samples. Existing defect identification and region merging algorithms often rely on fixed similarity thresholds or single-dimensional appearance features for static merging and processing, failing to deeply integrate the continuity of boundary structures between adjacent regions and the dynamic changes in the network's confidence in defect judgment. Furthermore, in the merging verification and classification stages, existing methods typically lack dynamic and objective iterative termination mechanisms, relying solely on a fixed number of iterations or simple feature splicing for network fitting, without inputting rigorous statistical measures and probability distribution monitoring. This mechanical processing mechanism easily leads to fragmented primitives belonging to the same complete defect being separated into isolated small defects, or to incorrectly merging regions that are essentially different but superficially similar. The lack of rigorous statistical measures makes the model's classification boundaries for complex defects in the feature space extremely blurry, unable to effectively identify and eliminate false positives and spurious defects. Ultimately, this results in insufficient stability and low overall accuracy of the soil sample defect identification algorithm, making it unreliable for use in complex real-world engineering geological scenarios. Summary of the Invention
[0005] To address the technical problem of low accuracy in soil sample defect identification caused by static merging and lack of dynamic termination mechanisms in existing technologies, this invention proposes an intelligent identification method for drilling and sampling soil defects based on machine vision.
[0006] This invention provides a machine vision-based intelligent identification method for defects in drilled soil samples, comprising: S1, acquiring soil core image, extracting initial defect primitives through a dual-branch convolutional neural network, the first branch outputting the segmentation mask of each primitive, and the second branch outputting the feature vector and the initial defect confidence; S2, initializing the iterative merging process, calculating the merging score with weighted terms for spatially adjacent candidate regions in each iteration, the weights of which are determined based on the harmonic mean of the initial defect confidence of the two regions, and the weights including the similarity of the feature vectors of the two regions, the structural feature continuity of the contact boundary, and the confidence obtained by re-inputting the confidence-weighted average of the feature vectors of the two regions into the classification head of the second branch of the dual-branch network. Gain: Based on the merge score, a normalized merge probability distribution is generated using the Softmax function. Candidate regions are randomly sampled according to the merge probability distribution and merged to generate composite regions. The feature vector of the composite region is updated using a confidence-weighted average algorithm, and its fusion confidence and the topology of the undirected graph are updated simultaneously. S3: Repeat the iteration until the information entropy of the merge probability distribution is lower than the preset first entropy threshold. The updated feature vector is used to calculate the Mahalanobis distance between each composite region and each centroid in the feature space constructed by the preset defect category centroid and covariance matrix. When the minimum Mahalanobis distance is less than the second distance threshold, the category corresponding to the minimum Mahalanobis distance is confirmed as the defect category of the composite region and the defect location and category are output.
[0007] This invention overcomes the regional error adhesion and fragmentation defects caused by traditional static merging that relies on fixed thresholds by dynamically adjusting the weighting terms using the harmonic mean of the initial defect confidence during the merging stage and performing random sampling and merging based on the probability distribution generated by Softmax. At the same time, an objective iteration termination mechanism is constructed by calculating the information entropy of the merging probability distribution, enabling the network to dynamically perceive and autonomously judge the degree of merging, thereby significantly improving the accuracy of soil sample defect identification in complex scenarios.
[0008] Preferably, the step of extracting initial defect primitives from the image using a dual-branch convolutional neural network includes: scaling the soil core image to be identified to a preset pixel size and normalizing it; inputting the normalized image into the shared feature extraction backbone of the dual-branch convolutional neural network; generating a global feature map using five convolutional groups connected end-to-end; inputting all generated global feature maps into the first branch, which consists of five in-series deconvolutional layers and is used to output segmentation masks for each primitive; extracting the region of interest features of each primitive from the global feature map based on the segmentation masks, converting them into fixed dimensions through region of interest pooling, and then inputting them into the second branch, which consists of three in-series fully connected layers and is used to output feature vectors representing the essential attributes of each primitive and the initial defect confidence score.
[0009] This invention acquires multi-scale global features by using a shared backbone of five convolutional groups connected end to end, and uses independent first and second branches to output segmentation masks and high-dimensional feature vectors respectively, ensuring comprehensive capture of the apparent contours and essential attributes of defect primitives and improving the accuracy of initial defect primitive extraction.
[0010] Preferably, the calculation of the combined score composed of weighted terms includes: calculating the cosine similarity of the feature vectors of the two regions as a first sub-item; calculating the inverse gradient of the curvature change of the contact boundary between the two regions as the structural feature continuity of the contact boundary to obtain a second sub-item; obtaining the confidence gain as a third sub-item; normalizing the first, second, and third sub-items respectively; multiplying the first, second, and third sub-items by their corresponding weights and summing them to obtain the combined score.
[0011] Preferably, the weights of the weighted terms are determined based on the harmonic mean of the initial defect confidence scores of the two regions, including: obtaining the first initial defect confidence score of the first spatially adjacent candidate region and the second initial defect confidence score of the matched and associated second candidate region; applying a calculation formula to derive weight allocation parameters, and using the weight allocation parameters as the weights of the weighted terms in the calculation of the combined score; the calculation formula is:
[0012] in, Assign parameters to the weights; This is the adjustment coefficient; The initial defect confidence level; The second initial defect confidence level; This is the fault tolerance constant.
[0013] This invention effectively avoids extreme value deviation caused by simple averaging by introducing a harmonic mean mechanism in the weight allocation to fuse the confidence levels of the two regions. At the same time, it sets a fault tolerance constant to prevent the operation from crashing when the denominator is zero, thereby enhancing the stability of the merging bias determination.
[0014] Preferably, the confidence gain is obtained by calculating the merged new feature vector using a fusion formula; the fusion formula is:
[0015] in, This is the new feature vector; The initial defect confidence level; The feature vector of the first candidate region; The second initial defect confidence level; The feature vector of the second candidate region; The new feature vector is input into the classification head of the second branch, and the new classification confidence is obtained through mapping calculation. The new classification confidence is then subtracted. and The arithmetic mean of the values is used to calculate the confidence gain, and the difference obtained from the calculation is designated as the confidence gain.
[0016] Preferably, the step of randomly sampling candidate region pairs according to the merging probability distribution to generate a composite region includes: constructing a continuous probability distribution interval based on the normalized merging probabilities of all candidate region pairs globally, and generating a random floating-point number within the interval; determining the specific candidate region pairs corresponding to the sub-intervals into which the random floating-point number falls as the region pairs to be merged, and clearing the boundary pixel coordinate distinctions at the adjacent boundaries of the internal space of the region pairs to be merged; and merging the region pairs to be merged to generate a complete single composite region tile mask with a single label of the same category and eliminating internal intersecting lines.
[0017] Preferably, the step of generating a normalized merging probability distribution based on the merging score using a Softmax function includes: calculating the normalized merging probability of each candidate region pair using a Softmax function that incorporates a temperature-adjusting hyperparameter; the calculation formula used is:
[0018] in, For the first Normalized merging probability of each candidate region pair; For the first The combined score of each candidate region pair; This refers to the temperature regulation hyperparameter. For the first The combined score of each candidate region pair; This represents the number of spatially adjacent candidate region pairs.
[0019] This invention introduces a temperature regulation hyperparameter into the Softmax function for calculating the normalized merging probability, which can effectively control the overflow of the exponential term caused by the absolute value of high potential energy, making the merging probability distribution smoother and preventing excessive polarization during the sampling and merging process.
[0020] Preferably, terminating the iteration when the information entropy of the merged probability distribution is lower than a preset first entropy threshold includes: calculating the information entropy representing the current distribution state based on the Shannon entropy principle using the information entropy formula; the method for calculating the information entropy is as follows:
[0021] in, The information entropy represents the current distribution state; This represents the total number of candidate region pairs; For the first Normalized merging probability of each candidate region pair; It is a constant with a very small offset; The system compares and determines whether the information entropy is strictly lower than a preset first entropy threshold. If the determination condition is met, the current iterative merging operation process is stopped.
[0022] This invention completely eliminates the risk of throwing non-numerical anomalies due to the probability of the input being zero when the input is minimized by pre-adding a minimal offset constant to the calculation formula of the Shannon entropy principle, thus ensuring the continuity and reliability of the objective iteration termination judgment mechanism.
[0023] Preferably, the calculation of the Mahalanobis distance between each composite region and each centroid is performed using the following method:
[0024] in, For absolute Mahalanobis distance variables; The one-dimensional feature difference row vector is obtained by subtracting the preset centroid feature vector of a specific defect category from the updated feature vector of the composite region. This is the pseudo-inverse matrix of the covariance matrix pre-calculated for each defect category; It is a column vector of differences.
[0025] Preferably, before initializing the iterative merging process, the method further includes: constructing an undirected graph data structure containing all initial defect primitives, where nodes represent initial defect primitives and edges represent spatial adjacency relationships; calling the kd-tree algorithm to determine the intersection of the segmentation mask boundaries of each primitive, and establishing candidate region pairs in a spatially adjacent state.
[0026] The beneficial effects of this invention are as follows: This invention utilizes a dual-branch convolutional neural network to extract segmentation masks, feature vectors, and confidence scores of defect primitives in images. It constructs a merging scoring mechanism that integrates feature similarity, structural continuity, and confidence gain. Weighting terms are assigned based on the harmonic mean of the initial confidence score. Combined with a probability distribution-based sampling strategy and an objective termination condition based on information entropy, this avoids misjudgments and omissions caused by traditional rigid merging methods, ensuring the scientific validity of composite region generation. By calculating Mahalanobis distance within the feature space to determine defect categories, it eliminates interference from correlations between feature variables, improving the accuracy of intelligent identification and location of defects in drilled soil samples.
[0027] Furthermore, this invention employs a shared feature extraction backbone comprising five convolutional groups connected end-to-end to acquire multi-scale global features. Independent first and second branches output segmentation masks and high-dimensional feature vectors respectively, ensuring comprehensive capture of the apparent contours and essential attributes of defect primitives and improving the accuracy of initial defect primitive extraction. A harmonic mean mechanism is introduced in weight allocation to fuse the confidence levels of two regions, avoiding extreme value bias caused by simple averaging. A fault tolerance constant is set to prevent computational crashes caused by a zero denominator, enhancing the stability of the merging bias judgment. A temperature-regulating hyperparameter is introduced into the Softmax function for calculating the normalized merging probability, controlling the exponential term overflow caused by high potential energy absolute values, resulting in a smoother merging probability distribution and preventing excessive polarization during the sampling and merging process. A minimal offset constant is pre-stacked in the Shannon entropy principle calculation formula, eliminating the risk of throwing non-numerical anomalies due to zero input caused by probability minimization, ensuring the continuity and reliability of the objective iteration termination judgment mechanism. Attached Figure Description
[0028] Figure 1 This is a flowchart of a machine vision-based intelligent identification method for defects in drilling and sampling soil samples, as described in this invention. Figure 2 This is a schematic diagram illustrating the information entropy change trend during the merging iteration process in this invention; Figure 3 This is a schematic diagram of feature distance distribution verification in this invention; Figure 4 This is a schematic diagram comparing ablation experiments in this invention. Detailed Implementation
[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0030] This invention discloses a machine vision-based intelligent identification method for defects in drilled soil samples, referring to... Figure 1 This includes steps S1-S3: S1. Acquire the image and extract the defect primitive features.
[0031] In an optional embodiment, an image of the soil core sample to be identified is obtained, and initial defect primitives in the image are extracted by a two-branch convolutional neural network. The first branch outputs the segmentation mask of each primitive, and the second branch outputs the feature vector representing the essential attributes of each primitive and the initial defect confidence.
[0032] Specifically, high-resolution color digital images of the drilling sampling site are read and converted into arrays in the RGB color space. Bilinear interpolation is then used to uniformly scale the images to a fixed resolution. The preprocessed image tensors are input into a two-branch convolutional neural network built on the PyTorch framework. The backbone of this network uses a ResNet50 structure to extract multi-scale feature maps. The first branch uses a feature pyramid network structure combined with a fully convolutional network to upsample and decode the feature maps, calculating the binarized segmentation mask for each initial defect primitive pixel by calculating the output of the Sigmoid activation function. The second branch compresses the multi-scale feature maps into one-dimensional vectors through global average pooling. This vector is then dimensionality-reduced through a fully connected layer containing a ReLU activation function to generate high-dimensional feature vectors representing the essential attributes of each primitive. These high-dimensional feature vectors are then connected in parallel to a multilayer perceptron network containing independent linear layers. After mapping to a scalar, the output of the Sigmoid function is the initial defect confidence score between 0 and 1.
[0033] The soil core image to be identified is scaled to 256×256 pixels and normalized to a pixel value of 0 to 1. The normalized image is then input into the shared feature extraction backbone of a two-branch convolutional neural network. Five convolutional groups connected end-to-end are used to generate a global feature map. Each convolutional group contains a 3×3 two-dimensional convolutional layer with a stride of 1 and a 2×2 two-dimensional max pooling layer with a stride of 2. All generated global feature maps are input into the first branch, which consists of five cascaded deconvolutional layers and is used to output the segmentation mask for each primitive. Based on the segmentation mask, the region of interest (ROI) features of each primitive are extracted from the global feature map and converted to a fixed dimension by ROI pooling before being input into the second branch. The second branch consists of three cascaded fully connected layers and is used to output the feature vector representing the essential attributes of each primitive and the initial defect confidence.
[0034] The structure of the dual-branch convolutional neural network includes a shared feature extraction backbone, a first branch, and a second branch. The network input is a normalized image, and the network output consists of segmentation masks for each primitive, feature vectors representing the essential attributes of each primitive, and initial defect confidence scores. A bilinear interpolation algorithm is used to scale the acquired high-resolution soil core images to a standard resolution of 256×256×3. A single-precision floating-point normalization mapping within the range of 0 to 1 is achieved through a division operation of 255 to eliminate the influence of numerical dimensions caused by differences in illumination intensity. The normalized soil core images are input into the shared feature extraction backbone constructed based on ResNet. The channel number variables in the five convolutional groups connected end-to-end are sequentially configured as 32, 64, 128, 256, and 512. After each 3×3 convolution feature extraction and 2×2 max pooling downsampling, the spatial resolution is halved layer by layer, generating a high-order global feature map with a spatial size of 8×8 and 512 channels. For the first branch, the stride variable for each of the five concatenated deconvolutional layers (i.e., transposed convolutional layers) is set to 2, and the convolutional kernel is set to 4×4. After layer-by-layer upsampling, the feature map is restored back to a spatial resolution of 256×256, and a single-channel binarized primitive segmentation mask is generated using the Sigmoid activation function. Region-of-interest pooling is performed using the boundary coordinates generated by the mask, cropping multi-dimensional features from the corresponding coordinate regions output by the feature extraction backbone, and uniformly scaling them to a fixed-dimensional feature block of 7×7 pixels. The 7×7×512 three-dimensional feature tensor is flattened and input into the second branch. The neuron node parameters of the three fully connected layers are configured with 1024, 512, and 128 dimensions, respectively, achieving dimensionality reduction from high-dimensional features to a low-dimensional semantic space. The linear classification head at the end outputs a feature vector with 128-dimensional representation and a single-channel initial defect confidence scalar value mapped by the Sigmoid function.
[0035] S2. Iteratively calculate and merge scores to generate composite regions.
[0036] In an optional embodiment, an iterative merging process is initialized. In each iteration, a merging score consisting of weighted terms is calculated for spatially adjacent candidate region pairs. The weights of the weighted terms are determined based on the harmonic mean of the initial defect confidence of the two regions. The weighted terms include the similarity of the feature vectors of the two regions, the structural feature continuity of the contact boundary, and the confidence gain obtained by re-inputting the confidence-weighted average of the feature vectors of the two regions into the second branch classification head of the dual-branch network. Based on the merging score, a normalized merging probability distribution is generated through the Softmax function, and region pairs are randomly sampled according to the probability distribution to perform merging and generate composite regions. The feature vectors are updated using a confidence-weighted averaging algorithm, and the confidence of the new region is updated to the fusion confidence while updating the topology of the undirected graph.
[0037] Specifically, an undirected graph data structure containing all initial defect primitives is constructed, where nodes represent primitives and edges represent spatial adjacency relationships. The kd-tree algorithm from the SciPy library is used to determine the intersection of mask boundaries to establish candidate pairs of spatially adjacent regions. For each pair of candidate regions, the corresponding initial defect confidence is obtained, and the harmonic mean is calculated by multiplying the product of two confidences by the sum of the confidences. The harmonic mean is then multiplied by the initial base weight coefficient to obtain the current weight of the weighting term. The merged score includes three weighted sub-items. The first sub-item uses the linear algebra module from the NumPy library to calculate the dot product of the feature vectors of the two regions and divides it by the product of their respective L2 norms to obtain the cosine similarity. The second sub-item uses the find-Contours function from the OpenCV library to extract the contact boundary between the two regions and calculates the inverse of the gradient of the curvature change as a measure of structural feature continuity. The third sub-item adds the feature vectors of the two regions according to their confidence ratios and performs a weighted average to obtain the fused vector. The forward propagation interface of the deep learning framework is invoked to input the fusion vector into the classification head of the second branch, recalculating the fusion confidence score and subtracting the average of the initial defect confidence scores of the two regions to obtain the confidence gain. After normalizing the first, second, and third sub-items, they are multiplied by their corresponding weights and summed to obtain the merge score. The merge scores of all adjacent region pairs are grouped into a one-dimensional array and input into the Softmax activation function for exponential operation and normalization, transforming it into a merge probability distribution where the sum of all elements is 1. The `choice` function in the `random` module of the NumPy library is called to randomly select a region pair according to the probability distribution using a roulette wheel algorithm, performing a Boolean OR operation on the mask image to merge spatial regions. Simultaneously, while generating a new composite region, the feature vector of the new region is updated to the previously calculated confidence-weighted average fusion vector, and the confidence score of the new region is updated to the fusion confidence score, while also updating the topology of the undirected graph.
[0038] In an optional embodiment, in each iteration, a combined score consisting of weighted terms is calculated for spatially adjacent candidate region pairs, the weights of which are determined based on the harmonic mean of the initial defect confidences of the two regions.
[0039] Specifically, obtain the first initial defect confidence of the first spatially adjacent candidate region. And the second initial defect confidence of the second candidate region matched and associated. To fully consider the combined influence of the confidence levels of the two regions in the evaluation and avoid extreme value bias caused by simple averaging, a weight allocation parameter is constructed based on the harmonic mean. An adjustment coefficient is introduced to adapt to the model's weight distribution requirements, and a fault tolerance constant is incorporated to prevent the denominator from being zero, thus deriving the weight allocation parameter. And assign weight parameters The generated weight values, used as weighting factors, are applied to the calculation of the combined score. The calculation method for the weight allocation parameters is as follows:
[0040] in, Assign parameters to the weights; This is the adjustment coefficient; The initial defect confidence level; The second initial defect confidence level; This is the fault tolerance constant.
[0041] In calculating the weight allocation parameters, a floating-point confidence variable is read from the end of the second branch of the two-branch network. and Due to the limitations of the network activation function, the values are all distributed within the open interval of 0 to 1. For example, when two adjacent regions both exhibit high-probability defect features, the extracted values are... It is 0.85. The value is 0.75. The harmonic mean is calculated to be approximately 0.7968. Multiplying this harmonic mean by the preset amplification factor of 1.5 yields the generated weight value. The value is approximately 1.195, thus assigning a higher weight to high-confidence regions in calculating the pooled score due to pooling bias. Adjustment coefficient. The value of is greater than 0, and preferably 1.5×2, where 2 comes from the harmonic mean formula and 1.5 is the confidence level adjustment amplification factor.
[0042] It should be noted that, in order to improve operational stability and prevent program crashes caused by a denominator of zero at the code level, a small fault-tolerance constant was input into the division operation. Its preferred value is set to 1e-6. The calculation module incorporates a boundary threshold constraint mechanism; if the calculated weight allocation parameters... If the value exceeds the set saturation limit of 1.4, the Clip function will be called to forcibly clamp the value to 1.4. This prevents extreme abnormal confidence combinations from causing a surge in potential energy values in local areas and ensures a stable probability distribution during the overall spatial feature reconstruction process.
[0043] In an optional embodiment, the confidence gain is obtained by re-inputting the confidence-weighted average of the feature vectors from the two regions into the classification head of the second branch of the dual-branch network.
[0044] Specifically, extract the feature vector of the first candidate region. and confidence level Extract the feature vector of the second candidate region. and confidence level To preserve the attributes of the dominant features with higher confidence during feature fusion, a formula for constructing the merged new feature vector is built based on the weighted mean:
[0045] in, This is the new feature vector; The confidence level of the first candidate region; The feature vector of the first candidate region; The confidence level of the second candidate region; This is the feature vector of the second candidate region.
[0046] The new feature vector is input into the classification head of the second branch, and the new classification confidence is obtained through mapping calculation. The new classification confidence is then subtracted. and The arithmetic mean of the values is used to calculate the confidence gain, and the difference obtained from the calculation is designated as the confidence gain.
[0047] The eigenvector blending step is implemented based on element-wise matrix calculation using the underlying tensor API, extracting eigenvectors. With feature vectors All are single-precision floating-point arrays of length 128, and the scalar confidence scores are... and As a weighting parameter, it performs scalar multiplication broadcasts with each of the 128 components in the array. For example, it sets the confidence level to be obtained. It is 0.8. The denominator is 0.6, and the sum of the denominators is 1.4. Substituting this into the formula yields the new eigenvector. This ensures that regions with more pronounced defective attributes dominate the direction of the feature space vector during feature fusion, maximizing the inheritance of strong morphological primitive attributes by the fused region. After fusion, the new 128-dimensional feature vector... Bypassing the backbone extraction layer of the network, the data is fed into the second branch's end multilayer perceptron (classification head), which is already in a fixed-weight state and has its gradient backpropagation blocked. A forward inference calculation is performed to obtain a new classification confidence level between 0 and 1. Assuming the predicted classification confidence level is 0.88, and the arithmetic mean of the original features is 0.7, a subtraction operation yields a difference of 0.18, which is then assigned to the confidence gain variable. When the confidence gain variable is greater than 0, it verifies from a high-dimensional feature level that the current fusion helps improve the model's confidence in the overall continuity of the defect.
[0048] In an optional embodiment, a normalized merging probability distribution is generated based on the merging score using a Softmax function, and merging is performed on randomly sampled regions according to the probability distribution to generate a composite region.
[0049] Specifically, the Softmax function is applied to the merge scores of all spatially adjacent candidate region pairs calculated independently to obtain the normalized merge probability of each region pair between 0 and 1. A continuous probability distribution interval is constructed based on the normalized merge probabilities of all global region pairs, and a random floating-point number is generated within the interval. The specific candidate region pairs corresponding to the sub-intervals in which the random floating-point number falls are determined as the region pairs to be merged. The boundary pixel coordinates at the spatial adjacency boundaries of the region pairs to be merged are cleared, and the region pairs to be merged are merged to generate a complete single composite region tile mask with a single label of the same category and no internal intersecting lines.
[0050] When constructing a continuous probability distribution interval, extract Each spatially adjacent candidate region is paired with the calculated merged score variable array. , , , To prevent excessive potential energy from causing exponential term overflow during Softmax calculations, a temperature-adjusting hyperparameter is introduced into the Softmax function to construct a normalization formula:
[0051] in, For the first Normalized merging probability of each region pair; For the first The combined score of each region pair; The preferred range for the temperature regulation hyperparameter is 1 to 2, and in this invention it is 1.5; For the first The combined score of each region pair; This represents the number of spatially adjacent candidate region pairs.
[0052] For example, the temperature-adjusted potential energy calculation results for three regions in a certain iteration. The values are mapped to 0.5, 0.3, and 0.2 respectively, thus generating contiguous cumulative probability sub-intervals of 0 to 0.5, 0.5 to 0.8, and 0.8 to 1 in memory. The Mersenne-Twister random algorithm is invoked to generate a random floating-point variable in the range of 0 to 1, such as 0.65. If the floating-point number falls within the second sub-interval, a memory address call instruction is issued, locking the second candidate region for entity graph theory merging. In the entity topology merging stage of the image spatial mask, a bitwise OR operation is performed on the set of Boolean pixels of the two selected regions within the 256×256 two-dimensional mask matrix to achieve a union. To remove pixel-level jagged edges and microscopic voids that may remain at the original adjacent seams due to mesh discretization, morphological closure is enforced, and the rectangular structuring element kernel size parameter used for dilation and erosion is configured to 3×3. After processing, a unique new integer ID label is assigned to the newly fused largest connected component.
[0053] S3. Iterate until the information entropy meets the condition and output the result.
[0054] In an optional embodiment, the iteration is repeated until the information entropy of the merged probability distribution is lower than a preset first entropy threshold, at which point the iteration is terminated. For each composite region formed, the Mahalanobis distance between the composite region and each centroid is calculated using the feature vector in the feature space constructed by the preset defect category centroid and covariance matrix. When the minimum Mahalanobis distance is less than a second distance threshold, the category corresponding to the minimum Mahalanobis distance is confirmed as the defect category of the composite region, and the location and category are output.
[0055] Specifically, a loop control structure is set up to continuously repeat the candidate region merging process from the previous steps. After each merging action, the `entropy` function of the SciPy library is called to calculate the information entropy of the current network topology merging probability distribution by multiplying the negative probability by the probability logarithm formula. The calculated information entropy is compared with a set first entropy threshold floating-point number. When the information entropy value drops below the first entropy threshold, the loop termination condition is triggered to end the entire merging iteration process. The feature vector centroids of various known defect samples, which are pre-calculated and saved based on historical manually labeled data using the expectation-maximization algorithm clustering, and the inverse matrix of the corresponding feature space covariance matrix are read. For each generated composite region, the updated feature vector is read, and the Mahalanobis distance between the feature vector and the centroids of each preset defect category in the corresponding covariance inverse matrix space is calculated. The minimum value in the Mahalanobis distance array is searched, and the minimum value is compared with a preset second distance threshold. If the value of the minimum value is strictly less than the second distance threshold, the category index corresponding to the minimum value is assigned to the current composite region as the confirmed defect category through dictionary mapping. The OpenCV library's boundingRect function is called to calculate the pixel coordinates of the minimum bounding rectangle based on the binary segmentation mask of the composite region. The bounding rectangle coordinates and the corresponding text-formatted defect category name are then output to the external device in JSON data format.
[0056] The iteration termination condition is jointly limited by information entropy, maximum merge score, maximum confidence gain, and the number of candidate region pairs. Specifically, in each iteration, region merging is sampled and performed one or more times based on the merging probability distribution, updating the current composite region set and the set of spatially adjacent candidate region pairs; the updated candidate region pair merging scores and merging probability distributions are recalculated, and it is determined whether the termination condition is met. The termination condition includes at least one of the following: the current number of candidate region pairs is zero, the maximum merge score of all current candidate region pairs is lower than a preset merge score threshold, the maximum confidence gain of all current candidate region pairs is lower than a preset gain threshold, etc.
[0057] In an optional embodiment, the iteration terminates when the information entropy of the merged probability distribution falls below a preset first entropy threshold.
[0058] Specifically, obtain the normalized merging probability set of all candidate region pairs in the current iteration, where the th... The probability of merging a pair of regions is , The total number of candidate region pairs. Based on Shannon entropy in information theory, and introducing a minimal offset constant to prevent non-numerical anomalies caused by zero input, an information entropy representing the current distribution state is constructed, calculated item by item, and accumulated:
[0059] in, The information entropy represents the current distribution state; This represents the total number of candidate region pairs; For the first The probability of merging a pair of regions; It is a constant with a very small offset.
[0060] Comparison and judgment of information entropy If the threshold is strictly lower than the preset first entropy threshold, and the condition is met, a termination command is sent to the system to stop the current iterative merging process, and the results of each composite region are returned and output.
[0061] In the calculation of evaluating the disorder of the system distribution, as the merging iteration progresses, the candidate regions affect the total number variable. It will exhibit a monotonically decreasing trend, with a few regions possessing extremely high potential energy dominating the probability distribution. To prevent non-numerical anomalies from being thrown due to zero input in floating-point operations, each extracted normalized probability variable... Pre-stacked minimal offset constant For example, 1e-9. When the iteration reaches the later stages, the number of remaining candidate region pairs on the graph... The value decreases to 3, and the normalized merging probability polarization distribution is: 0.85 0.1 The value is 0.05. The processor substitutes each term into the formula and performs the accumulation. After binary logarithmic conversion and summation, the current information entropy is obtained. The first entropy threshold is 0.72. It is preferably limited to between 0.5 and 1 to find a balance between avoiding excessive global merging of regions and preventing fragmentation of identification. Assuming the first entropy threshold loaded in the configuration file is 0.8, since the real-time calculated entropy of 0.72 satisfies the condition of being less than 0.8, the main control process will respond immediately and forcibly block the current While loop through a built-in interrupt function.
[0062] Reference Figure 2 This presents the information entropy during the merging and iteration process. As the number of remaining candidate logs decreases, the broken line shows a gradual decline from left to right, eventually crossing the preset threshold. This indicates that as the merging process progresses, the uncertainty of the distribution gradually decreases, and the information entropy... When the temperature drops to a preset threshold, the iteration process can automatically terminate, effectively avoiding excessive merging or fragmentation of regions.
[0063] Freeze all composite region mask matrices and their corresponding 128-dimensional feature vectors to complete the data encapsulation of entity feature results, and transfer them to the end category comparison mapping output stage.
[0064] In an optional embodiment, the Mahalanobis distance between the formed composite regions and each centroid is calculated using feature vectors in a feature space constructed from preset defect category centroids and covariance matrices.
[0065] Specifically, the preset centroid feature vectors of each defect category are obtained from the feature library of typical defect categories, and the inverse matrix of the covariance matrix pre-calculated for each defect category is called. For each composite region, the feature vector of the composite region is obtained, and the preset centroid feature vector of the specific defect category is subtracted from the feature vector to obtain a one-dimensional feature difference row vector. The one-dimensional feature difference row vector, the inverse matrix of the covariance matrix, and the transpose column vector of the one-dimensional feature difference row vector are successively multiplied by continuous matrix multiplication, and the square root of the scalar result obtained after multiplication is calculated to obtain the Mahalanobis distance between the composite region and the centroid of the specific defect category.
[0066] The implementation of the feature distance metric relies on offline data statistics and pre-loaded models. In the early offline training phase, parameters were extracted from the features of hundreds of thousands of labeled samples for each specific soil sample defect category, such as cracks, pores, and inclusions. The arithmetic mean of the features for each category was calculated, and the mean was serialized and saved as a 1×128 dimension preset centroid feature vector, denoted as vector. Simultaneously, the 128×128-dimensional covariance matrix for each category is calculated based on the sample variance distribution. To address the potential issues of high collinearity and ill-conditioned matrix invertibility arising from high-dimensional feature data, a singular value decomposition algorithm is used to pre-calculate the pseudo-inverse matrix of the covariance matrix, denoted as . This data is then cached in the device's memory data dictionary. During the inference phase, a high-dimensional representation feature of a specific composite region is extracted and assigned to a 1×128 dimension vector. The centroid vector of a specific category, such as cracks. Perform element-wise subtraction at the bottom layer to generate a one-dimensional feature difference row vector, also with dimensions of 1×128. The formula for calculating the absolute Mahalanobis distance variable is constructed based on Mahalanobis distance in multivariate statistical analysis:
[0067] in, For absolute Mahalanobis distance variables; This is a one-dimensional feature difference row vector; It is a pseudo-inverse matrix; It is a column vector of differences.
[0068] When performing tensor multiplication, prioritize processing the difference row vectors. With a 128×128 dimension pseudo-inverse matrix The multiplication outputs a 1×128-dimensional intermediate row vector containing scale weights and orientation correlation, which is then multiplied by the 128×1-dimensional difference column vector generated by matrix transpose. Perform a dot product operation to obtain a floating-point squared Mahalanobis distance scalar representing the spatial offset, and then use the square root function to obtain the absolute Mahalanobis distance variable. For example, for the same composite region, calculate the distance between the composite region and the hole type. The distance to the mixed category is 2.85. The value is 7.4, which is the distance from the crack category. The value is 8.1. With the second distance threshold variable preferably ranging from 4 to 6, for example, the second distance threshold variable is preset to 5, acting as a confidence boundary filter. Since 2.85 is the minimum value and less than the preset distance threshold of 5, the logic judge immediately issues an instruction to confirm the composite area as a hole category, and packages the mask outline and positioning coordinates and outputs them to the front-end interface.
[0069] Reference Figure 3 This study demonstrates the distribution of distances from feature vectors belonging to different defect categories to the centroids of each category. Within each defect category, the Mahalanobis distance from the feature vector to its own category centroid is significantly lower than the distance cutoff line, while the Mahalanobis distances from the feature vector to the centroids of other categories are significantly higher than the distance cutoff line. This significant differentiation in distance values proves that the feature vectors of cracks, pores, and inclusions are significantly less distant from their own category centroids than from the centroids of other categories. By using a pre-set specific distance threshold, accurate identification and classification of different defect categories can be achieved.
[0070] The experiment used the PyTorch deep learning framework for model construction and network parameter training. The dataset consisted of 10,000 high-resolution soil core images, covering three typical defect categories: cracks, pores, and inclusions. 8,000 images were used for model training, and 2,000 were used for model validation and testing. All images were uniformly scaled to 256×256 pixels and normalized to pixel values between 0 and 1. The network optimizer employed the AdamW algorithm, with an initial learning rate of 0.001 and a batch size of 16. A total of 100 training generations of forward inference and backpropagation were performed.
[0071] The baseline model employs an arithmetic mean weighting of confidence scores, a fixed number of adjacent region iterative mergings, and a classification method based on Euclidean distance. The average crossover union (CUB) for defect identification is 76.4%, the overall classification accuracy is 81.2%, and the average processing time per image is 45ms. By inputting a weighting strategy based on the harmonic mean of the initial defect confidence scores of two regions into the baseline model, the average CUB increases to 81.7%, and the classification accuracy reaches 85.6%. Further adding a termination merging mechanism based on information entropy distribution increases the average CUB to 84.3%, and the processing time per image decreases to 32ms. After deploying a complete scheme including a high-dimensional feature space Mahalanobis distance metric, the average CUB on the test set reaches 89.1%, the overall classification accuracy climbs to 94.5%, and the average processing time stabilizes at 34ms.
[0072] Reference Figure 4 The evolution of the model from the baseline model through different ablation models to the complete solution reveals the changing trends of relevant performance indicators and processing time overhead. The average intersection-over-union ratio (AUC) and overall classification accuracy show a progressively increasing trend in each evolution stage, while the processing time overhead experienced a trajectory of initial slight increase, followed by a significant decrease, and finally stabilizing at a low level. With the sequential addition of harmonic mean weighting, information entropy termination strategy, and Mahalanobis distance metric, the overall recognition accuracy of the model has steadily improved, while the processing time overhead of the algorithm has been effectively reduced and controlled, fully verifying the positive contribution of each mechanism module to the overall system performance.
[0073] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A machine vision-based intelligent identification method for defects in drilled soil samples, characterized in that, include: S1. Obtain soil core image and extract initial defect primitives using a two-branch convolutional neural network. The first branch outputs the segmentation mask for each primitive, and the second branch outputs the feature vector and initial defect confidence. S2. Construct an undirected graph data structure containing all initial defect primitives, where nodes represent initial defect primitives and edges represent spatial adjacency relationships. Initialize the iterative merging process. In each iteration, calculate the merging score with weighted terms for spatially adjacent candidate regions. The weights are determined based on the harmonic mean of the initial defect confidence of the two regions. The weights include the similarity of the feature vectors of the two regions, the structural continuity of the contact boundary, and the result of re-inputting the confidence-weighted average of the feature vectors of the two regions into the second branch classification head of the two-branch network. The confidence gain is calculated based on the merge score. A normalized merge probability distribution is generated using the Softmax function. Candidate regions are randomly sampled according to the merge probability distribution and merged to generate composite regions. The feature vector of the composite region is updated using a confidence-weighted average algorithm, and its fusion confidence and the topology of the undirected graph are updated simultaneously. S3: The iteration is repeated until the information entropy of the merge probability distribution is lower than the preset first entropy threshold. The updated feature vector is used to calculate the Mahalanobis distance between each composite region and each centroid in the feature space constructed by the preset defect category centroid and covariance matrix. When the minimum Mahalanobis distance is less than the second distance threshold, the category corresponding to the minimum Mahalanobis distance is confirmed as the defect category of the composite region and the defect location and category are output.
2. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The step of extracting initial defect primitives from an image using a dual-branch convolutional neural network includes: scaling the soil core image to be identified to a preset pixel size and normalizing it; inputting the normalized image into the shared feature extraction backbone of the dual-branch convolutional neural network; generating a global feature map using five convolutional groups connected end-to-end; inputting all generated global feature maps into the first branch, which consists of five cascaded deconvolutional layers and is used to output a segmentation mask for each primitive; extracting the region of interest features of each primitive from the global feature map based on the segmentation mask, converting them to a fixed dimension through region of interest pooling, and then inputting them into the second branch, which consists of three cascaded fully connected layers and is used to output feature vectors representing the essential attributes of each primitive and the initial defect confidence score.
3. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The calculation of the combined score, which consists of weighted terms, includes: calculating the cosine similarity of the feature vectors of the two regions as the first sub-term; calculating the inverse gradient of the curvature change of the contact boundary between the two regions to obtain the structural feature continuity of the contact boundary, and taking the structural feature continuity as the second sub-term; obtaining the confidence gain as the third sub-term; normalizing the first, second, and third sub-terms respectively; multiplying the first, second, and third sub-terms by their corresponding weights and summing them to obtain the combined score.
4. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The weights of the weighted terms are determined based on the harmonic mean of the initial defect confidence scores of the two regions, including: obtaining the first initial defect confidence score of the first spatially adjacent candidate region and the second initial defect confidence score of the matched and associated second candidate region; applying a calculation formula to derive weight allocation parameters, and using these weight allocation parameters as the weights of the weighted terms in the calculation of the combined score; the calculation formula is: in, Assign parameters to the weights; This is the adjustment coefficient; The initial defect confidence level; The second initial defect confidence level; This is the fault tolerance constant.
5. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The confidence gain is obtained by applying a fusion formula to calculate the merged new feature vector; the fusion formula is: in, This is the new feature vector; The initial defect confidence level; The feature vector of the first candidate region; The second initial defect confidence level; The feature vector of the second candidate region; The new feature vector is input into the classification head of the second branch, and the new classification confidence is obtained through mapping calculation. The new classification confidence is then subtracted. and The arithmetic mean of the values is used to calculate the confidence gain, and the difference obtained from the calculation is designated as the confidence gain.
6. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The step of randomly sampling candidate region pairs according to the merging probability distribution and merging them to generate composite regions includes: constructing a continuous probability distribution interval based on the normalized merging probability of all candidate region pairs globally, and generating a random floating-point number within the interval; determining the specific candidate region pairs corresponding to the sub-intervals into which the random floating-point number falls as the region pairs to be merged, and clearing the boundary pixel coordinate distinctions at the adjacent boundaries of the internal space of the region pairs to be merged; and merging the region pairs to be merged to generate a complete single composite region tile mask with a single label of the same category and eliminating internal intersecting lines.
7. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The process of generating a normalized merging probability distribution based on the merging score using the Softmax function includes: calculating the normalized merging probability of each candidate region pair using the Softmax function with an introduced temperature-adjusting hyperparameter; the calculation formula used is: in, For the first Normalized merging probability of each candidate region pair; For the first The combined score of each candidate region pair; This refers to the temperature regulation hyperparameter. For the first The combined score of each candidate region pair; This represents the number of spatially adjacent candidate region pairs.
8. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The iteration terminates when the information entropy of the merged probability distribution falls below a preset first entropy threshold, including: calculating the information entropy representing the current distribution state based on the Shannon entropy principle using the information entropy formula; the method for calculating the information entropy is as follows: in, The information entropy represents the current distribution state; This represents the total number of candidate region pairs; For the first Normalized merging probability of each candidate region pair; It is a constant with a very small offset; The system compares and determines whether the information entropy is strictly lower than a preset first entropy threshold. If the determination condition is met, the current iterative merging operation process is stopped.
9. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, The calculation of the Mahalanobis distance between each composite region and each centroid is as follows: in, For absolute Mahalanobis distance variables; The one-dimensional feature difference row vector is obtained by subtracting the preset centroid feature vector of a specific defect category from the updated feature vector of the composite region. This is the pseudo-inverse matrix of the covariance matrix pre-calculated for each defect category; It is a column vector of differences.
10. The intelligent identification method for defects in drilling and sampling soil samples based on machine vision according to claim 1, characterized in that, Before initializing the iterative merging process, the method further includes: calling the kd-tree algorithm to determine the intersection of the segmentation mask boundaries of each primitive and establishing candidate region pairs in a spatially adjacent state.
Citation Information
Patent Citations
Rock core image crack identification and crack inclination angle calculation method based on deep learning
CN117291871A
Surface defect detection method based on deep learning and global difference information
CN116228754A
Infrared thermal imaging building facade defect intelligent diagnosis method based on multi-modal fusion
CN120635610A