Method and device for identifying anomalies in rock mass based on large model

Through a large-scale model-based method for identifying anomalies in rock masses, automatic detection and early warning of cracks and fragments in video data are achieved, solving the problem of traditional manual inspections being time-consuming, labor-intensive and inaccurate, and improving detection accuracy and timeliness.

CN120564111BActive Publication Date: 2025-09-30CHINA COAL RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511064488.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-30
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional crack and fragment detection methods rely on manual inspections, which are time-consuming and labor-intensive, and lack accuracy and timeliness. Existing video detection methods have low recognition accuracy in complex environments and lack effective hazard assessment.

Method used

A large-scale model-based rock mass anomaly recognition method is adopted to automatically detect and warn of cracks and fragments in video data through video acquisition, random sampling, mask prediction, mask splicing and feature extraction.

Benefits of technology

It improves the accuracy and timeliness of detection, reduces the difficulty and danger of manual inspections, and can accurately identify cracks and fragments in complex environments and issue timely warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564111B_ABST
    Figure CN120564111B_ABST
Patent Text Reader

Abstract

This application proposes a large-scale model-based method and device for identifying anomalies in rock masses, relating to the technical fields of target detection and risk identification. The method includes: capturing video within a target area to obtain video data; randomly sampling the to-be-segmented areas of the video frame to obtain prompt information corresponding to different features; performing multiple mask predictions based on the prompt information and the video frame using a large-scale target image segmentation model to generate a candidate mask set; obtaining a mask combination based on the candidate mask set, and splicing the masks within the mask combination to obtain a target mask map; extracting crack or fragment features from the target mask map to obtain a feature information set corresponding to the video frame; and performing risk identification on the target area based on the feature information set. Thus, this solution can achieve automatic detection, identification, and early warning of cracks and fragments in video data, improve detection accuracy and timeliness, and reduce the difficulty and danger of manual inspections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of target detection and risk identification, and in particular to a method and device for identifying anomalies in rock masses based on a large model. Background Art

[0002] In infrastructure such as buildings, bridges, and tunnels, as well as in natural environments like mines and mountains, the presence of cracks and debris often signals potential danger. For example, cracks in coal rock can indicate a threat to structural stability, while the rapid ejection of debris can signal impending danger.

[0003] Traditional crack and fragment detection methods rely primarily on manual inspections, which are not only time-consuming and labor-intensive but also subject to subjective judgment and fatigue, making it difficult to guarantee the accuracy and timeliness of test results. This is especially true in complex environments, such as underground and at height, where manual inspections are significantly more difficult and dangerous.

[0004] With the development of computer vision technology, automated inspection using video surveillance has become possible. However, existing video-based inspection methods have low accuracy in identifying cracks and fragments in complex environments and lack an effective indicator system to assess their hazard, making them difficult to meet the needs of practical applications. Summary of the Invention

[0005] The purpose of this application is to solve one of the technical problems in the related art at least to a certain extent.

[0006] To this end, the first purpose of this application is to propose a large-scale model-based method for identifying anomalies in rock masses, so as to realize automatic detection, identification and early warning of cracks and fragments in video data.

[0007] The second purpose of this application is to propose a device for identifying anomalies in rock masses based on a large model.

[0008] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a method for identifying anomalies in rock masses based on a large model, including: collecting video in the target area to obtain video data, wherein the video data includes multiple video frames; identifying the area to be segmented of the video frame, and randomly sampling the area to be segmented to obtain prompt information corresponding to different features; calling the target image segmentation large model, and performing multiple mask predictions based on the prompt information and the video frame through the target image segmentation large model to generate a candidate mask set; grouping the candidate mask set to obtain one or more mask combinations, and splicing the masks in the mask combination to obtain a target mask map corresponding to the mask combination; extracting crack or fragment features from the target mask map to obtain a feature information set corresponding to the video frame; and identifying risks in the target area based on the feature information set.

[0009] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a device for identifying anomalies in rock masses based on a large model, including: a video acquisition module, used to acquire video in a target area to obtain video data, wherein the video data includes multiple video frames; a random sampling module, used to identify the area to be segmented of the video frame, and randomly sample the area to be segmented to obtain prompt information corresponding to different features; a mask prediction module, used to call the target image segmentation large model, and perform multiple mask predictions based on the prompt information and the video frame through the target image segmentation large model to generate a candidate mask set; a mask splicing module, used to group the candidate mask set to obtain one or more mask combinations, and splice the masks in the mask combination to obtain a target mask map corresponding to the mask combination; a feature extraction module, used to extract crack or fragment features from the target mask map to obtain a feature information set corresponding to the video frame; a risk identification module, used to identify risks in the target area based on the feature information set.

[0010] The large-scale model-based anomaly identification method and device provided in the present application collects video data containing video frames in the target area and determines the area to be segmented of the video frame, so as to determine the prompt information corresponding to different features according to the area to be segmented. Furthermore, by calling the target image segmentation large model, the model performs multiple mask predictions based on the prompt information and the video frame to obtain a candidate mask set. Then, one or more mask combinations can be determined based on the candidate mask set, and a target mask map can be obtained based on the mask combination. By performing feature extraction on the target mask map, a feature information set containing crack or fragment features can be obtained, so as to issue a risk warning for the target area based on the feature information set. Therefore, this solution can realize automatic detection, identification and warning of cracks and fragments in video data, improve the accuracy and timeliness of detection, and reduce the difficulty and danger of manual inspections.

[0011] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0013] Figure 1 A schematic flow chart of a method for identifying anomalies in rock masses based on a large model provided in an embodiment of the present application;

[0014] Figure 2 A schematic flow chart of another method for identifying anomalies in rock masses based on a large model provided in an embodiment of the present application;

[0015] Figure 3 A schematic flow chart of a fine-tuning process of a target image segmentation large model in a large-model-based rock mass anomaly recognition method provided in an embodiment of the present application;

[0016] Figure 4 A schematic structural diagram of a device for identifying anomalies in rock masses based on a large model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0018] The following describes a method and device for identifying anomalies in a rock mass based on a large model according to an embodiment of the present application with reference to the accompanying drawings.

[0019] Figure 1 This is a flow chart of a method for identifying anomalies in rock mass based on a large model according to an embodiment of the present application, such as Figure 1 As shown, the method for identifying anomalies in rock mass based on a large model in an embodiment of the present application includes but is not limited to the following steps:

[0020] S101, capturing video within a target area to obtain video data, where the video data includes multiple video frames.

[0021] It should be noted that the execution entity of the large-scale model-based rock mass anomaly identification method provided in the embodiments of this application is an electronic device, which can be a terminal device. Optionally, the terminal device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), while the non-mobile electronic device can be a personal computer (PC), television, etc. This embodiment of the application does not impose any specific limitations.

[0022] In some embodiments, multiple image acquisition devices may be provided in the target area to capture video within the target area using the image acquisition devices, thereby obtaining video data including multiple video frames. For example, the image acquisition device may be a camera.

[0023] In some embodiments, a video may be captured of a target object within the target area. Alternatively, a video may be captured of a target object such as a rock or a mine within the target area to obtain video data containing the target object.

[0024] In some embodiments, video data can be obtained by collecting video within the target area at a set period or at a set video length.

[0025] S102 , identifying the area to be segmented in the video frame, and randomly sampling the area to be segmented to obtain prompt information corresponding to different features.

[0026] In some embodiments, before identifying the region to be segmented in the video frame, the video data may be preprocessed to improve the image quality and feature expression capability of the video frame. Optionally, the video data may be preprocessed such as image enhancement and denoising to obtain high-quality video frames.

[0027] For example, preprocessing operations such as histogram equalization, sharpening, Gaussian filtering, and median filtering are performed on the video data to obtain high-quality video frames, thereby identifying the area to be segmented in the preprocessed video frames.

[0028] In some embodiments, the area to be segmented of the video frame can be identified based on the target object, that is, by identifying whether the video frame contains the target object, and determining the video frame containing the target object, and setting the set range around the target object in the video frame as the area to be segmented.

[0029] In some embodiments, the region to be segmented in the video frame can be identified based on pixel features of the video frame. Alternatively, the region to be segmented in the video frame can be identified by identifying image features of the video frame and then using the image features. Alternatively, the region to be segmented in the video frame can be identified using an image segmentation algorithm.

[0030] In some embodiments, multiple sampling points are randomly selected from the area to be segmented, and feature extraction is performed on each sampling point to obtain different features, and corresponding prompt information is generated based on the different features. The features may be features of cracks or fragments.

[0031] Optionally, the prompt information may be a text description of the feature. For example, the feature may include: length, width, shape, and the corresponding prompt information may be: length feature, width feature, shape feature.

[0032] S103: calling the target image segmentation large model, and performing multiple mask predictions based on the prompt information and the video frame through the target image segmentation large model to generate a candidate mask set.

[0033] In some embodiments, by calling the target image segmentation model and inputting the prompt information and video frames into the target image segmentation model, the model performs multiple target recognitions based on the prompt information and video frames, that is, multiple mask predictions, thereby obtaining multiple predicted masks and generating a candidate mask set based on the masks.

[0034] It can be understood that the target image segmentation model can output a binary mask based on the input video frame and prompt information to mark the position of the target at the pixel level, where 1 represents the target and 0 represents the background.

[0035] S104 , grouping the candidate mask sets to obtain one or more mask combinations, and concatenating the masks within the mask combinations to obtain a target mask map corresponding to the mask combinations.

[0036] In some embodiments, the candidate mask set can be grouped according to the prediction result score of the mask. Alternatively, the prediction result score is determined based on the intersection-of-union ratio of each two masks in the candidate mask set, and the candidate mask set is further grouped to obtain one or more mask combinations. Alternatively, any two masks in the candidate mask set can be used as a mask pair, and the intersection-of-union ratio of the mask pair is calculated as the prediction result score.

[0037] In some embodiments, the mask pairs can be sorted based on the prediction result scores, and candidate mask pairs can be determined based on the sorting results, thereby grouping the candidate mask pairs to obtain one or more mask combinations. Alternatively, the mask pairs can be sorted from largest to smallest, and the top K mask pairs from the sorting results can be selected as candidate mask pairs. In other words, the top K mask pairs with the highest prediction result scores can be selected as candidate mask pairs, where K is a natural number greater than 1.

[0038] In some embodiments, it is determined whether there are overlapping masks in the candidate mask pairs, and the candidate mask pairs with overlapping masks are used as masks in the mask combination. For example, the candidate mask pairs include: M 1,2 、 M 2,3 、 M 3,4 、 M 5,6 、 M 6,7 , where masks 1, 2, 3, and 4 overlap, masks 5, 6, and 7 overlap, and masks 1, 2, 3, and 4 do not overlap with masks 5, 6, and 7, then M 1,2 、 M 2,3 、 M 3,4 As a mask combination, M 5,6 、 M 6,7 As another mask combination.

[0039] Furthermore, by splicing the masks within the mask combination, a target mask image corresponding to the mask combination can be obtained, thereby improving the integrity and accuracy of the target mask image. Optionally, the masks can be directly spliced ​​together, and weight values ​​can be assigned to different masks to splice them according to the size of the weight values.

[0040] S105 , extracting crack or fragment features from the target mask image to obtain a feature information set corresponding to the video frame.

[0041] In some embodiments, crack features include the path length of the crack centerline, the maximum crack width, and the main direction of the crack; while fragment features include the fragment area, the fragment center of mass, and the fragment motion direction. In other words, the feature information set corresponding to the video frame includes crack features and fragment features.

[0042] Alternatively, the crack or fragment characteristics can be determined by extracting pixel features from the target mask image. Pixel features include pixel position, pixel number, and so on. For example, taking the path length of a crack centerline as an example, the starting and ending positions of the target mask image can be determined based on pixel positions, and the path length of the crack centerline can be calculated based on the starting and ending positions.

[0043] For another example, taking the fragment area as an example, the fragment area can be determined based on the number of pixels.

[0044] S106: Identify risks of the target area based on the feature information set.

[0045] In some embodiments, different risk identification conditions can be set, and it can be determined whether the feature information in the feature information set meets any risk identification condition. If so, it can be determined that the target area has a risk; otherwise, it can be determined that the target area does not have a risk.

[0046] In some embodiments, after determining that the target area is at risk, the risk level of the target area can be determined based on the feature information set. Optionally, the number of characteristic information in the feature information set that meets the risk identification condition can be obtained, and the risk level can be determined based on the number.

[0047] For example, if the number that meets the risk identification conditions is within the first range, the risk level is determined to be level A; if the number that meets the risk identification conditions is within the second range, the risk level is determined to be level B, where the first range is smaller than the second range, and level A is smaller than level B.

[0048] In some embodiments, after determining that a target area has risks, a risk warning can be issued for the target area, providing strong support for ensuring the safety of personnel and property.

[0049] In the large-scale model-based method for identifying anomalies in rock masses provided in an embodiment of the present application, video data containing video frames in the target area is collected, and the area to be segmented of the video frames is determined, so as to determine prompt information corresponding to different features according to the area to be segmented. Furthermore, by calling the target image segmentation large model, the model performs multiple mask predictions based on the prompt information and the video frames to obtain a candidate mask set. Then, one or more mask combinations can be determined based on the candidate mask set, and a target mask map can be obtained based on the mask combination. By performing feature extraction on the target mask map, a feature information set containing crack or fragment features can be obtained, so as to perform risk warning on the target area based on the feature information set. Therefore, this solution can realize automatic detection, identification and warning of cracks and fragments in video data, improve the accuracy and timeliness of detection, and reduce the difficulty and danger of manual inspections.

[0050] Figure 2 This is a flow chart of a method for identifying anomalies in rock mass based on a large model according to an embodiment of the present application, such as Figure 2 As shown, the method for identifying anomalies in rock mass based on a large model in an embodiment of the present application includes but is not limited to the following steps:

[0051] S201 , capturing video within a target area to obtain video data, where the video data includes multiple video frames.

[0052] S202 , identifying the area to be segmented in the video frame, and randomly sampling the area to be segmented to obtain prompt information corresponding to different features.

[0053] S203: calling the target image segmentation large model, and performing multiple mask predictions based on the prompt information and the video frame through the target image segmentation large model to generate a candidate mask set.

[0054] In the embodiment of the present application, steps S201-S203 can be implemented by any one of the methods in the embodiments of the present application, which is not limited here and will not be described in detail.

[0055] S204: Taking any two masks in the candidate mask set as a mask pair, and obtaining an intersection-over-union ratio of the two masks in the mask pair.

[0056] S205 : Determine the prediction result score of the mask pair according to the intersection-over-union ratio of the mask pair.

[0057] In some embodiments, any two masks in the candidate mask set are selected as a mask pair, and an IoU calculation is performed based on the mask pair, so that the IoU of the mask pair can be used as the prediction result score of the mask pair.

[0058] Optionally, for the first of multiple mask predictions i Second and j Second prediction, calculate the intersection of the two masks in the mask pair IoU i,j The formula is as follows:

[0059] (1)

[0060] in, M i Indicates the i The predicted mask, M j Indicates the j The predicted mask is M i ∩ M j express M i andM j The number of intersection pixels, M i ∪ M j express M i and M j The number of pixels in the union.

[0061] Furthermore, the prediction result score S i,j Intersection-Union IoU i,j In other words:

[0062] (2)

[0063] S206 , sorting all mask pairs based on the prediction result scores, and screening out candidate mask pairs according to the sorting results.

[0064] In some embodiments, all mask pairs can be sorted from largest to smallest based on the prediction result scores. That is, mask pairs with larger prediction result scores are placed at the front, and mask pairs with smaller prediction result scores are placed at the back. The top K mask pairs can then be selected from the sorted results as candidate mask pairs. K is a natural number greater than 1.

[0065] Optionally, the prediction result scores may be sorted first, and a sorted mask pair sequence may be determined based on the sorted prediction result scores, and then the first K mask pairs may be selected from the mask pair sequence as candidate mask pairs.

[0066] For example, if K=3, the ranking of the prediction results is: S (1,2) ≥ S (2,3) ≥…≥ S (n-1,n) , then the corresponding mask pair sequence is: M (1,2) , M (2,3) ,…, M (n-1,n) , select the first three mask pairs from the mask pair sequence as: M (1,2) 、 M (2,3) 、 M (3,4) .

[0067] In some embodiments, in order to eliminate the situation where the prediction result scores are abnormal due to the instability of the large model, before screening the candidate mask pairs according to the sorting results, the candidate mask set can be regenerated when the prediction result scores are abnormal to improve the accuracy and stability of the candidate mask set.

[0068] In some embodiments, scenarios that cause large-scale model instability can be pre-determined as scenarios that meet set requirements, so that when the prediction result score is abnormal, the candidate mask set is regenerated. Among them, abnormal weather can be used as a scenario that meets set requirements, and significant changes in video data scenes can also be used as a scenario that meets set requirements.

[0069] Optionally, a prediction result score less than a set value may be regarded as a prediction result score abnormality, and a difference between prediction result scores that is too large or too small may be regarded as a prediction result score abnormality.

[0070] In other words, in response to an abnormal prediction score, a candidate mask set is regenerated in scenarios that meet the set requirements. The candidate mask set is generated by re-calling the target image segmentation model and performing multiple mask predictions based on the prompt information and video frames using the target image segmentation model.

[0071] S207 , performing mask overlap identification on the candidate mask pairs to group the candidate mask pairs to obtain one or more mask combinations, where masks in the mask combinations have overlapping parts.

[0072] In some embodiments, it can be determined whether any two candidate mask pairs have the same mask. If so, it can be determined that the two candidate mask pairs have mask overlap, and the candidate mask pairs with mask overlap can be regarded as a mask combination.

[0073] For example, among the candidate mask pairs, there is mask overlap between candidate mask pair A and candidate mask pair B, and there is mask overlap between candidate mask pair C and candidate mask pair D. There is no mask overlap between candidate mask pair A, candidate mask pair B and candidate mask pair C, candidate mask pair D. Then, candidate mask pair A and candidate mask pair B can be used as mask combination 1, and candidate mask pair C and candidate mask pair D can be used as mask combination 2.

[0074] For another example, the candidate mask pairs include: M 1,2 、 M 2,3 、 M 3,4 、 M 5,6 、 M 6,7 , where the candidate mask pair M1,2 、 M 2,3 、 M 3,4 There is mask overlap, candidate mask pair M 5,6 、 M 6,7 There is mask overlap, M 1,2 、 M 2,3 、 M 3,4 and M 5,6 、 M 6,7 If there is no mask overlap, then M 1,2 、 M 2,3 、 M 3,4 As a mask combination, M 5,6 、 M 6,7 As another mask combination.

[0075] S208 , concatenating the masks in the mask combination to obtain a target mask image corresponding to the mask combination.

[0076] In some embodiments, the masks in the mask combination can be directly spliced ​​to obtain the target mask image corresponding to the mask combination. l mask combinations, within which there are n masks, the formula for determining the target mask map is:

[0077] (3)

[0078] in, M l,final represents the target mask map, M j Represents the jth mask among n masks.

[0079] S209: Extract crack or fragment features from the target mask image to obtain a feature information set corresponding to the video frame.

[0080] In some embodiments, the feature information set corresponding to the video frame includes crack features or fragment features, wherein the crack features include but are not limited to the path length of the crack centerline, the maximum width of the crack, and the main direction of the crack; the fragment features include but are not limited to the fragment area, the center of mass of the fragment, and the movement direction of the fragment.

[0081] In some embodiments, crack features or fragment features may be calculated based on the target mask image, thereby extracting crack or fragment features from the target mask image to generate a feature information set corresponding to the video frame.

[0082] In some embodiments, the path length of the crack centerline can be calculated based on the starting and ending positions of the target mask image. The starting and ending positions of the target mask image can be determined based on the pixel positions of the target mask image, and the path length of the crack centerline associated with the target mask image can be determined based on the starting and ending positions.

[0083] Alternatively, the formula for determining the path length of the fracture centerline is as follows:

[0084] (4)

[0085] in, L represents the path length of the crack centerline, ( x i , y i ) indicates the starting point, ( x i+1 , y i+1 ) indicates the end position, r Indicates the ratio between the size of the crack or fragment in the video frame and its actual size.

[0086] Optionally, the sample crack or fragment can be photographed in advance to determine the first size of the crack or fragment in the image, and the actual physical second size of the sample crack or fragment can be measured. By calculating the ratio of the first size to the second size, the r .

[0087] In some embodiments, the crack region can be determined based on the target mask image, and the vertical width within the region can be calibrated using the crack centerline as a reference to determine the maximum crack width. Alternatively, the vertical width can be calibrated using circles of different diameters to calculate the maximum crack width. The calculation formula is as follows:

[0088] (5)

[0089] in, W max represents the maximum width of the crack, d i Indicates the diameter of a circle.

[0090] In some embodiments, the main direction of the crack can be determined based on the path length of the crack centerline and the maximum width of the crack. Alternatively, the main direction of the crack can be determined by calculating the ratio of the path length of the crack centerline and the maximum width of the crack using a trigonometric function relationship. The calculation formula is as follows:

[0091] (6)

[0092] in, θ Indicates the main direction of the crack.

[0093] In some embodiments, the crack features further include: crack length per unit area and crack width per unit area. The crack length per unit area and the crack width per unit area can be calculated based on the total crack path length and total crack width in the video frame, as well as the crack area in the video frame.

[0094] In some embodiments, the total path length of the crack corresponding to the video frame can be determined based on the path length of the crack centerline corresponding to the target mask image, and the total width of the crack corresponding to the video frame can be determined based on the maximum width of the crack corresponding to the target mask image.

[0095] That is to say, by l The total crack path length can be obtained by summing the path lengths of the crack center lines corresponding to the target mask images of each mask combination. l The total crack width can be obtained by summing up the maximum crack widths corresponding to the target mask images of each mask combination.

[0096] In some embodiments, the crack area in the video frame is determined by obtaining the number of pixels in the target mask image and using the number of pixels in the target mask image. Alternatively, the number of crack pixels can be obtained from the number of pixels in the target mask image and the pixel size can be determined, thereby determining the crack area based on the number of crack pixels and the pixel size. Alternatively, the number of pixels can be directly used as the crack area in the video frame.

[0097] Furthermore, the crack length per unit area may be determined based on the crack area and the total crack path length, and the crack width per unit area may be determined based on the crack area and the total crack width.

[0098] Alternatively, the formula for determining the crack length per unit area is as follows:

[0099] (7)

[0100] in, ρ represents the crack length per unit area, ∑ L i represents the total crack path length,A region Represents the crack area.

[0101] Alternatively, the total crack width can be substituted for the total crack path length in the above formula (7), and the crack width per unit area can be calculated.

[0102] In some embodiments, the number of pixels corresponding to the target mask image may be determined, and the area of ​​the fragment may be determined based on the number of pixels corresponding to the target mask image. Optionally, the formula for calculating the area of ​​the fragment may be as follows:

[0103] (8)

[0104] in, A represents the area of ​​the fragment, N pixel Indicates the number of pixels, r Indicates the ratio between the size of the crack or fragment in the video frame and its actual size.

[0105] In some embodiments, the fragment centroid can be determined based on the center position of the mask in the mask combination corresponding to the target mask image. Alternatively, the fragment centroid can be determined based on the center position of the mask in the mask combination. Alternatively, the number of fragment pixels can be obtained, and the fragment centroid can be calculated based on the center position and the number of fragment pixels. Calculating the fragment centroid ( x centroid , y centroid ) is as follows:

[0106] (9)

[0107] in,( x i , y i ) represents the center position of the mask, N Indicates the number of fragment pixels.

[0108] In some embodiments, the movement direction of the fragment can be calculated based on the starting position and the ending position of the target mask image. That is, by determining the starting position and the ending position of the target mask image, and then determining the movement direction of the fragment based on the starting position and the ending position, the calculation formula is as follows:

[0109] (10)

[0110] in, Indicates the direction of movement of the fragments, ( x start , y start) indicates the starting point, ( x end , y end ) indicates the end point position.

[0111] It is understandable that due to the generative nature of large models, the calculated feature information may be unstable. The accuracy of the feature information set can be improved by determining the abnormal feature information in the feature information set and eliminating the abnormal feature information.

[0112] In some embodiments, a sliding average may be performed on the features in the feature information set to determine abnormal feature information and eliminate the abnormal feature information.

[0113] Optionally, by calculating the mean of the feature information μ t and standard deviation σ t , and determine the set value corresponding to the feature information, so as to compare the difference between the set value and the mean with the standard deviation, and determine the abnormal feature information according to the comparison result, and remove the abnormal feature information.

[0114] Alternatively, the formula for comparing the standard deviation based on the difference between the set value and the mean is as follows:

[0115] (11)

[0116] in, y t Indicates the set value, α represents a constant, and α =3. That is, when the difference between the set value and the mean is greater than 3 times the standard deviation, the feature information is considered to be abnormal feature information.

[0117] S210: Identify risks of the target area based on the feature information set.

[0118] In some embodiments, for any feature information in the feature information set, a risk identification condition that any feature information needs to meet can be obtained, and when any feature information meets the corresponding risk identification condition, it can be identified that the target area has a risk.

[0119] For example, for the crack length per unit area in the feature information set, the first length along the set direction is determined, and the first length is greater than or equal to the first set risk value, which is used as the risk identification condition that the crack length per unit area must meet.

[0120] If the first set risk value is 0.7, the risk identification condition that the crack length per unit area must meet is: ,in, θ i To set the direction, is the first length.

[0121] That is, by determining a first length of the crack length per unit area in the feature information set along a set direction, in response to the first length being greater than or equal to a first set risk value, it is determined that the target area is at risk.

[0122] For another example, for the crack width per unit area in the feature information set, the first width along the set direction is determined, and the first width is greater than or equal to the second set risk value, which is used as the risk identification condition that the crack width per unit area must meet.

[0123] If the second set risk value is 0.5, the risk identification condition that the crack width per unit area must meet is: ,in, , is the first width.

[0124] That is, by determining a first width of the crack width per unit area in the feature information set along a set direction, in response to the first width being greater than or equal to a second set risk value, it is determined that the target area is at risk.

[0125] For another example, for the movement direction of the fragments in the feature information set, the maximum area of ​​the fragments along the movement direction within the unit area is determined, and the ratio of the maximum area to the image area is obtained, and the ratio is greater than or equal to the third set risk value, which is used as the risk identification condition that the movement direction of the fragments must meet.

[0126] If the third set risk value is 2.8, the risk identification conditions that the movement direction of the fragments must meet are: ,in, Indicates the j The fragments are moving in the direction The projected area on Indicates the ratio of the maximum area to the image area.

[0127] That is, by determining the maximum area of ​​the fragments within a unit area along the moving direction, and based on the ratio of the maximum area to the image area, in response to the ratio being greater than or equal to the third set risk value, it is determined that the target area is at risk.

[0128] In some embodiments, whether a target area is at risk can also be determined based on the number of fragments and the distance between the centroids of the fragments and the crack. This is accomplished by determining the number of fragments surrounding a crack indicated by the feature information set and calculating the shortest distance from the centroids of the fragments to the crack. If the feature information set indicates that a set number of fragments are surrounding the crack, and the shortest distance from the centroids of the fragments to the crack is greater than the width of the crack, the target area is determined to be at risk.

[0129] For example, if , determine that there are risks in the target area. Among them, C Indicates a crack. d (x j ,C) represents the j The center of mass of the fragment x j To the rift C The shortest distance, W C Indicates the width of the crack, 1 {•} Represents an indicator function (1 if the condition is met, 0 otherwise).

[0130] In some embodiments, after determining that a target area is at risk, the risk level can be determined based on the risk identification conditions satisfied by different feature information, and a risk warning can be issued for the target area based on the risk level to ensure the safety of personnel and property.

[0131] Optionally, the number of feature information in the feature information set that meets the risk identification condition may be obtained, and the risk level may be determined based on the number.

[0132] In the large-scale model-based anomaly identification method for rock masses provided in an embodiment of the present application, the intersection-over-union ratio of mask pairs in a candidate mask set is determined as the prediction result score of the mask pair, and the mask pairs are sorted based on the prediction result score to screen out candidate mask pairs according to the sorting results, and mask overlap identification is performed on the candidate mask pairs to determine the mask combination, and a target mask map is obtained based on the mask combination. By performing feature extraction on the target mask map, a feature information set containing crack or fragment features can be obtained to perform risk warnings on the target area based on the feature information set. Therefore, this solution can utilize the powerful feature learning ability of the large model to accurately identify cracks and fragments in the video under complex environments, thereby improving the accuracy and reliability of recognition. By processing video data in real time, changes in cracks and fragments can be discovered in a timely manner, and early warnings can be quickly issued when risks arise, providing strong support for ensuring the safety of people and property.

[0133] Based on the above embodiments, the present invention can explain the fine-tuning process of the target image segmentation model. Figure 3As shown, the fine-tuning process of the target image segmentation model in the embodiment of the present application includes but is not limited to the following steps:

[0134] S301 : Acquire a rock mass image set, and segment the rock mass images in the rock mass image set to obtain segmentation and annotation results of the rock mass images.

[0135] In some embodiments, any image can be selected from a rock image library as a rock image set, and / or rock images can be downloaded from the Internet as a rock image set, and / or images can be collected based on an image acquisition device as a rock image set, and / or AI images created by users through software or work can be used as a rock image set.

[0136] In some embodiments, multiple image segmentation algorithms can be obtained and used to segment the rock mass images in the rock mass image collection, thereby obtaining segmentation and annotation results for the rock mass images. Optionally, during the rock mass image segmentation process, features of cracks or fragments can be annotated to obtain segmentation and annotation results containing these features. For example, the length, width, and shape of the cracks or fragments can be annotated.

[0137] Among them, the image segmentation algorithms include but are not limited to U-NEt, FCN, Mask R-CNN and other image segmentation algorithms.

[0138] In some embodiments, different image segmentation algorithms can be used to segment and annotate the same feature, and the annotation results can be fused to obtain a segmentation and annotation result. Optionally, the fusion method is 0-1 fusion. For example, if five algorithms are used to annotate the length feature, if any of the algorithms recognizes the length feature, the final segmentation and annotation result is determined to include the length feature.

[0139] S302 , based on the segmentation and annotation results of the rock mass images and the manual annotation results of the rock mass images, effective image data is screened from the rock mass image collection to construct a fine-tuning data set required for the image segmentation large model.

[0140] In some embodiments, it can be determined whether the first recognition range corresponding to the segmentation and annotation results of the rock image exceeds the second recognition range of the manual annotation results of the rock image. If the first recognition range does not exceed the second recognition range, the rock image corresponding to the segmentation and annotation results can be determined to be valid image data; otherwise, the rock image is determined to be invalid image data.

[0141] Furthermore, based on the segmentation and annotation results, the image annotations can be matched with the valid image data to construct the fine-tuning dataset required for the large image segmentation model. Alternatively, for any annotation, the segmentation and annotation results can be used to determine whether the valid image data contains the annotation. If so, the annotation is determined to match the valid image data. By matching the annotations with the valid image data one by one, a fine-tuning dataset can be obtained.

[0142] S303, based on the fine-tuning data set, fine-tune the image segmentation large model until the fine-tuning training end condition is met and the training is terminated to obtain the target image segmentation large model.

[0143] In some embodiments, valid image data and its matching annotations are randomly selected from the fine-tuning dataset and input into a large image segmentation model, which then outputs a corresponding prediction mask. The matching annotations of the valid image data are used as reference results, and a model loss is calculated based on the prediction mask and the reference results. Model parameters are then updated via backpropagation based on the model loss, thereby fine-tuning the large image segmentation model. Training is terminated when the fine-tuning training termination conditions are met, resulting in a target large image segmentation model.

[0144] In some embodiments, the training end condition may be that the number of training times reaches a set number; the training end condition may also be that the model loss reaches a set value; the training end condition may also be that the accuracy of the model output result reaches an accuracy threshold.

[0145] In the large-scale model-based anomaly identification method for rock masses provided in an embodiment of the present application, a fine-tuning dataset is constructed by obtaining the segmentation and annotation results corresponding to a rock mass image set and filtering valid image data from the rock mass image set based on the segmentation and annotation results. The fine-tuning dataset is further used to fine-tune the image segmentation large model until the fine-tuning training end conditions are met and the training is terminated, thereby obtaining a target image segmentation large model. The fine-tuned target image segmentation large model can maintain the accuracy of model recognition when external conditions change or the scene changes, and can accurately identify cracks and fragments in videos in complex environments, thereby improving the accuracy and reliability of recognition and helping to ensure the stability of risk identification and early warning.

[0146] The above-mentioned embodiments correspond to the large-model-based anomaly identification methods in rock masses. An embodiment of the present application also proposes a large-model-based anomaly identification device in rock masses. Since the large-model-based anomaly identification device in rock masses proposed in the embodiment of the present application corresponds to the large-model-based anomaly identification methods in rock masses proposed in the above-mentioned embodiments, the implementation method of the above-mentioned large-model-based anomaly identification method in rock masses is also applicable to the large-model-based anomaly identification device in rock masses proposed in the embodiment of the present application, and will not be described in detail in the following embodiments.

[0147] In order to implement the above embodiment, the present application also proposes a device for identifying anomalies in rock masses based on a large model.

[0148] Figure 4 A schematic structural diagram of a device for identifying anomalies in rock masses based on a large model is provided in an embodiment of the present application.

[0149] like Figure 4 As shown, the large-scale model-based rock mass anomaly identification device 400 includes:

[0150] The video acquisition module 401 is used to acquire video data within the target area, where the video data includes multiple video frames.

[0151] The random sampling module 402 is used to identify the area to be segmented in the video frame and randomly sample the area to be segmented to obtain prompt information corresponding to different features;

[0152] The mask prediction module 403 is used to call the target image segmentation model and perform multiple mask predictions based on the prompt information and the video frame through the target image segmentation model to generate a candidate mask set;

[0153] The mask splicing module 404 is used to group the candidate mask sets to obtain one or more mask combinations, and splice the masks in the mask combinations to obtain a target mask image corresponding to the mask combinations;

[0154] The feature extraction module 405 is used to extract crack or fragment features from the target mask image to obtain a feature information set corresponding to the video frame;

[0155] The risk identification module 406 is used to identify risks of the target area according to the feature information set.

[0156] In a possible implementation of an embodiment of the present application, the mask splicing module 404 is further used to: take any two masks in the candidate mask set as a mask pair, and obtain the intersection-over-union ratio of the two masks in the mask pair; determine the prediction result score of the mask pair based on the intersection-over-union ratio of the mask pair; sort all mask pairs based on the prediction result score, and filter out candidate mask pairs based on the sorting result; perform mask overlap identification on the candidate mask pairs to group the candidate mask pairs to obtain one or more mask combinations, where there are overlapping parts between the masks in the mask combination.

[0157] In a possible implementation of an embodiment of the present application, the feature extraction module 405 is also used to: determine the starting position and the end position of the target mask image; determine the path length of the crack centerline associated with the target mask image based on the starting position and the end position; determine the area where the crack is located based on the target mask image, and calibrate the vertical width within the area with the crack centerline as the reference to determine the maximum width of the crack; determine the main direction of the crack based on the path length of the crack centerline and the maximum width of the crack.

[0158] In a possible implementation of an embodiment of the present application, the feature extraction module 405 is also used to: determine the total path length of the crack corresponding to the video frame based on the path length of the crack centerline corresponding to the target mask image; determine the total width of the crack corresponding to the video frame based on the maximum width of the crack corresponding to the target mask image; determine the crack area in the video frame based on the number of pixels of the target mask image; determine the crack length per unit area based on the crack area and the total path length of the crack; determine the crack width per unit area based on the crack area and the total width of the crack.

[0159] In a possible implementation of an embodiment of the present application, the feature extraction module 405 is further used to: determine the area of ​​the fragment based on the number of pixels corresponding to the target mask image; determine the center position of the mask in the mask combination, and determine the center of mass of the fragment based on the center position; determine the starting position and end position of the target mask image, and determine the movement direction of the fragment based on the starting position and end position.

[0160] In a possible implementation of the embodiment of the present application, the mask splicing module 404 is further configured to: in response to an abnormal prediction result score, regenerate a candidate mask set in a scenario that meets set requirements.

[0161] In a possible implementation of the embodiment of the present application, the feature extraction module 405 is further configured to perform a sliding average on the features in the feature information set to determine abnormal feature information therefrom and eliminate the abnormal feature information.

[0162] In a possible implementation of an embodiment of the present application, the risk identification module 406 is further used to: determine a first length of the crack length within a unit area in the feature information set along a set direction, and in response to the first length being greater than or equal to a first set risk value, determine that there is a risk in the target area; determine a first width of the crack width within a unit area in the feature information set along a set direction, and in response to the first width being greater than or equal to a second set risk value, determine that there is a risk in the target area; determine the maximum area of ​​the fragments within the unit area along the direction of movement, and based on the ratio of the maximum area to the image area, in response to the ratio being greater than or equal to a third set risk value, determine that there is a risk in the target area; in response to the feature information set indicating that there are a set number of fragments around the crack, and the shortest distance from the center of mass of the fragment to the crack is greater than the width of the crack, determine that there is a risk in the target area.

[0163] In a possible implementation of an embodiment of the present application, the mask prediction module 403 is also used to: obtain a rock image set, and segment the rock images in the rock image set to obtain segmentation and annotation results of the rock images; based on the segmentation and annotation results of the rock images and the manual annotation results of the rock images, screen valid image data from the rock image set to construct a fine-tuning data set required for the image segmentation large model; based on the fine-tuning data set, fine-tune the image segmentation large model until the fine-tuning training end condition is met and the training is terminated to obtain the target image segmentation large model.

[0164] In the large-scale model-based rock mass anomaly identification device provided in an embodiment of the present application, video data containing video frames in the target area is collected, and the area to be segmented of the video frame is determined, so as to determine the prompt information corresponding to different features according to the area to be segmented. Furthermore, by calling the target image segmentation large model, the model performs multiple mask predictions based on the prompt information and the video frame to obtain a candidate mask set. Then, one or more mask combinations can be determined based on the candidate mask set, and a target mask map can be obtained based on the mask combination. By extracting features from the target mask map, a feature information set containing crack or fragment features can be obtained, so as to issue a risk warning for the target area based on the feature information set. Therefore, this solution can realize automatic detection, identification and warning of cracks and fragments in video data, thereby improving the accuracy and timeliness of detection and reducing the difficulty and danger of manual inspections.

[0165] It should be noted that the above explanation of the embodiment of the method for identifying anomalies in rock masses based on a large model is also applicable to the device for identifying anomalies in rock masses based on a large model in this embodiment, and will not be repeated here.

[0166] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0167] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0168] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0169] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" is any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0170] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0171] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0172] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0173] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for identifying anomalies in rock mass based on a large model, characterized in that: The method comprises: Capturing video within the target area to obtain video data, wherein the video data includes a plurality of video frames; Identifying a region to be segmented in the video frame, and randomly sampling the region to be segmented to obtain prompt information corresponding to different features; Calling a target image segmentation model, and performing multiple mask predictions based on the prompt information and the video frame by the target image segmentation model to generate a candidate mask set; Grouping the candidate mask sets to obtain one or more mask combinations, and concatenating the masks within the mask combinations to obtain a target mask map corresponding to the mask combinations; Extracting crack or fragment features from the target mask image to obtain a feature information set corresponding to the video frame; Performing risk identification on the target area according to the feature information set; The method further comprises: Determining the starting position and the ending position of the target mask image; Determining a path length of a crack centerline associated with the target mask image based on the starting position and the end position; Determine the region where the crack is located according to the target mask image, and calibrate the vertical width within the region with the center line of the crack as a reference to determine the maximum width of the crack; determining the main direction of the crack according to the path length of the crack centerline and the maximum width of the crack; The method further comprises: Determining the total path length of the crack corresponding to the video frame according to the path length of the crack centerline corresponding to the target mask image; Determining a total crack width corresponding to the video frame according to a maximum crack width corresponding to the target mask image; Determining the crack area in the video frame according to the number of pixels of the target mask image; Determining the crack length per unit area based on the crack area and the total crack path length; Determining the crack width per unit area according to the crack area and the total crack width; The method further comprises: Determining the area of ​​the fragment according to the number of pixels corresponding to the target mask image; Determining a center position of a mask in the mask combination, and determining a centroid of a fragment based on the center position; Determining a starting position and an end position of the target mask image, and determining a moving direction of the fragment according to the starting position and the end position; The performing risk identification on the target area according to the feature information set includes at least one of the following operations: determining a first length of a crack length per unit area in the feature information set along a set direction, and determining that a risk exists in the target area in response to the first length being greater than or equal to a first set risk value; determining a first width of a crack width per unit area in the feature information set along a set direction, and determining that a risk exists in the target area in response to the first width being greater than or equal to a second set risk value; determining a maximum area of ​​the fragments within a unit area along the direction of motion, and determining, based on a ratio of the maximum area to the image area, that the target area is at risk in response to the ratio being greater than or equal to a third set risk value; In response to the feature information set indicating that a set number of fragments exist around the crack, and the shortest distance from the centroid of the fragments to the crack is greater than the width of the crack, it is determined that the target area is at risk.

2. The method according to claim 1, characterized in that The grouping of the candidate mask sets to obtain one or more mask combinations includes: Taking any two masks in the candidate mask set as a mask pair, and obtaining the intersection-over-union ratio of the two masks in the mask pair; Determining a prediction result score of the mask pair according to an intersection-over-union ratio of the mask pair; Sorting all mask pairs based on the prediction result scores, and screening candidate mask pairs according to the sorting results; Mask overlap identification is performed on the candidate mask pairs to group the candidate mask pairs to obtain one or more mask combinations, where masks in the mask combinations have overlapped parts.

3. The method according to claim 2, characterized in that Before sorting all mask pairs based on the prediction result scores and selecting candidate mask pairs according to the sorting results, the method further includes: In response to the prediction result score being abnormal, the candidate mask set is regenerated in a scenario that meets set requirements.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Perform a sliding average on the features in the feature information set to determine abnormal feature information and eliminate the abnormal feature information.

5. The method according to any one of claims 1 to 3, characterized in that The fine-tuning process of the target image segmentation model includes: Acquire a rock mass image set, and segment the rock mass images in the rock mass image set to obtain segmentation and annotation results of the rock mass images; Based on the segmentation and annotation results of the rock mass images and the manual annotation results of the rock mass images, valid image data is screened from the rock mass image collection to construct a fine-tuning data set required for a large image segmentation model; Based on the fine-tuning data set, fine-tuning training is performed on the large image segmentation model until the fine-tuning training end condition is met and the training is terminated to obtain the target large image segmentation model.

6. A device for identifying anomalies in rock mass based on a large model, characterized in that: The apparatus is configured to implement the method according to any one of claims 1 to 5, and the apparatus includes: A video acquisition module, configured to acquire video data within a target area, wherein the video data includes a plurality of video frames; A random sampling module is used to identify the area to be segmented in the video frame and randomly sample the area to be segmented to obtain prompt information corresponding to different features; A mask prediction module is used to call a target image segmentation large model, and perform multiple mask predictions based on the prompt information and the video frame through the target image segmentation large model to generate a candidate mask set; a mask splicing module, configured to group the candidate mask sets to obtain one or more mask combinations, and splice the masks within the mask combinations to obtain a target mask map corresponding to the mask combinations; A feature extraction module is used to extract crack or fragment features from the target mask image to obtain a feature information set corresponding to the video frame; A risk identification module is used to identify risks of the target area based on the feature information set.