Gas hazard identification method and system based on multi-modal large model
By using a multimodal large model to identify potential spatial conflicts in the installation and connection of gas equipment, and by constructing a two-dimensional topological relationship model and combining it with national standard logical rules, the problem of the inability to detect the spatial relationship between the physical installation and connection of gas equipment in existing technologies has been solved, thus achieving efficient identification of potential hazards in gas equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENENG NATURAL GAS CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-21
Smart Images

Figure CN122432871A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gas pipeline technology, and more specifically, to a method and system for identifying gas hazard based on a multimodal large model. Background Technology
[0002] Small and medium-sized gas equipment structures such as gas pipeline interfaces, valves, and flanges are high-risk areas for potential hazards in gas transmission and distribution scenarios. These devices undertake key functions such as pipeline connection, opening and closing control, and sealing and pressure bearing, and are the key inspection targets for daily safety inspections of gas systems. In the industry, manual on-site inspection is the main inspection method.
[0003] In the existing technology, relevant patents have been researched in the fields of gas pipeline network data inspection and multimodal hazard identification. For example, invention patent CN202311575145.X discloses a method, system, and chip for inspecting abnormal data in gas pipeline network topology data. This method acquires and standardizes the topology data of gas pipeline network point tables and line tables, filters and repairs abnormal data through pipe diameter mutation, ID, and coordinate anomaly detection, and automatically updates pipeline network attributes and topology data, reducing data processing costs. Another example is invention patent CN202510850375.5, which discloses a multimodal collaborative gas pipeline network hazard identification method and system. This method collects pipeline network pressure, temperature, and gas concentration sensor data, and inputs them into an intelligent model after wavelet packet decomposition, entropy feature extraction, and normalization to achieve the location and hierarchical identification of pipeline network leakage and corrosion hazards.
[0004] While the aforementioned technical solutions possess corresponding design advantages, they also suffer from the following technical shortcomings: Existing technologies do not rely on on-site acquired equipment images to detect the spatial relationships of the physical installation and connections of gas equipment. CN202311575145.X only performs anomaly detection and repair on gas pipeline network topology data, without addressing spatial structural hazard identification in the image dimension of gas equipment; CN202510850375.5 relies solely on pressure, temperature, and gas concentration sensor data for pipeline hazard identification, without analyzing the image features and installation connection spatial relationships of small and medium-sized gas equipment. Neither of these methods can achieve image-based spatial structural hazard identification of gas pipeline interfaces, valves, flanges, and other equipment. Therefore, we propose a gas hazard identification method and system based on a multimodal large model. Summary of the Invention
[0005] The purpose of this invention is to provide a gas hazard identification method and system based on a multimodal large model, so as to solve the problem that the existing technologies mentioned in the background do not rely on on-site collected equipment images to detect the spatial relationship of the physical installation and connection of gas equipment.
[0006] To address the aforementioned technical problems, one objective of this invention is to provide a gas hazard identification method based on a multimodal large model, comprising the following steps: S1. Image Acquisition and Quality Inspection: Acquire on-site inspection images of gas pipelines and related equipment, calculate the gradient variance of the on-site inspection images using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a reacquisition prompt for low-quality images, and output the sharpness score and image quality level to step S3. S2, Hazard Target Detection: Input qualified on-site inspection images into the Transformer-based target detection model, extract multi-scale features through a convolutional backbone network, perform global modeling through a Transformer with deformable attention mechanism, and output the category information of gas equipment components and hazards and bounding box detection results using an ensemble prediction mechanism, and output the detection results to step S3; S3. Multimodal Semantic Understanding and Spatial Conflict Identification: Multimodal input is constructed by fusing on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. The hazard semantic analysis is completed through the visual-text cross-attention fusion feature of the multimodal large model. Based on the detection results, the categories, coordinates, directions, and spatial adjacency relationships of gas equipment components are extracted and a two-dimensional topological relationship model is constructed. Consistency verification is performed by combining the standard equipment structure logic rules preset by the national standard. The verification threshold is adaptively adjusted through a dynamic threshold self-calibration mechanism to identify spatial conflict hazards related to component installation, connection, and structure. At the same time, zero-sample identification of untrained hazard types is achieved based on the pre-trained knowledge of the multimodal large model. Comprehensive hazard analysis data is generated and output to step S4. S4. Structured Result Output and Process Linkage: The comprehensive hazard analysis data is processed in a unified structure to output a structured gas hazard identification report containing hazard details, risk level, cause analysis, treatment suggestions and confidence level. Gas operation and maintenance parameters are adjusted based on the structured gas hazard identification report.
[0007] As a further improvement to this technical solution, the image acquisition and quality detection in S1 includes the following steps: S11. Collect on-site inspection images of gas pipelines, valves, pressure regulators, flanges and pipeline interfaces, and complete image format unification and encoding conversion processing; S12. The Laplacian operator is used to perform convolution operation on the preprocessed on-site inspection image to calculate the gradient variance data of the entire image. S13. Perform linear normalization on the calculated gradient variance data, map it to the numerical range of 0 to 100, and generate an image sharpness score. S14. Divide the sharpness score according to the preset three-level threshold: a score of 70 or above is judged as sharp, a score between 30 and 70 is judged as medium, and a score less than 30 is judged as blurry. S15. When an image is determined to be blurry, generate an image re-acquisition prompt and terminate the subsequent detection process of the current image. S16. Output the sharpness score and image quality level corresponding to the clear and medium level images to step S3.
[0008] As a further improvement to this technical solution, the hazard target detection in S2 includes the following steps: S21. The on-site inspection image determined to be clear and of medium quality in step S1 is input into the target detection model based on Transformer, wherein the target detection model based on Transformer is the RF-DETR model. S22. Extract multi-scale visual features from on-site inspection images using a convolutional backbone network; S23. Global feature modeling of the image is completed using a Transformer with deformable attention mechanism; S24. Use an ensemble prediction mechanism to output the category information and bounding box detection results of gas equipment components and potential hazards; S25. Output the category information and bounding box detection results to step S3.
[0009] As a further improvement to this technical solution, the multimodal semantic understanding and spatial contradiction recognition in S3 includes the following steps: S31. Multimodal input data is constructed by integrating on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. The multimodal input data is organized in a structured combination manner and input into a large multimodal model. The features are fused through a visual-text cross-attention mechanism to complete the semantic analysis of hazards. S32. Based on the results of hazard target detection, extract the category, spatial coordinates, orientation attributes and spatial adjacency relationships between gas equipment components, and construct a two-dimensional topological relationship model to describe the structural relationships of the equipment. S33. Combine the standard equipment structure logic rules preset by the national standard to perform consistency verification on the two-dimensional topology relationship model, and construct a dynamic threshold self-calibration mechanism based on the distribution of test results or historical statistical data to adaptively adjust the verification threshold in order to identify potential spatial contradictions in component installation, connection and structure. S34. Based on the pre-trained knowledge of the multimodal large model, and combined with the description of the hazard features in the prompt words, zero-sample identification of untrained hazard types is achieved. The semantic analysis results of the hazard and the spatial contradiction identification results are then integrated to generate comprehensive hazard analysis data and output to step S4.
[0010] As a further improvement to this technical solution, the process of fusion of structured multimodal input and visual-text cross-attention in S31 includes the following steps: S31.1. Structure and stitch together on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert tips according to the gas inspection scenario to form multimodal input data; S31.2 Input the multimodal input data into the visual encoder and the text encoder respectively to generate the corresponding visual feature sequence and text feature sequence; S31.3. Employ a visual-text cross-attention mechanism to achieve bidirectional association and fusion of visual and textual features; S31.4. Based on bidirectional cross-modal fusion features, complete the semantic parsing, type determination and hazard degree quantification of gas hazards.
[0011] As a further improvement to this technical solution, the process of constructing the two-dimensional topological relationship model of S32 includes the following steps: S32.1 Extract the category and center coordinates of gas equipment components from the hazard target detection results. Installation direction angle and the dimensions of the outer frame; S32.2, Using Euclidean distance Computing components With components The two-dimensional spatial spacing is used to determine whether there is a spatial adjacency relationship between components when the spacing is less than the preset adjacency threshold. S32.3 Construct a two-dimensional topology graph using components as nodes and spatial adjacency relationships as edges, combined with component category and direction attributes; S32.4. Annotate the two-dimensional topology with structured attributes to form a two-dimensional topology model of gas equipment that is compatible with national standard verification.
[0012] As a further improvement to this technical solution, the process of verifying the national standard structure consistency and dynamic threshold self-calibration of S33 includes the following steps: S33.1. Perform edge set matching between the actual topology graph and the national standard topology graph, based on accuracy. Recall rate Calculate F1 matching degree Among them, accuracy Recall represents the proportion of correctly matched edges in the actual topological edge set. This indicates the proportion of correctly matched edges in the national standard topology edge set; S33.2, Based on the average of historical matching data with standard deviation ,use The principle is to dynamically calibrate the verification threshold, and the calculation formula is: ; in: This represents the mean of the historical structural matching degree. The standard deviation of the historical structural matching degree. The threshold for verifying the structural matching degree after adaptive calibration; S33.3, when At that time, it was determined that there were potential problems with component installation, connection, and structural space.
[0013] As a further improvement to this technical solution, the zero-sample identification and comprehensive hazard analysis data generation process of S34 includes the following steps: S34.1 Based on the pre-training knowledge of multimodal large models, text feature vectors are used. With visual feature vectors Cosine similarity between Complete the matching calculation between text features and visual features; S34.2. Complete zero-sample identification of untrained hazard types based on similarity thresholds; S34.3, Using weighted confidence levels The confidence scores of the three results—hidden danger semantic analysis, spatial contradiction identification, and zero-shot identification—are normalized and fused. Used to characterize the final credibility of the comprehensive hazard analysis results; S34.4 Generate comprehensive hazard analysis data based on the final fusion confidence level and output it to step S4.
[0014] As a further improvement to this technical solution, the process of linking the structured result output of S4 with the process includes the following steps: S4.1. Perform unified structured processing on the comprehensive hazard analysis data, classify and archive the data according to hazard type, hazard location and risk level, standardize the data format and complete the equipment information associated with the hazard; S4.2 Based on the structured hazard data, generate a data set including hazard details, risk level, cause analysis, treatment recommendations, and final confidence level. A structured gas hazard identification report, which includes hazard details such as the equipment component where the hazard is located, its specific location, and the manifestation of the hazard. S4.3. Associate and match the risk level in the structured gas hazard identification report with the gas operation and maintenance parameters. Based on the preset control rules corresponding to different risk levels, dynamically adjust the gas transmission pressure, inspection frequency and other gas operation and maintenance parameters to achieve linkage response between hazard identification and operation and maintenance processes.
[0015] The second objective of this invention is to provide a gas hazard identification system based on a multimodal large model, used to execute the aforementioned gas hazard identification method based on a multimodal large model, including: The image acquisition and quality detection unit is used to acquire on-site inspection images of gas pipelines and equipment, calculate the gradient variance using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a re-acquisition prompt for low-quality images, and output the sharpness score and image quality level. The hidden danger target detection unit is used to input qualified inspection images into the RF-DETR target detection model, extract multi-scale visual features through a convolutional backbone network, complete global modeling using a Transformer with deformable attention, and output the category and bounding box detection results of equipment components and hidden dangers through an ensemble prediction mechanism. A multimodal semantic understanding and spatial contradiction identification unit is used to complete multimodal data fusion, gas equipment structural hazard identification, and zero-sample hazard detection, generating comprehensive hazard analysis data. It includes a multimodal input and cross-modal fusion module, a topology modeling and national standard structure verification module, and a zero-sample identification and comprehensive data generation module, wherein: The multimodal input and cross-modal fusion module is used to fuse on-site inspection images, hazard target detection results, clarity scores, image quality levels and gas expert prompts to construct structured multimodal input data. It completes the association and fusion of visual and text features and hazard semantic analysis through a visual-text cross-attention mechanism. The topology modeling and national standard structure verification module extracts component features based on the detection results to construct a two-dimensional topology model, and combines it with national standard equipment structure rules to complete consistency verification. A dynamic threshold self-calibration mechanism identifies potential spatial conflicts related to component installation, connection, and structural elements. The zero-shot identification and comprehensive data generation module achieves zero-shot identification of untrained hazard types based on the pre-trained knowledge of a multimodal large model, and generates and outputs comprehensive hazard analysis data by integrating semantic analysis and spatial contradiction identification results. The structured result output and process linkage unit is used to perform structured processing on the comprehensive hidden danger analysis data, output a structured gas hidden danger identification report containing hidden danger details, risk level, cause analysis, treatment suggestions and confidence level, and adjust gas operation and maintenance parameters according to the report.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention can perform image quality detection and hidden danger detection based on on-site inspection images of small and medium-sized gas equipment such as gas pipeline interfaces, valves, and flanges. By constructing a two-dimensional topological relationship model of the equipment and combining it with the national standard structural logic rules for consistency verification, it can effectively identify spatial contradictions and hidden dangers in the installation, connection, and structure of gas equipment components. This makes up for the shortcomings of existing technologies that cannot detect the spatial relationship of the physical installation and connection of gas equipment based on on-site equipment images. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the method steps of the present invention; Figure 2 This is a schematic diagram of the method steps for detecting hidden danger targets in this invention; Figure 3 Overall architecture diagram of the intelligent gas hazard identification system of the present invention; Figure 4 The overall processing flowchart of the intelligent gas hazard identification method of the present invention; Figure 5 Architecture diagram of the RF-DETR target detection network in this invention; Figure 6 A schematic diagram illustrating the principle of multimodal feature fusion in this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] like Figures 1-6 As shown, this embodiment provides a gas hazard identification method based on a multimodal large model, including: S1. Image Acquisition and Quality Inspection: Acquire on-site inspection images of gas pipelines and related equipment, calculate the gradient variance of the on-site inspection images using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a reacquisition prompt for low-quality images, and output the sharpness score and image quality level to step S3. In this embodiment, the image acquisition and quality detection in S1 includes the following steps: S11. Collect on-site inspection images of gas pipelines, valves, pressure regulators, flanges and pipeline interfaces, and complete image format unification and encoding conversion processing; Specifically, using common preprocessing methods in the field of digital image processing, the acquired raw images undergo format unification, encoding conversion, and grayscale conversion: Images output from different acquisition devices are uniformly converted into a standard bitmap format to complete encoding standardization and eliminate data format differences; Convert a color image to a grayscale image to obtain a single-channel grayscale image. This provides standard input for subsequent convolution calculations.
[0020] S12. The Laplacian operator is used to perform convolution operation on the preprocessed on-site inspection image to calculate the gradient variance data of the entire image. Specifically, a 3×3 standard Laplacian convolution kernel is used to perform convolution operations on the grayscale on-site inspection image obtained by S11, extracting image gradient features and calculating the gradient variance of the entire on-site inspection image. The specific implementation is as follows: 3×3 standard Laplacian convolution kernel: ; Laplacian convolution operation formula: ; in: : Grayscale on-site inspection image in pixel coordinates Laplacian convolution response value at the location; : Grayscale on-site inspection image in pixel coordinates The pixel grayscale value at that location; : 3×3 standard Laplacian convolution kernel; : The horizontal and vertical coordinates of the pixels in the on-site inspection image; : Kernel traversal offset, with values of -1, 0, and 1.
[0021] Convolution response mean calculation: ; in: The mean value of the Laplacian convolution response of all pixels in the entire on-site inspection image; : Pixel width of the on-site inspection image; : Pixel height of the on-site inspection image.
[0022] Gradient variance calculation: ; in: The gradient variance of the entire on-site inspection image is used to characterize the image's clarity.
[0023] S13. Perform linear normalization on the calculated gradient variance data, mapping it to a numerical range of 0 to 100, to generate an image sharpness score; the linear normalization calculation formula is as follows: ; : Clarity score of on-site inspection images, with a value range of [0,100]; : The gradient variance calculation result of the current on-site inspection images; The minimum value of the gradient variance of on-site inspection images in a gas inspection scenario; : The maximum value of the gradient variance of on-site inspection images in the gas inspection scenario.
[0024] S14. Divide the sharpness score according to the preset three-level threshold: a score of 70 or above is judged as sharp, a score between 30 and 70 is judged as medium, and a score less than 30 is judged as blurry. Specifically, the sharpness is scored according to a preset three-level threshold. Determine the level: when At that time, the on-site inspection images were determined to be of a high clarity level; when At that time, the on-site inspection images were determined to be of medium quality. when At that time, the on-site inspection image was determined to be blurry.
[0025] S15. When an image is determined to be blurry, an image re-acquisition prompt message is generated and the subsequent detection process of the current image is terminated. The steps of hazard target detection, multimodal semantic understanding and spatial contradiction recognition are no longer executed.
[0026] S16. Output the sharpness scores and image quality levels corresponding to the clear and medium-level images to step S3 as input data for multimodal data fusion and hazard analysis.
[0027] It should be added that this embodiment uses the Laplacian operator to quantify image sharpness, generates a standard sharpness score by linear normalization, and completes automatic image quality grading by combining a three-level threshold. It blocks the process for blurry images and outputs a re-acquisition prompt, and connects the quality data of valid images to the subsequent multimodal analysis stage. This ensures the quality of input data from the source, reduces invalid calculations, and improves the stability and reliability of the overall hidden danger identification process.
[0028] S2, Hazard Target Detection: Input qualified on-site inspection images into the Transformer-based target detection model, extract multi-scale features through a convolutional backbone network, perform global modeling through a Transformer with deformable attention mechanism, and output the category information of gas equipment components and hazards and bounding box detection results using an ensemble prediction mechanism, and output the detection results to step S3; In this embodiment, the hazard target detection in S2 includes the following steps: S21. The on-site inspection image determined to be clear and of medium quality in step S1 is input into the target detection model based on Transformer, wherein the target detection model based on Transformer is the RF-DETR model. Specifically, the Transformer-based target detection model used in this embodiment is the RF-DETR (ResNet-FeatureDeformableDETR) model, whose network structure is as follows: Figure 5 As shown, this is an end-to-end Transformer target detection model optimized for gas inspection scenarios. The model takes the standardized, uniformly formatted on-site inspection image (completed in step S1) as a fixed-size input. It incorporates a ResNet-50 convolutional backbone network, multi-scale feature pyramids (P3–P6), a Transformer encoder with deformable attention, a Transformer decoder, and an ensemble prediction head. It requires no non-maximum suppression (NMS) post-processing throughout the process and can directly achieve one-time detection output of equipment components such as gas pipelines, valves, pressure regulators, flanges, and pipe interfaces, as well as corresponding hazards. Addressing the characteristics of gas inspection images—large differences in target scale, a high proportion of small-target hazards, and complex background interference—this model improves the detection accuracy of minute hazards and densely packed components through multi-scale feature perception and deformable attention global modeling. The ensemble prediction mechanism ensures that only a unique detection result is output for the same equipment component / hazard, avoiding duplicate detections and false detections. This provides standardized, high-confidence detection data for subsequent multimodal semantic understanding and spatial contradiction recognition.
[0029] S22. Extract multi-scale visual features from on-site inspection images using a convolutional backbone network; Specifically, the convolutional backbone network built into the RF-DETR model is used to perform layered convolution and downsampling processing on the input on-site inspection images, extracting multi-scale visual features of the images layer by layer, and generating multi-scale feature maps adapted to the detection of targets of different sizes. The multi-scale feature maps cover different receptive field ranges of large, medium and small, and completely preserve the core visual features such as texture, outline and location of gas equipment components and hidden dangers, providing a multi-dimensional feature foundation for subsequent global feature modeling.
[0030] S23. Global feature modeling of the image is completed using a Transformer with deformable attention mechanism; Specifically, the multi-scale visual features extracted by S22 are input into a Transformer structure equipped with a deformable attention mechanism. Global feature modeling of the image is completed through deformable attention, achieving efficient association between local image features and global contextual information. The specific implementation is as follows: Deformable attention feature aggregation operation formula: ; in: This represents the global modeling features output after deformable attention operations; This represents the query vector in the Transformer attention module; Indicates the coordinates of a reference point on the feature map; This represents the input multi-scale visual feature vector; Indicates the number of sampling points for the deformable attention mechanism; Indicates the first Attention weight coefficients corresponding to each sampling point; Indicates the first The coordinate offset of each sampling point relative to the reference point; Indicates the first The characteristic modulation coefficients corresponding to each sampling point.
[0031] After the above deformable attention operation, the association and modeling of global image features are completed, and the detection feature sequence fused with global context information is output.
[0032] S24. Use an ensemble prediction mechanism to output the category information and bounding box detection results of gas equipment components and potential hazards; Specifically, an ensemble prediction mechanism is used to predict targets based on the detection features after global modeling. The optimal matching between the target and the prediction result is achieved through bipartite graph matching, and the detection results of gas equipment components and hidden dangers are directly output. The detection results include category information and bounding box coordinate information. The category information covers equipment components such as gas pipelines, valves, pressure regulators, flanges, and pipeline interfaces, as well as the corresponding hidden danger types. The bounding box represents the position range of the target in the image in a standardized coordinate form. The ensemble prediction mechanism can ensure that only a unique detection result is output for the same target.
[0033] S25. Output the category information and bounding box detection results to step S3.
[0034] Specifically, the gas equipment components and hazard category information and boundary box detection results obtained in S24 are standardized and organized to form a unified format of detection result data, which is directly output to the subsequent multimodal semantic understanding and spatial contradiction identification stages as the core input data for multimodal data fusion and comprehensive hazard analysis.
[0035] It should be added that this embodiment uses the RF-DETR model to complete the detection of potential hazards, realizes multi-scale feature extraction through a convolutional backbone network, improves the efficiency of global feature modeling by combining a deformable attention mechanism, and ensures the uniqueness of detection results by relying on an ensemble prediction mechanism. The entire process from feature extraction and global modeling to result output is optimized, providing accurate and standardized target detection data for subsequent multimodal hazard analysis, and improving the detection stability and reliability of the overall solution.
[0036] S3. Multimodal Semantic Understanding and Spatial Conflict Identification: Multimodal input is constructed by fusing on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. The hazard semantic analysis is completed through the visual-text cross-attention fusion feature of the multimodal large model. Based on the detection results, the categories, coordinates, directions, and spatial adjacency relationships of gas equipment components are extracted and a two-dimensional topological relationship model is constructed. Consistency verification is performed by combining the standard equipment structure logic rules preset by the national standard. The verification threshold is adaptively adjusted through a dynamic threshold self-calibration mechanism to identify spatial conflict hazards related to component installation, connection, and structure. At the same time, zero-sample identification of untrained hazard types is achieved based on the pre-trained knowledge of the multimodal large model. Comprehensive hazard analysis data is generated and output to step S4. In this embodiment, the multimodal semantic understanding and spatial contradiction recognition in S3 includes the following steps: Furthermore, in step S31, multimodal input data is constructed by fusing on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. This multimodal input data is organized using a structured combination method and input into a large multimodal model. The semantic analysis of the hazard is completed by fusing features through a visual-text cross-attention mechanism. The process of fusing the structured multimodal input with visual-text cross-attention in step S31 includes the following steps: S31.1. The on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts are structurally stitched together according to the gas inspection scenario to form multimodal input data. At the same time, the fixed order of data fields is maintained during the stitching process. The on-site inspection images serve as the basic visual data, the hazard target detection results, clarity scores, and image quality levels serve as quantitative auxiliary data, and the gas expert prompts serve as professional semantic guidance data. After the various data are combined according to the preset fields, a structured data carrier that adapts to the input standard of the multimodal large model is formed, with no data redundancy and format disorder, providing standardized input for subsequent feature encoding and fusion.
[0037] S31.2 Input the multimodal input data into the visual encoder and the text encoder respectively, and generate the corresponding visual feature sequence and text feature sequence through the feature extraction operation of the encoder; Specifically, the visual encoder extracts global and local features from on-site inspection images and outputs a fixed-dimensional visual feature sequence; the text encoder extracts semantic features from text and numerical data such as hazard target detection results, clarity scores, image quality levels, and gas expert tips and outputs a text feature sequence that matches the dimension of the visual feature sequence.
[0038] The formula for generating visual feature sequences is as follows: ; in: This represents the sequence of visual features output by the visual encoder. Represents the first to the second visual feature sequence One visual feature vector; This represents the total number of vectors representing the visual feature sequence.
[0039] The formula for generating text feature sequences is as follows: ; in: This represents the text feature sequence output by the text encoder; Represents the first to the second feature sequence in the text. One text feature vector; This represents the total number of vectors representing the text feature sequences.
[0040] S31.3. Employ a visual-text cross-attention mechanism to achieve bidirectional association and fusion of visual and textual features; Specifically, a visual-text cross-attention mechanism is employed to achieve bidirectional association and fusion of visual feature sequences and text feature sequences. Cross-modal feature interaction is achieved through text-guided visual attention and visual-guided text attention, respectively, ultimately resulting in fused cross-modal features. This bidirectional fusion method enhances the association accuracy between visual and text data and is the core improvement of this embodiment. The specific implementation is as follows: Text-guided visual attention computation: ; in: This indicates the attention output result of text features guiding visual features; This represents a query vector generated from a sequence of text features. This represents a key vector generated from a sequence of visual features; This represents a value vector generated from a sequence of visual features. This represents the dimension of the feature vector in the attention mechanism; This represents the normalized exponential activation function, used to normalize the attention weights to the 0-1 range.
[0041] Visually Guided Text Attention Calculation: ; in: This represents the attention output result where visual features guide text features; This represents a query vector generated from a sequence of visual features; This represents a key vector generated from a sequence of text features; This represents a value vector generated from a sequence of text features. This represents the dimension of the feature vector in the attention mechanism; This represents the normalized exponential activation function.
[0042] Bidirectional cross-modal fusion feature generation: ; in: This represents the cross-modal fusion feature after visual-text bidirectional cross-attention fusion; This indicates a feature vector concatenation operation, which concatenates the two sets of attention outputs into a single fused feature according to the feature dimension.
[0043] Furthermore, this embodiment also provides the following calculation example: Define the dimension of the attention feature vector Text query vector With visual key vector The dot product result is 256. Substituting this into the text-guided visual attention formula, we can calculate: First calculate After normalization using the Softmax function, standardized attention weights are obtained. These weights are then compared with the visual value vector. Weighted calculation yields the text-guided visual attention result. ; The same computational logic was used to complete the visually guided text attention results. The operation integrates the two sets of attention results through concatenation, ultimately yielding a bidirectional cross-modal fusion feature. .
[0044] It should be added that this embodiment abandons the traditional one-way cross-modal feature fusion method and adopts visual-text bidirectional cross-attention to achieve mutual guidance and association between visual data and text data, so as to deeply bind image features with gas professional text semantics, thereby improving the accuracy and completeness of extracting hidden danger semantic features.
[0045] S31.4. Based on bidirectional cross-modal fusion features, complete the semantic parsing, type determination and hazard degree quantification of gas hazards.
[0046] Specifically, relying on the bidirectional cross-modal fusion features generated by S31.3, the full-dimensional semantic analysis, accurate type determination and quantitative hazard assessment of gas hazards are completed within the multimodal large model. The entire process relies on the visual-text association information of the fusion features, and standardized analysis can be completed without additional manual intervention.
[0047] In the semantic parsing stage, the fusion feature integrates the visual details of the on-site inspection images, the location information of the hazard target detection results, the auxiliary information of the image quality level and clarity score, and the professional logic of the gas expert's prompts, to fully restore the actual state of the gas equipment and the manifestation of the hazard, forming a concrete semantic description of the hazard.
[0048] In the type determination stage, based on cross-modal fusion features, various hidden danger feature databases in gas inspection scenarios are matched to accurately distinguish between equipment component abnormalities and hidden danger types, ensuring that the determination results are highly consistent with the actual situation on site.
[0049] In the hazard quantification stage, the hazard hazard level is standardized and quantified by combining the location, manifestation, equipment correlation, and the impact of image quality on the analysis confidence level, and outputting quantitative results that can be directly used for subsequent hazard classification and disposal.
[0050] After the above processing, a complete semantic analysis result of the hidden danger is obtained. This result serves as the core data, providing solid semantic support for subsequent spatial contradiction identification, zero-sample hidden danger identification, and comprehensive hidden danger analysis.
[0051] Further, in step S32, based on the hazard target detection results, the categories, spatial coordinates, directional attributes, and spatial adjacency relationships between gas equipment components are extracted to construct a two-dimensional topological relationship model describing the structural relationships of the equipment. The process of constructing the two-dimensional topological relationship model in step S32 includes the following steps: S32.1 Extract the category and center coordinates of gas equipment components from the hazard target detection results. Installation direction angle and the dimensions of the outer frame; Specifically, the categories and center coordinates of gas equipment components are extracted from the hazard detection results. Installation direction angle The structured features of the equipment components are extracted using the outer frame size parameters, providing basic data for subsequent topology modeling.
[0052] Category information: Covers all types of gas equipment components, including gas pipelines, valves, pressure regulators, flanges, and pipe interfaces, and is directly taken from the category output of the target detection results; center coordinates The coordinates of the bounding box of the component in the target detection result are calculated using the following formula: ; ; in: Indicates the first The center horizontal coordinate of each gas equipment component; Indicates the first The center ordinate of each gas equipment component; , Indicates the first The coordinates of the top-left corner of the bounding box of each component; , Indicates the first The coordinates of the bottom right corner of the outer bounding box of each component.
[0053] Installation direction angle : Characterizes the installation orientation of the component. For directional components such as pipes and flanges, it is calculated by the angle between the long side of the outer frame and the horizontal axis of the image. Outline frame size parameters: including the width of the component's outer bounding box. ,high The calculation formula is: ; ; in: Indicates the first The width of the outer bounding box of each component; Indicates the first The height of the outer bounding box of each component.
[0054] S32.2, Using Euclidean distance Computing components With components The two-dimensional spatial spacing is used to determine whether there is a spatial adjacency relationship between components when the spacing is less than the preset adjacency threshold. Specifically, Euclidean distance is used. Computing components With components The two-dimensional spatial spacing is used to determine whether a spatial adjacency relationship exists between components when the spacing is less than a preset adjacency threshold, thus providing a basis for edge construction of the topology graph. The specific implementation is as follows: Euclidean distance calculation formula: ; in: Representation Component With components The two-dimensional Euclidean distance between them; , Representation Component The center coordinates; , Representation Component The center coordinates.
[0055] Adjacency determination rules: ; in: This indicates a preset adjacency threshold, which is preset based on the resolution of the gas inspection image and the standard size of the equipment components. It is used to determine whether there is a physical connection or structural association between components.
[0056] It should be added that this embodiment uses Euclidean distance to quantify the spatial positional relationship between components and uses a preset threshold to automatically determine the adjacency relationship, replacing manual subjective judgment, providing objective adjacency relationship data for topology modeling, and ensuring the accuracy of the topology structure.
[0057] S32.3 Construct a two-dimensional topology graph using components as nodes and spatial adjacency relationships as edges, combined with component category and direction attributes; Specifically, using gas equipment components as nodes and spatial adjacency relationships between components as edges, and combining component category and orientation attributes, a two-dimensional topology graph is constructed to represent the structural relationships of the equipment, thus completing the basic construction of the topology. The specific implementation is as follows: Node definition: Each gas equipment component is associated with a node in the topology graph. Node attributes include component category, center coordinates, installation direction angle, and outer frame size parameters. Edge definition: A pair of components determined to have a spatial adjacency relationship is represented by an edge in the topology graph. The edge attributes include the identifiers of the two nodes and the Euclidean distance. ; Topology graph structure: Constructed in the form of an undirected graph, with edges representing physical connections and structural relationships between nodes, fully restoring the actual structural layout of the gas equipment site.
[0058] S32.4. Annotate the two-dimensional topology with structured attributes to form a two-dimensional topology model of gas equipment adapted for national standard verification, providing a standardized structural data carrier for subsequent national standard conformity verification. Among these: Structured attribute annotation: Add structured attributes such as device number, installation location, and scene information to the nodes and edges of the topology graph to improve the information dimension of the two-dimensional topology relationship model; National Standard Adaptation and Optimization: Based on the structural logic rules of relevant national standards for gas equipment installation, the node attributes and edge relationships of the topology graph are standardized and regulated to ensure that the structural logic of the topology model is fully adapted to the national standard requirements, and can be directly used for subsequent equipment structure compliance verification.
[0059] Furthermore, S33, combined with the standard equipment structure logic rules preset by the national standard, performs consistency verification on the two-dimensional topological relationship model, and constructs a dynamic threshold self-calibration mechanism based on the distribution of test results or historical statistical data to adaptively adjust the verification threshold in order to identify potential spatial contradictions in component installation, connection and structure. Specifically, the standard equipment structure logic rules preset by the national standard can be formulated based on the mandatory clauses in GB50028-2020 "Urban Gas Design Standard" regarding the installation spacing, connection method, orientation requirements, and structural matching relationships of gas pipelines, valves, pressure regulators, flanges, and pipeline interfaces. These include: coaxial connection requirements between gas pipelines and flanges, flow-direction installation requirements between valves and pipelines, vertical assembly requirements between pressure regulators and pipeline interfaces, minimum safety distance requirements between components, and sealing connection structure requirements for pipeline interfaces. The aforementioned national standard clauses clearly define the standard spatial topological structure relationships of gas equipment components, providing a legal technical basis for the consistency verification in this step.
[0060] The process of verifying the national standard structural consistency and performing dynamic threshold self-calibration of S33 includes the following steps: S33.1. Perform edge set matching between the actual topology graph and the national standard topology graph, based on accuracy. Recall rate Calculate F1 matching degree Among them, accuracy Recall represents the proportion of correctly matched edges in the actual topological edge set. This indicates the proportion of correctly matched edges in the national standard topology edge set; Specifically, accuracy Calculation formula: ; in: Precision rate represents the proportion of correctly matched edges in the actual topological edge set; This represents the true instance, i.e., the number of edges in the actual topological edge set that correctly match the national standard topological edge set; This represents a false positive, i.e., the number of edges in the actual topological edge set that do not match the national standard topological edge set.
[0061] Recall rate Calculation formula: ; in: Recall rate represents the proportion of correctly matched edges in the national standard topology edge set; This represents the true instance, i.e., the number of edges that correctly match the actual topological edge set in the national standard topological edge set. This indicates a false negative, which is the number of edges in the national standard topology edge set that are not matched by the actual topology edge set.
[0062] F1 matching degree Calculation formula: ; in: The F1 matching degree is used to comprehensively quantify the consistency between the actual topology and the national standard structure. The value range is [0,1], and the closer the value is to 1, the higher the structural consistency.
[0063] Furthermore, this embodiment also provides the following calculation example: The actual topology edge set of a certain gas equipment contains 8 edges, while the national standard topology edge set contains 7 edges; the number of correctly matched edges is... The number of incorrectly matched edges in the actual topology The number of unmatched edges in the national standard .
[0064] Substitute into the precision formula: ; Substitute into the recall formula: ; Substitute into the F1 matching formula: This means that the actual topology of the device matches the F1 degree of the national standard structure by 0.80.
[0065] S33.2, Based on the average of historical matching data for the same station, equipment model, and inspection scenario. with standard deviation ,use The principle is to dynamically calibrate the verification threshold, and the calculation formula is: ; in: This represents the mean of the historical structural matching degree. The standard deviation of the historical structural matching degree. The threshold for verifying the structural matching degree after adaptive calibration; Specifically, historical average matching degree The calculation formula is as follows: ; in: This represents the mean of the historical structural matching degree, which is the arithmetic mean of the F1 matching degree of multiple historical gas equipment topologies. This indicates the number of samples in the historical matching data; Indicates the first F1 matching degree of the historical gas equipment topology.
[0066] Historical matching standard deviation The calculation formula is as follows: ; in: The standard deviation of historical structural matching degree is used to characterize the dispersion of historical matching degree data; This indicates the number of samples in the historical matching data; Indicates the first F1 matching degree of the historical gas equipment topology.
[0067] This embodiment also provides the following calculation example: Take 100 sets of historical gas equipment topology data of the same type and model at the same gas station, and calculate the average historical matching degree. Historical matching standard deviation ; Substitute into the dynamic threshold formula: That is, the structural matching degree verification threshold after adaptive calibration is 0.83.
[0068] It should be added that this embodiment abandons the traditional fixed threshold verification method and adopts a 3-valued method based on the statistical distribution of historical matching data. The principle is to construct a dynamic threshold self-calibration mechanism, which can adaptively adjust the threshold according to the historical verification data of different sites and equipment, thereby improving the adaptability and accuracy of structural verification and avoiding misjudgment and omission caused by fixed thresholds.
[0069] S33.3, when At that time, it was determined that there were potential problems with component installation, connection, and structural space.
[0070] Specifically, the F1 matching degree calculated based on S33.1 The dynamic verification threshold obtained from S33.2 calibration Determining potential conflicts and hidden dangers in execution space: when At that time, it was determined that the gas equipment had potential spatial conflicts related to component installation, connection, and structural issues. when At that time, it was determined that the structure of the gas equipment met the national standards and there were no potential spatial conflicts.
[0071] It is understandable that this embodiment achieves automated and adaptive verification of the structural compliance of gas equipment through the entire process of topological edge set matching, F1 matching quantification, 3σ dynamic threshold calibration, and automatic hazard judgment. This solves the industry pain points of low efficiency and large subjective error of traditional manual verification, as well as poor adaptability and high misjudgment rate of fixed threshold verification.
[0072] Furthermore, in step S34, based on the pre-trained knowledge of the multimodal large model, and combined with the description of the hazard characteristics in the prompt words, zero-shot identification of untrained hazard types is achieved. The semantic analysis results of the hazard and the spatial contradiction identification results are then fused to generate comprehensive hazard analysis data, which is output to step S4. The zero-shot identification and comprehensive hazard analysis data generation process in step S34 includes the following steps: S34.1 Based on the pre-training knowledge of multimodal large models, text feature vectors are used. With visual feature vectors Cosine similarity between The matching calculation of text features and visual features is completed to provide a quantitative similarity basis for zero-shot recognition; the specific implementation is as follows: Cosine similarity calculation formula: ; in: The cosine similarity between the text feature vector and the visual feature vector is represented, with a value range of [−1, 1]. The closer the value is to 1, the higher the degree of matching between the text features and the visual features. This represents the text feature vector generated from the descriptions of potential hazards in the gas expert's warnings; This represents a visual feature vector generated from on-site inspection images and hazard target detection results. Represents text feature vectors With visual feature vectors The dot product operation is calculated using the following formula: ,in The dimension of the feature vector; The L2 norm of the text feature vector u is represented by the formula: ; The L2 norm of the visual feature vector v is represented by the formula: .
[0073] This embodiment also provides the following calculation example: Set the feature vector dimension d=3, text feature vector Visual feature vectors ; First, calculate the vector dot product: ; Calculate the L2 norm of the text feature vectors: ; Calculate the L2 norm of visual feature vectors: ; Substituting into the cosine similarity formula: The cosine similarity between the text features and the visual features is approximately 0.985, indicating a very high degree of matching.
[0074] S34.2. Complete zero-sample identification of untrained hazard types based on similarity thresholds; Specifically, zero-sample identification of untrained hazard types is completed based on the cosine similarity threshold, without the need to retrain the model for new hazard types. Generalized hazard identification is achieved by relying on the pre-trained knowledge of the multimodal large model and the guidance of prompt words.
[0075] Zero-shot identification criteria: ; in: This represents the preset cosine similarity threshold, which is preset according to the hazard identification requirements of the gas inspection scenario. It is used to determine whether the matching of text features and visual features meets the hazard identification requirements. This represents the text-visual feature cosine similarity calculated by S34.1.
[0076] It should be added that this embodiment abandons the limitation of traditional deep learning relying on a large number of labeled samples. By matching the pre-trained knowledge of a multimodal large model with cosine similarity, it achieves zero-sample identification of untrained hazard types, solves the industry pain point of difficulty in identifying new hazards and small-sample hazards in the gas industry, and improves the generalization ability of hazard identification.
[0077] S34.3, Using weighted confidence levels The confidence scores of the three results—hidden danger semantic analysis, spatial contradiction identification, and zero-shot identification—are normalized and fused. Used to characterize the final credibility of the comprehensive hazard analysis results; Specifically, the formula for calculating the weighted confidence fusion is as follows: ; in: This represents the final weighted confidence level of the comprehensive hazard analysis results, with a value range of [0,1]. The closer the value is to 1, the higher the confidence level of the comprehensive analysis results. The weight coefficients for the three results—hidden danger semantic analysis, spatial contradiction identification, and zero-shot identification—are respectively, satisfying the following conditions: The weighting coefficients can be adaptively adjusted according to the actual needs of the gas inspection scenario; The confidence level of the semantic analysis results of potential hazards is output by the semantic parsing stage; The confidence level of the spatial conflict identification result is output by the spatial conflict hazard assessment stage; This represents the confidence level of the zero-shot identification result, output by the zero-shot identification process, and is expressed as cosine similarity. .
[0078] This embodiment also provides the following calculation example: Set weight coefficients The confidence levels for each item are as follows: (Semantic analysis confidence) (Confidence level for spatial contradiction identification) (Zero-sample identification confidence, i.e., cosine similarity); Substitute into the weighted confidence formula: The final weighted confidence level of the comprehensive hazard analysis results is approximately 0.93.
[0079] S34.4 Generate comprehensive hazard analysis data based on the final fusion confidence level and output it to step S4.
[0080] Specifically, based on the final fusion confidence level The system integrates the results of semantic analysis of potential hazards, spatial contradiction identification, and zero-shot identification to generate standardized comprehensive hazard analysis data. This data is then output to subsequent structured result output and process linkage stages. Specifically: Comprehensive hazard analysis data includes: hazard type, hazard location, hazard manifestation, risk level, confidence levels of each analysis item, and final weighted confidence level. Structured information such as processing suggestions provides complete and traceable core data support for subsequent report generation and operation and maintenance parameter linkage.
[0081] S4. Structured Result Output and Process Linkage: The comprehensive hazard analysis data is processed in a unified structure to output a structured gas hazard identification report containing hazard details, risk level, cause analysis, treatment suggestions and confidence level. Gas operation and maintenance parameters are adjusted based on the structured gas hazard identification report.
[0082] In this embodiment, the process of linking the structured result output of S4 with the process includes the following steps: S4.1 The comprehensive hazard analysis data is processed in a unified structure, classified and archived according to hazard type, hazard location, and risk level, the data format is standardized, and the equipment information associated with the hazard is supplemented, providing a standardized data foundation for subsequent report generation and process linkage. Among these: Data standardization: The multi-source results of hidden danger semantic analysis, spatial contradiction identification, and zero-shot identification are standardized according to preset fields to unify the format, eliminate data heterogeneity, and ensure that the data fields, formats, and codes are completely consistent; Categorized archiving: Structured data is categorized and stored according to the type of hazard (such as pipeline corrosion, abnormal flange connection, valve sealing failure, etc.), the location of the hazard (such as site number, pipeline section number, equipment coordinates, etc.), and the risk level (high, medium, low) to facilitate quick retrieval and access; Information Completion: Complete the basic information of the equipment associated with the hidden danger, including equipment number, installation date, site, maintenance records, etc., to improve the full-dimensional information of the hidden danger data and provide support for cause analysis and treatment suggestions.
[0083] S4.2 Based on the structured hazard data, generate a data set including hazard details, risk level, cause analysis, treatment recommendations, and final confidence level. A structured gas hazard identification report, which includes hazard details such as the equipment component where the hazard is located, its specific location, and the manifestation of the hazard. Specifically, the structured gas hazard identification report includes: Details of the hidden danger: Identify the gas equipment components (such as pipelines, valves, flanges, etc.) where the hidden danger is located, the specific location (such as the middle section of pipeline #1 in the station, the flange interface, etc.), and the manifestation of the hidden danger (such as pipeline corrosion, reverse installation of flanges, poor valve sealing, etc.). The information is directly taken from the results of previous target detection, topology analysis, and semantic analysis. Risk level: Based on the quantification of the severity of hazards according to the comprehensive hazard analysis data, combined with the final confidence level. The risk levels are divided into high, medium, and low levels; among them The comprehensive weighted confidence level calculated by S34.3 has a value range of [0,1] and is used to characterize the final credibility of the hazard analysis results; Cause analysis: Based on the results of multimodal semantic analysis and combined with the basic information of the equipment, the causes of the hidden dangers are identified (such as aging of the anti-corrosion layer due to humid environment, abnormal flange connection due to improper installation, etc.). Recommendations: Based on gas industry standards and expert advice, we will generate targeted recommendations (such as rust removal and repair of the anti-corrosion layer within 7 days, reinstallation of flanges, replacement of seals, etc.). Final confidence level The comprehensive weighted confidence level calculated by S34.3 is included in the report to quantify the credibility of the hazard analysis results and provide data support for operation and maintenance decisions.
[0084] S4.3 Associate and match the risk levels in the structured gas hazard identification report with gas operation and maintenance parameters. Based on preset control rules corresponding to different risk levels, dynamically adjust gas delivery pressure, inspection frequency, and other gas operation and maintenance parameters to achieve a coordinated response between hazard identification and operation and maintenance processes. Specifically: Association matching: Establish a mapping relationship between risk level and gas operation and maintenance parameters, and match the high, medium and low risk levels in the report with operation and maintenance parameters such as gas transmission pressure, inspection frequency and maintenance priority. Preset control rules: Based on gas industry operation and maintenance standards, preset control rules for different risk levels are provided, as shown in the following example: High-risk hazards: Immediately reduce gas transmission pressure, increase inspection frequency to once a day, and prioritize maintenance work; Medium-risk hazards: Maintain gas transmission pressure, increase inspection frequency to once every 15 days, and arrange maintenance work according to plan; Low-risk hazards: Maintain the existing gas transmission pressure and inspection frequency, and conduct routine inspections and monitoring; Linked Execution: The system automatically matches control rules based on the risk level reported and issues control instructions to the gas dispatch and operation system to complete the dynamic adjustment of gas operation and maintenance parameters without manual intervention, thus realizing the automated linkage of operation and maintenance processes.
[0085] This embodiment also provides a gas hazard identification system based on a multimodal large model. The gas hazard identification method based on the aforementioned multimodal large model includes: The image acquisition and quality detection unit is used to acquire on-site inspection images of gas pipelines and equipment, calculate the gradient variance using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a re-acquisition prompt for low-quality images, and output the sharpness score and image quality level. The hidden danger target detection unit is used to input qualified inspection images into the RF-DETR target detection model, extract multi-scale visual features through a convolutional backbone network, complete global modeling using a Transformer with deformable attention, and output the category and bounding box detection results of equipment components and hidden dangers through an ensemble prediction mechanism. A multimodal semantic understanding and spatial contradiction identification unit is used to complete multimodal data fusion, gas equipment structural hazard identification, and zero-sample hazard detection, generating comprehensive hazard analysis data. It includes a multimodal input and cross-modal fusion module, a topology modeling and national standard structure verification module, and a zero-sample identification and comprehensive data generation module, wherein: The multimodal input and cross-modal fusion module is used to fuse on-site inspection images, hazard target detection results, clarity scores, image quality levels and gas expert prompts to construct structured multimodal input data. It completes the association and fusion of visual and text features and hazard semantic analysis through a visual-text cross-attention mechanism. The topology modeling and national standard structure verification module extracts component features based on the detection results to construct a two-dimensional topology model, and combines it with national standard equipment structure rules to complete consistency verification. A dynamic threshold self-calibration mechanism identifies potential spatial conflicts related to component installation, connection, and structural elements. The zero-shot identification and comprehensive data generation module achieves zero-shot identification of untrained hazard types based on the pre-trained knowledge of a multimodal large model, and generates and outputs comprehensive hazard analysis data by integrating semantic analysis and spatial contradiction identification results. The structured result output and process linkage unit is used to perform structured processing on the comprehensive hidden danger analysis data, output a structured gas hidden danger identification report containing hidden danger details, risk level, cause analysis, treatment suggestions and confidence level, and adjust gas operation and maintenance parameters according to the report.
[0086] Those skilled in the art will understand that the process of implementing all or part of the steps of the above embodiments can be carried out by hardware or by a program instructing the relevant hardware.
[0087] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A method for identifying gas hazard based on a multimodal large model, characterized in that, Includes the following steps: S1. Image Acquisition and Quality Inspection: Acquire on-site inspection images of gas pipelines and related equipment, calculate the gradient variance of the on-site inspection images using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a reacquisition prompt for low-quality images, and output the sharpness score and image quality level to step S3. S2, Hazard Target Detection: Input qualified on-site inspection images into the Transformer-based target detection model, extract multi-scale features through a convolutional backbone network, perform global modeling through a Transformer with deformable attention mechanism, and output the category information of gas equipment components and hazards and bounding box detection results using an ensemble prediction mechanism, and output the detection results to step S3; S3. Multimodal Semantic Understanding and Spatial Conflict Identification: Multimodal input is constructed by fusing on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. The hazard semantic analysis is completed through the visual-text cross-attention fusion feature of the multimodal large model. Based on the detection results, the categories, coordinates, directions, and spatial adjacency relationships of gas equipment components are extracted and a two-dimensional topological relationship model is constructed. Consistency verification is performed by combining the standard equipment structure logic rules preset by the national standard. The verification threshold is adaptively adjusted through a dynamic threshold self-calibration mechanism to identify spatial conflict hazards related to component installation, connection, and structure. At the same time, zero-sample identification of untrained hazard types is achieved based on the pre-trained knowledge of the multimodal large model. Comprehensive hazard analysis data is generated and output to step S4. S4. Structured Result Output and Process Linkage: The comprehensive hazard analysis data is processed in a unified structure to output a structured gas hazard identification report containing hazard details, risk level, cause analysis, treatment suggestions and confidence level. Gas operation and maintenance parameters are adjusted based on the structured gas hazard identification report.
2. The gas hazard identification method based on a multimodal large model according to claim 1, characterized in that, The image acquisition and quality detection in S1 includes the following steps: S11. Collect on-site inspection images of gas pipelines, valves, pressure regulators, flanges and pipeline interfaces, and complete image format unification and encoding conversion processing; S12. The Laplacian operator is used to perform convolution operation on the preprocessed on-site inspection image to calculate the gradient variance data of the entire image. S13. Perform linear normalization on the calculated gradient variance data, map it to the numerical range of 0 to 100, and generate an image sharpness score. S14. Divide the sharpness score according to the preset three-level threshold: a score of 70 or above is judged as sharp, a score between 30 and 70 is judged as medium, and a score less than 30 is judged as blurry. S15. When an image is determined to be blurry, generate an image re-acquisition prompt and terminate the subsequent detection process of the current image. S16. Output the sharpness score and image quality level corresponding to the clear and medium level images to step S3.
3. The gas hazard identification method based on a multimodal large model according to claim 1, characterized in that, The hazard detection in S2 includes the following steps: S21. The on-site inspection image determined to be clear and of medium quality in step S1 is input into the target detection model based on Transformer, wherein the target detection model based on Transformer is the RF-DETR model. S22. Extract multi-scale visual features from on-site inspection images using a convolutional backbone network; S23. Global feature modeling of the image is completed using a Transformer with deformable attention mechanism; S24. Use an ensemble prediction mechanism to output the category information and bounding box detection results of gas equipment components and potential hazards; S25. Output the category information and bounding box detection results to step S3.
4. The gas hazard identification method based on a multimodal large model according to claim 1, characterized in that, The multimodal semantic understanding and spatial contradiction recognition in S3 include the following steps: S31. Multimodal input data is constructed by integrating on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert prompts. The multimodal input data is organized in a structured combination manner and input into a large multimodal model. The features are fused through a visual-text cross-attention mechanism to complete the semantic analysis of hazards. S32. Based on the results of hazard target detection, extract the category, spatial coordinates, orientation attributes and spatial adjacency relationships between gas equipment components, and construct a two-dimensional topological relationship model to describe the structural relationships of the equipment. S33. Combine the standard equipment structure logic rules preset by the national standard to perform consistency verification on the two-dimensional topology relationship model, and construct a dynamic threshold self-calibration mechanism based on the distribution of test results or historical statistical data to adaptively adjust the verification threshold in order to identify potential spatial contradictions in component installation, connection and structure. S34. Based on the pre-trained knowledge of the multimodal large model, and combined with the description of the hazard features in the prompt words, zero-sample identification of untrained hazard types is achieved. The semantic analysis results of the hazard and the spatial contradiction identification results are then integrated to generate comprehensive hazard analysis data and output to step S4.
5. The gas hazard identification method based on a multimodal large model according to claim 4, characterized in that, The process of fusion of structured multimodal input and visual-text cross-attention in S31 includes the following steps: S31.
1. Structure and stitch together on-site inspection images, hazard target detection results, clarity scores, image quality levels, and gas expert tips according to the gas inspection scenario to form multimodal input data; S31.2 Input the multimodal input data into the visual encoder and the text encoder respectively to generate the corresponding visual feature sequence and text feature sequence; S31.
3. Employ a visual-text cross-attention mechanism to achieve bidirectional association and fusion of visual and textual features; S31.
4. Based on bidirectional cross-modal fusion features, complete the semantic parsing, type determination and hazard degree quantification of gas hazards.
6. The gas hazard identification method based on a multimodal large model according to claim 4, characterized in that, The process of constructing the two-dimensional topological relationship model of S32 includes the following steps: S32.1 Extract the category and center coordinates of gas equipment components from the hazard target detection results. Installation direction angle and the dimensions of the outer frame; S32.2, Using Euclidean distance Computing components With components The two-dimensional spatial spacing is used to determine whether there is a spatial adjacency relationship between components when the spacing is less than the preset adjacency threshold. S32.3 Construct a two-dimensional topology graph using components as nodes and spatial adjacency relationships as edges, combined with component category and direction attributes; S32.
4. Annotate the two-dimensional topology with structured attributes to form a two-dimensional topology model of gas equipment that is compatible with national standard verification.
7. The gas hazard identification method based on a multimodal large model according to claim 4, characterized in that, The process of verifying the national standard structural consistency and performing dynamic threshold self-calibration of S33 includes the following steps: S33.
1. Perform edge set matching between the actual topology graph and the national standard topology graph, based on accuracy. Recall rate Calculate F1 matching degree Among them, accuracy Recall represents the proportion of correctly matched edges in the actual topological edge set. This indicates the proportion of correctly matched edges in the national standard topology edge set; S33.2, Based on the average of historical matching data with standard deviation ,use The principle is to dynamically calibrate the verification threshold, and the calculation formula is: ; in: This represents the mean of the historical structural matching degree. The standard deviation of the historical structural matching degree. The threshold for verifying the structural matching degree after adaptive calibration; S33.3, when At that time, it was determined that there were potential problems with component installation, connection, and structural space.
8. The gas hazard identification method based on a multimodal large model according to claim 4, characterized in that, The zero-sample identification and comprehensive hazard analysis data generation process of S34 includes the following steps: S34.1 Based on the pre-training knowledge of multimodal large models, text feature vectors are used. With visual feature vectors Cosine similarity between Complete the matching calculation between text features and visual features; S34.
2. Complete zero-sample identification of untrained hazard types based on similarity thresholds; S34.3, Using weighted confidence levels The confidence scores of the three results—hidden danger semantic analysis, spatial contradiction identification, and zero-shot identification—are normalized and fused. Used to characterize the final credibility of the comprehensive hazard analysis results; S34.4 Generate comprehensive hazard analysis data based on the final fusion confidence level and output it to step S4.
9. The gas hazard identification method based on a multimodal large model according to claim 1, characterized in that, The process of linking the structured results output of S4 with the process includes the following steps: S4.
1. Perform unified structured processing on the comprehensive hazard analysis data, classify and archive the data according to hazard type, hazard location and risk level, standardize the data format and complete the equipment information associated with the hazard; S4.2 Based on the structured hazard data, generate a data set including hazard details, risk level, cause analysis, treatment recommendations, and final confidence level. A structured gas hazard identification report, which includes hazard details such as the equipment component where the hazard is located, its specific location, and the manifestation of the hazard. S4.
3. Associate and match the risk level in the structured gas hazard identification report with the gas operation and maintenance parameters. Based on the preset control rules corresponding to different risk levels, dynamically adjust the gas operation and maintenance parameters to achieve a linkage response between hazard identification and operation and maintenance processes.
10. A gas hazard identification system based on a multimodal large model, wherein the computer program of the gas hazard identification system based on a multimodal large model is used to execute the steps of the gas hazard identification method based on a multimodal large model as described in any one of claims 1-9, characterized in that, include: The image acquisition and quality detection unit is used to acquire on-site inspection images of gas pipelines and equipment, calculate the gradient variance using the Laplacian operator and normalize it to obtain a sharpness score, classify the image quality level according to a preset threshold, output a re-acquisition prompt for low-quality images, and output the sharpness score and image quality level. The hidden danger target detection unit is used to input qualified inspection images into the RF-DETR target detection model, extract multi-scale visual features through a convolutional backbone network, complete global modeling using a Transformer with deformable attention, and output the category and bounding box detection results of equipment components and hidden dangers through an ensemble prediction mechanism. A multimodal semantic understanding and spatial contradiction identification unit is used to complete multimodal data fusion, gas equipment structural hazard identification, and zero-sample hazard detection, generating comprehensive hazard analysis data. It includes a multimodal input and cross-modal fusion module, a topology modeling and national standard structure verification module, and a zero-sample identification and comprehensive data generation module, wherein: The multimodal input and cross-modal fusion module is used to fuse on-site inspection images, hazard target detection results, clarity scores, image quality levels and gas expert prompts to construct structured multimodal input data. It completes the association and fusion of visual and text features and hazard semantic analysis through a visual-text cross-attention mechanism. The topology modeling and national standard structure verification module extracts component features based on the detection results to construct a two-dimensional topology model, and combines it with national standard equipment structure rules to complete consistency verification. A dynamic threshold self-calibration mechanism identifies potential spatial conflicts related to component installation, connection, and structural elements. The zero-shot identification and comprehensive data generation module achieves zero-shot identification of untrained hazard types based on the pre-trained knowledge of a multimodal large model, and generates and outputs comprehensive hazard analysis data by integrating semantic analysis and spatial contradiction identification results. The structured result output and process linkage unit is used to perform structured processing on the comprehensive hidden danger analysis data, output a structured gas hidden danger identification report containing hidden danger details, risk level, cause analysis, treatment suggestions and confidence level, and adjust gas operation and maintenance parameters according to the report.