Classified detection method for home-entry security check gas equipment
By acquiring images through the camera of a mobile device and using a visual feature extraction network and a hierarchical detection head component, combined with a set of spatial relationship rules for gas equipment, the limitations of traditional manual inspection are overcome, efficient and accurate automatic classification of gas equipment is achieved, and recognition accuracy and robustness are significantly improved.
Patent Information
- Application Number
- CN202510929997.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional manual visual inspection of gas equipment is inefficient and easily affected by subjective judgment, lighting conditions and environmental interference, making it difficult to achieve efficient and accurate gas equipment classification and safety hazard identification. Existing image recognition technology has insufficient recognition accuracy and poor robustness in complex environments.
The mobile device camera is used to acquire images, and a visual feature extraction network consisting of a backbone network and an attention module is used for preprocessing and feature extraction. The hierarchical detection head component is combined for device identification, and a set of spatial relationship rules for gas equipment is constructed for context perception and result optimization.
It has achieved efficient, accurate and automated classification detection of gas equipment in complex environments, significantly improved recognition accuracy and robustness, and overcome the limitations of traditional manual detection.
Smart Images

Figure CN120766031A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of device classification, and more particularly, to a household security check gas device classification detection method BACKGROUND
[0002] With the development of society and the improvement of residents' safety awareness, the safety inspection of household gas equipment has become an important link to ensure public safety and family property safety. Traditional gas equipment security check relies on manual visual inspection. This method is not only inefficient, but also susceptible to subjective judgment, lighting conditions, equipment obstruction and other factors, resulting in the risk of missed detection and misjudgment of the detection results, making it difficult to achieve efficient and accurate classification of various gas equipment and identification of safety hazards. Especially in complex home environments, there are many types of gas equipment, such as stoves, water heaters, wall-mounted furnaces, gas meters, pipes, valves, etc. And due to differences in brand and model, the appearance of the same type of equipment may differ significantly, showing high intra-class variance. While some devices of different categories may be highly similar in appearance under certain angles, showing low inter-class variance, which greatly increases the difficulty of manual identification. In addition, household security checks are usually conducted in user kitchens or balconies, etc. These scenes are cluttered, lighting conditions are variable and uncontrollable, and equipment may be partially obscured. These environmental factors greatly interfere with image recognition, making it difficult for existing general image recognition techniques to achieve ideal detection accuracy and robustness in this specific application scenario.
[0003] In view of the above challenges, there is an urgent need for an automated detection method that can overcome the interference of complex environments and achieve multi-target accurate identification and classification of gas equipment. However, some existing general target detection or image classification techniques often show insufficient recognition accuracy and poor robustness when faced with this specific application of gas equipment with high target diversity and strong environmental complexity, making it difficult to meet the high reliability requirements of household security checks. They identify each device as an independent individual, ignoring the inherent spatial and functional relationships between gas equipment and their environment, making it difficult to effectively assist in judgment and correct results when the appearance of the equipment is blurred, partially obscured or has low confidence.
[0004] In order to effectively solve the above problems, an optimized household security check gas device classification detection method is urgently needed. SUMMARY
[0005] In view of the above limitations of existing methods, according to an aspect of the present application, a household security check gas device classification detection method is provided, which comprises: acquiring a live original image collected by a mobile device camera; After image preprocessing of the on-site original image, it is input into a visual feature extraction network containing a backbone network and an attention module to obtain an on-site device focused visual feature map; The on-site device focused visual feature map is input into a hierarchical detection head component containing a first level detection module and a second level detection module to obtain a list of initial detection results, each initial detection result in the list of initial detection results containing a final device type label, a bounding box coordinate and a confidence level; A set of gas device spatial relationship rules is constructed, and the list of initial detection results is context-aware and result-optimized based on the set of gas device spatial relationship rules to obtain a list of optimized detection results; Based on the list of optimized detection results and the on-site original image, a visual image with a bounding box and a device type label is generated.
[0006] Compared with the prior art, the household security gas device classification detection method provided by the present application first acquires an on-site image through a mobile device camera, and uses an attention-enhanced visual feature extraction network to preprocess and extract features from the image, so that it can adaptively focus on the target device area, effectively cope with complex environmental disturbances such as cluttered background, variable lighting and partial occlusion, and solve the problem that traditional manual detection is easily affected by the environment. Secondly, a hierarchical detection head component is used for device recognition, first detecting the device category, and then performing secondary recognition on the categories that need to be further subdivided, thereby effectively solving the problem of a large number of gas devices with similar or large differences in appearance, and significantly improving the precision of fine-grained classification. Finally, a set of gas device spatial relationship rules is constructed, and the initial detection results are context-aware and optimized based on this. By using the inherent spatial and functional association between gas devices, the detection results with low confidence or possible errors are corrected, overcoming the defect of existing general AI models that ignore context information, leading to unstable recognition. This method realizes efficient and accurate automatic classification and detection of gas devices, and significantly overcomes the limitations of traditional manual detection and the recognition difficulties in complex environments. BRIEF DESCRIPTION OF DRAWINGS
[0007] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application taken in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of embodiments of the present application and constitute a part of the specification, together with the description of embodiments of the present application, to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0008] Figure 1 A flowchart of the household security gas device classification detection method according to the embodiments of the present application.
[0009] Figure 2 Data flow diagram of the household security inspection gas equipment classification detection method according to the embodiment of the present application.
[0010] Figure 3 Flow chart of step S3 in the household security inspection gas equipment classification detection method according to the embodiment of the present application.
[0011] Figure 4 Flow chart of step S4 in the household security inspection gas equipment classification detection method according to the embodiment of the present application. DETAILED DESCRIPTION
[0012] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.
[0013] In view of the problems in the above background art, the present application proposes a household security inspection gas equipment classification detection method. Figure 1 Flow chart of the household security inspection gas equipment classification detection method according to the embodiment of the present application. Figure 2 Data flow diagram of the household security inspection gas equipment classification detection method according to the embodiment of the present application. As shown in Figure 1 and Figure 2 The household security inspection gas equipment classification detection method according to the embodiment of the present application includes: S1, acquiring a field original image collected by a camera of a mobile device; S2, after image preprocessing of the field original image, inputting it into a visual feature extraction network containing a backbone network and an attention module to obtain a field device focused visual feature map; S3, inputting the field device focused visual feature map into a hierarchical detection head component containing a first-level detection module and a second-level detection module to obtain a list of initial detection results, each initial detection result in the list of initial detection results containing a final device type label, a bounding box coordinate and a confidence; S4, constructing a gas equipment spatial relationship rule set, and based on the gas equipment spatial relationship rule set, performing context awareness and result optimization on the list of initial detection results to obtain a list of optimized detection results; S5, based on the list of optimized detection results and the field original image, generating a visualized image with a bounding box and a device type label.
[0014] In step S1, a raw image of the scene captured by the mobile device's camera is obtained. It should be understood that the traditional method of relying on manual visual inspection during home inspections of gas equipment has many inherent limitations. For example, detection efficiency is low, results are easily influenced by subjective judgment, and in complex and changing home environments, such as poor lighting conditions, partial obstruction of the device, and cluttered backgrounds, the accuracy and robustness of manual identification are difficult to guarantee. Furthermore, the wide variety of gas equipment types, with significant differences in appearance between similar devices and similar appearances between different categories, greatly increase the difficulty of manual identification and lead to the risk of missed detections or misjudgments. To overcome these challenges and achieve efficient and accurate classification of gas equipment and automated identification of safety hazards, this technical solution introduces an automated detection method based on image processing and deep learning. Acquiring raw images of the scene captured by the mobile device's camera is the starting point and foundation of the entire automated inspection process. Using portable, easy-to-use mobile devices, complex visual information from the scene can be converted into digital image data that can be processed by computers, laying the data foundation for subsequent processing.
[0015] The specific implementation process of step S1 is as follows: The security inspector carries a mobile device pre-installed with a dedicated security inspection application, such as a smartphone or tablet equipped with a high-resolution camera. When the security inspector arrives at a user's home and needs to inspect gas equipment, they launch the dedicated security inspection application. The application activates the mobile device's built-in camera module and displays the live image captured by the camera in real time. The security inspector aims the mobile device's camera at the gas equipment area to be inspected, such as a gas stove, gas water heater, gas meter, and its connecting pipes in the kitchen. During this process, the mobile device's camera's autofocus function automatically adjusts the focal length based on scene depth information to ensure that the target device is clearly visible in the image. Simultaneously, the auto-exposure and auto-white balance functions automatically adjust the image brightness and color based on the lighting conditions, such as natural light and indoor lighting, to obtain a high-quality image. Once the security inspector confirms that the device to be inspected is fully and clearly displayed in the center of the screen, they can capture one or more original images of the scene by clicking the capture button on the application interface or triggering other pre-set image capture commands. These images are stored in the local storage space of the mobile device in their original format without significant compression or post-processing, for example, high-resolution JPEG or PNG format, and the resolution can be preset to 1920x1080 pixels to retain sufficient detail information.
[0016] In step S2, after image preprocessing, the raw on-site image is input into a visual feature extraction network comprising a backbone network and an attention module to obtain a visual feature map of the on-site device focus. Accordingly, considering that raw on-site images captured from mobile device cameras often suffer from issues such as varying resolution, complex lighting conditions, cluttered backgrounds, and partial occlusion of devices, directly inputting these raw images into a deep learning model can lead to unstable model training, slow convergence, and difficulty in effectively extracting features critical for device identification. Image preprocessing standardizes and optimizes the raw image, eliminating or reducing noise and irrelevant information, making it more suitable as input for the deep learning network, thereby improving the efficiency and accuracy of feature extraction. The visual feature extraction network, comprising a backbone network and an attention module, automatically learns and extracts high-level, semantically rich visual features from the preprocessed image. Using an attention mechanism, it adaptively focuses on key areas of the gas equipment and suppresses background interference, effectively addressing the challenge of accurate multi-target identification in complex environments and laying a solid foundation for subsequent device classification.
[0017] The specific implementation process of image preprocessing of the original image on the scene in step S2 is as follows: the original image on the scene is first preprocessed, and its primary task is to unify the image size. Since different mobile devices or shooting conditions may cause the original image resolution to be different, in order to adapt to the subsequent visual feature extraction network's requirement for a fixed input size, the original image will be uniformly scaled to a preset standard size, such as 224x224 pixels or 300x300 pixels. This scaling process uses algorithms such as bilinear interpolation to avoid introducing obvious distortion while maintaining the image content. Next, pixel value normalization is performed. The pixel values of the original image are between 0 and 255, representing the brightness information of the image. In order to make the training process of the neural network more stable and efficient, these pixel values need to be normalized to a specific range, such as between 0 and 1, or standardized by subtracting the mean and dividing by the standard deviation. For example, the RGB value of each pixel can be divided by 255 to change its range to [0,1]. If the visual feature extraction network uses a backbone network pre-trained on the ImageNet dataset, it will be normalized using the mean and standard deviation of the ImageNet dataset. For example, for the RGB channels, the mean can be set to [0.485, 0.456, 0.406] and the standard deviation can be set to [0.229, 0.224, 0.225]. This normalization operation helps the model converge faster and improves its generalization ability.
[0018] Further, in order to efficiently and accurately extract visual features related to gas equipment from complex scenes, and adaptively focus on key areas, thereby effectively dealing with challenges such as cluttered background, light changes and equipment occlusion, it is necessary to extract visual features from the preprocessed image.
[0019] In particular, in one example of the present application, the backbone network is EfficientNet-B3, and the attention module is CBAM attention module. It is worth mentioning that the EfficientNet series model realizes the significant reduction of model parameter quantity and calculation quantity while maintaining high accuracy through the composite scaling method, that is, systematically unifying the depth, width and resolution of the network. This is crucial for security systems deployed on mobile devices or resource-constrained environments. It ensures that the model can achieve faster inference speed and lower resource consumption while ensuring recognition accuracy, thereby improving user experience and system practicality. Integrating the CBAM attention module is to further enhance the feature extraction capability and robustness of the network. In complex home security scenes, images may have problems such as cluttered background, uneven lighting, and partial equipment occlusion. CBAM can make the network adaptively learn and focus on the key feature channels and spatial regions related to gas equipment in the image through its channel attention and spatial attention mechanisms, effectively suppressing the interference of irrelevant background information, and extracting visual features with better discrimination and anti-interference ability. The introduction of this attention mechanism significantly improves the recognition accuracy and stability of the model in real complex environments.
[0020] The specific implementation process of inputting the preprocessed image into the visual feature extraction network containing the backbone network and the attention module in step S2 is as follows: the architecture of EfficientNet-B3 is mainly stacked by a series of Mobile Inverted BottleNeck Convolution (MBConv) modules, each of which contains a depth separable convolution, an SE attention mechanism and a residual connection, so that it can efficiently extract multi-scale features and enhance the feature expression capability. After a specific level of the EfficientNet-B3 backbone network, CBAM (Convolutional Block Attention Module) is embedded. CBAM is a lightweight general module designed to improve the feature expression capability of the convolutional neural network. It contains two sequential sub-modules: a channel attention module and a spatial attention module. The channel attention module first performs global average pooling and global maximum pooling on the input feature map, then sends the two pooling results into a shared multi-layer perceptron (MLP), adds the outputs of the MLP and generates channel attention weights through the Sigmoid activation function. The weights are then multiplied with the original feature map channel by channel to emphasize important feature channels. Then, the spatial attention module receives the channel attention weighted feature map, performs global average pooling and global maximum pooling on the channel dimension, concatenates the two results, and generates spatial attention weights through a standard convolution layer (for example, a convolution kernel size of 7x7) and a Sigmoid activation function. The weights are then multiplied with the channel weighted feature map pixel by pixel to emphasize important spatial positions. It is worth noting that the weights and bias parameters of this visual feature extraction network are first obtained by pre-training on a large general image dataset such as ImageNet to learn general visual features. Subsequently, these pre-trained weights are fine-tuned on a specially constructed gas equipment image dataset to make the network better adapt to the specific feature distribution of the gas equipment recognition task. During the fine-tuning process, the weights and biases of the network are iteratively updated according to the loss function (such as classification loss and positioning loss) through the back propagation algorithm and the optimizer (such as Adam or SGD) until the model performance converges. After processing by the EfficientNet-B3 backbone network and the CBAM attention module, the final on-site device focused visual feature map is output.
[0021] In step S3, the on-site device focus visual feature map is input into a hierarchical detection head component including a first-level detection module and a second-level detection module to obtain a list of initial detection results, each initial detection result in the list of initial detection results including a final device type label, a bounding box coordinate, and a confidence. It should be understood that a conventional single detection model is difficult to achieve ideal fine-grained classification accuracy in a gas device identification task with high intra-class variance and low inter-class variance. Therefore, the present application introduces a hierarchical detection head component to decompose the complex device identification task into more manageable sub-problems. The first-level detection module is responsible for identifying the large class of devices, reducing the difficulty of initial classification; while the second-level detection module focuses on fine-grained identification of specific large classes that need further subdivision, thereby significantly improving the accuracy of fine-grained classification. This hierarchical processing strategy not only improves the recognition ability of the model, but also enhances the interpretability of the detection results, enabling it to more accurately output a list of initial detection results including a final device type label, a bounding box coordinate, and a confidence.
[0022] Specifically, in one example of the present application, Figure 3 A flowchart of step S3 in the household security gas device classification and detection method according to an embodiment of the present application. As shown in Figure 3 S3, the on-site device focus visual feature map is input into a hierarchical detection head component including a first-level detection module and a second-level detection module to obtain a list of initial detection results, including: S31, inputting the on-site device focus visual feature map into the first-level detection module to obtain a list of device large class detection results, each device large class detection result in the list of device large class detection results including a device large class bounding box, a device large class label, and a confidence; S32, in response to the device large class label being a specific large class that does not need further subdivision, taking the device large class label as the final device type label; S33, in response to the device large class label being a specific large class that needs further subdivision, extracting a specific large class device ROI image from the on-site device focus visual feature map; S34, inputting the specific large class device ROI image into the second-level detection module to obtain a device subdivision type label as the final device type label. In particular, in one example of the present application, the first-level detection module is a Yolo detection head structure, and the second-level detection module is a fine-grained classifier.
[0023] The implementation process of step S3 is as follows: first, S31 is executed. The field device focused visual feature map is sent into the first level detection module. In this embodiment, the first level detection module adopts the detection head structure of Yolo. The detection head internally contains a series of convolution operations, which gradually process the feature map and extract higher level semantic information. Finally, through one or more output convolution layers, the Yolo detection head maps the feature map to a multi-dimensional tensor. This tensor is divided into a grid (for example, 13x13 or 26x26) in the spatial dimension, and each grid unit is responsible for predicting the gas equipment whose center falls within the unit. For each grid unit, it will predict several bounding boxes, for example, 3 or 5, each containing five basic information: the center point coordinates (x, y) of the bounding box, the width (w), the height (h), and an object confidence, indicating the probability of the existence of a gas equipment in the bounding box and the accuracy of the positioning of the box. In addition, each bounding box will also predict a probability distribution of a set of equipment class labels, such as stoves, water heaters, gas meters, pipes, etc. In order to process gas equipment of different sizes, the Yolo detection head will make predictions on different scale feature maps of the backbone network, that is, adopt a multi-scale detection strategy, so that the model can simultaneously identify large, medium and small equipment in the image. After the Yolo detection head completes the prediction, a post-processing step is performed to optimize the results. First, according to the preset confidence threshold, for example, 0.7, the prediction boxes with too low confidence are filtered out. Then, the non-maximum suppression algorithm is applied. Non-maximum suppression is a commonly used post-processing technique to eliminate multiple overlapping prediction boxes for the same gas equipment, and only keep the one with the highest confidence. For example, if multiple bounding boxes are highly overlapping and all predict stoves, non-maximum suppression will only keep the one with the highest confidence according to their confidence, and suppress the low confidence boxes with a high overlap threshold (for example, the intersection over union IoU threshold is 0.5). Finally, a list of equipment class detection results is obtained. Each element in the list represents a detected gas equipment class and contains its precise equipment class bounding box, that is, the pixel coordinates on the original image, the corresponding equipment class label, and the confidence of the detection result. It is worth noting that the weights and bias parameters of the Yolo detection head are obtained by end-to-end training on a labeled dataset containing a large number of gas equipment images.
[0024] In particular, when the field device focus visual feature map obtained based on the CBAM attention mechanism is subjected to hierarchical detection based on different category densities, the dynamic probability boundary constructed by the attention mechanism will present multi-scale coupled oscillation of the category probability distribution during hierarchical detection, which makes it necessary to suppress the high-dimensional multi-granularity coupled oscillation response of the field device focus visual feature map while strengthening the local image semantic feature distribution through the attention mechanism, so as to avoid causing image semantic regression multi-scale focusing misalignment when positioning the category probability through hierarchical detection, and affecting the hierarchical accuracy of the initial detection result under different density label probability distributions.
[0025] Based on this, when performing S31, one preferred embodiment of the present application inputs the field device focus visual feature map into the first-level detection module to obtain a list of device category detection results, including: performing category offset fractal probability boundary calculation on each feature value in the field device focus visual feature map to obtain a category probability feature probability boundary value corresponding to each feature value, that is, ; wherein, is each feature value in the field device focus visual feature map, is the category probability feature probability boundary value corresponding to each feature value, that is, for each feature value of the field device focus visual feature map , first construct the fractal probability boundary corresponding to the image semantics through category offset to determine the feature probability boundary performance in the form of category probability.
[0026] Based on the category probability feature probability boundary value corresponding to each feature value, construct a feature trajectory phase evolution gradient value corresponding to each feature value, that is, ; wherein, is the feature trajectory phase evolution gradient value corresponding to each feature value, that is, calculate the partial derivative of the fractal probability boundary with respect to the feature value itself to obtain the feature trajectory phase evolution gradient.
[0027] Take the feature trajectory phase evolution gradient value corresponding to each feature value as a fractal probability boundary constraint hyperparameter to obtain a multi-scale interactive reconstruction inverse mapping core, that is, ; wherein, is the multi-scale interactive reconstruction inverse mapping core, that is., through the feature trajectory phase evolution gradient as the fractal probability boundary constraint hyperparameter, the implicit oscillation response is converted into the multi-scale interactive reconstruction inverse mapping core.
[0028] The multi-scale interactive reconstruction inverse mapping core is taken as a density reference benchmark of high-dimensional multi-granularity coupled oscillation, and each feature value in the field device focused visual feature map is interactively constrained to obtain an improved field device focused visual feature map, that is: ; wherein, each feature value in the improved field device focused visual feature map, thereby, taking the multi-scale interactive reconstruction inverse mapping core as a density reference benchmark of high-dimensional multi-granularity coupled oscillation, and further ensuring that the phase evolution conforms to multi-scale interactive constraints through the response calculation of the feature value microscopic level to the macroscopic reconstruction space while the implicit phase evolution of the fractal probability boundary representing different category probabilities evolves; The improved field device focused visual feature map is input into the first-level detection module to obtain a list of device category detection results, so that the field device focused visual feature map The image semantic regression multi-scale focusing misalignment is caused when positioning the category probability through hierarchical detection, thereby improving the hierarchical accuracy of the initial detection result under different density label probability distributions. In particular, the implementation of this step is the same as the implementation of the above-mentioned field device focused visual feature map input into the first-level detection module.
[0029] Then, S32 is performed. Each detection result in the list of device category detection results is judged based on a pre-configured device classification hierarchy mapping table. For example, the mapping table can define gas meters, gas pipelines, valves, etc. as categories that do not need to be further subdivided, because they do not need to be distinguished more carefully in security checks; while stoves, water heaters, etc. are defined as categories that need to be further subdivided, because they can have different models such as embedded / table, strong exhaust / balanced, etc. which are important for security checks.
[0030] Each detection result in the list of device category detection results is judged. In response to a certain device category label being pre-set as a specific category that does not need to be further subdivided, such as a gas meter category or a gas pipeline category, the device category label is directly determined as the final device type label of the detection result. This means that for this type of device, the first-level detection has provided sufficient classification granularity.
[0031] However, S33 is performed. In response to a certain device category label being preset as a specific category that needs further subdivision, such as a stove category or a water heater category, a corresponding specific category device ROI image is extracted from the on-site device focused feature map. This extraction process is based on the bounding box of the device category output by the first-level detection module, and through region of interest pooling or region of interest alignment, the feature region corresponding to the device category is accurately cropped or aligned from the original feature map to form a fixed-size feature map, i.e., a specific category device ROI image.
[0032] Finally, S34 is performed. The extracted specific category device ROI image is input into the second-level detection module, i.e., the fine-grained classifier. The fine-grained classifier is an independent neural network model, and its architecture can be a small convolutional neural network (e.g., containing several convolutional layers, pooling layers, and fully connected layers, and finally outputting class probabilities through a Softmax activation function) or a multi-layer perceptron, which is specifically used to distinguish subcategories under a specific category. For example, for a ROI image of a stove category, the fine-grained classifier can further identify whether it is an embedded stove or a table stove. The weight and bias parameters of the fine-grained classifier are obtained by training on an image dataset containing subcategory labels under the specific category, and the training goal is to minimize the subcategory classification error. The class label output by the fine-grained classifier is the device subcategory label of the detection result, and it is used as the final device type label of the detection result. After the above hierarchical processing, a list of initial detection results is finally obtained, each of which contains an accurate final device type label, a bounding box coordinate, and a confidence.
[0033] In step S4, a set of gas device spatial relationship rules is constructed, and the list of initial detection results is context-aware and result-optimized based on the set of gas device spatial relationship rules to obtain a list of optimized detection results. Accordingly, although the foregoing steps can effectively identify gas devices, in a complex environment, the initial detection results may still have low confidence, missed detection, or misjudgment. Gas devices do not exist in isolation in actual installation, and there are inherent spatial and functional associations between them and the environment that conform to physical logic, such as a gas pipeline connected to a stove or a water heater. Thus, in this application, by utilizing these prior context information, the results identified by the visual model can be cross-verified and corrected, especially when a device has low confidence, if it meets the preset spatial relationship rules with a related device with high confidence, its confidence can be improved, thereby correcting potential errors, improving the robustness and accuracy of the overall detection, and overcoming the limitations of traditional independent target detection that ignores context information.
[0034] Specifically, in one example of the present application, Figure 4The figure is a flow chart of step S4 in the household security inspection gas equipment classification detection method according to the embodiment of the present application. As shown in Figure 4 step S4, a set of gas equipment spatial relationship rules is constructed, and the list of initial detection results is contextually perceived and optimized based on the set of gas equipment spatial relationship rules to obtain a list of optimized detection results, including: S41, extracting a first initial detection result and a second initial detection result from the list of initial detection results; S42, extracting a gas equipment spatial relationship rule that is adapted to the first initial detection result and the second initial detection result from the set of gas equipment spatial relationship rules; S43, correcting the confidence in the initial detection result with lower confidence among the first initial detection result and the second initial detection result based on the gas equipment spatial relationship rule to obtain a corrected confidence.
[0035] The specific implementation process of step S4 is as follows: first, a set of gas equipment spatial relationship rules is constructed. This rule set is pre-established based on industry standards, safety specifications for gas equipment installation, and statistical analysis and expert experience of a large number of actual security inspection scene data. It contains common spatial relative position relationships and connection relationships between different types of gas equipment. For example, the rule can be defined as: the gas stove is located below the gas pipeline, and there is a connection relationship between them in the horizontal direction; the gas meter is located on the gas pipeline and close to the entry point; the gas water heater is located above the gas pipeline and is connected to it. Each rule not only contains the logical relationship between the device types, but also may contain spatial predicates such as above, below, close to, connected, and corresponding spatial distance or overlap threshold. These rules are stored in a structured form, for example, they can be a knowledge graph or a relational database, so as to facilitate subsequent query and matching.
[0036] Next, from the list of initial detection results, a first initial detection result and a second initial detection result are extracted. This is achieved by traversing all detection result pairs in the list. For any two detection results in the list, they are designated as the first and second initial detection results respectively, and their device type labels, bounding box coordinates and confidence are obtained.
[0037] Then, from the pre-constructed set of gas equipment spatial relationship rules, a gas equipment spatial relationship rule that is adapted to the device type labels of the current first initial detection result and the second initial detection result is extracted. For example, if the first detection result is a gas stove and the second detection result is a gas pipeline, it is queried whether there is a rule about the relationship between the gas stove and the gas pipeline in the rule set. Once the adapted rule is found, the actual spatial relationship between the two detection results, such as relative position, distance, overlapping area, etc., is calculated according to the bounding box coordinates of the two detection results, and compared with the spatial conditions defined in the rule to confirm whether the rule is satisfied in the current scene.
[0038] Finally, based on the gas equipment spatial relationship rules, the confidence of the initial detection result with lower confidence in the first initial detection result and the second initial detection result is corrected to obtain a corrected confidence. Specifically, if there is an adaptive rule that meets the spatial condition, the confidences of the two detection results are compared, and the confidence of the detection result with lower confidence is corrected. After iterative correction of all related detection result pairs, a list of optimized detection results is finally obtained. That is, in the actual home security scene, the gas equipment may be affected by complex environmental factors such as uneven lighting, partial occlusion, and cluttered background, resulting in unclear visual features, so that the detection confidence of the deep learning model is low. However, there are often fixed spatial and functional relationships between gas equipment, for example, gas pipelines must be connected to gas stoves or water heaters. If a gas stove is detected with high confidence, but the gas pipeline connected to it has low confidence for some reason, then according to their strong correlation, it can be inferred that the presence of the gas pipeline is highly reliable. Therefore, by correcting the low-confidence result, the context information can be used to make up for the shortcomings of visual recognition, reduce missed detection and misjudgment, and improve the robustness and accuracy of the overall detection, so as to more comprehensively and reliably evaluate the safety status of the gas equipment.
[0039] Specifically, in one example of the present application, based on the gas equipment spatial relationship rules, the confidence of the initial detection result with lower confidence in the first initial detection result and the second initial detection result is corrected to obtain a corrected confidence, comprising: setting the confidence in the first initial detection result as a first confidence, setting the confidence in the second initial detection result as a second confidence, and the first confidence is greater than the second confidence; the confidence in the second initial detection result is corrected as follows: ; wherein, the first confidence is the second confidence is is a confidence improvement factor, which can be set according to the actual application scene and experience, for example, 0.2, is a minimum value function, is the corrected confidence. That is, when two gas equipment with spatial correlation are identified, one of which, such as the first initial detection result, has a higher confidence, the formula will try to improve the second confidence. The magnitude of the improvement is determined by the second confidence multiplied by a factor greater than 1, wherein is a confidence improvement factor. At the same time, in order to avoid excessive correction, the corrected confidence will not exceed the confidence of the higher confidence associated with it. This ensures the rationality of the correction process, that is, the low-confidence result is improved while not exceeding the reliability of its high-confidence associated object.
[0040] In step S5, based on the list of optimized detection results and the original image of the scene, a visual image with a bounding box and device type label is generated. It should be understood that although the aforementioned steps can accurately identify gas equipment and optimize its confidence, its output is in the form of a data list, which is not convenient for security personnel to quickly and accurately grasp the situation on the scene. In order to convert abstract machine recognition results into intuitive visual information that is easy for humans to understand and verify, this application greatly improves the interpretability and usability of automated detection results by directly marking the bounding box and type label of the device on the original image. This not only helps security personnel to quickly check the test results and conduct on-site reviews, but also serves as an important safety inspection record, providing reliable visual evidence for subsequent audits, traceability and safety management, thereby effectively improving the efficiency and reliability of gas equipment security inspections.
[0041] The specific implementation process of step S5 is as follows: first, the original image of the scene collected in step S1 is obtained as the base map, and the list of optimized detection results outputted in step S4 is obtained.
[0042] Next, iterate over each detection result in the list of optimized detection results. For each entry in the list, first extract its bounding box coordinates. These coordinates are in pixels and define the position and size of the device in the original image, for example, the x, y coordinates of the upper left corner and the x, y coordinates of the lower right corner. Use an image processing library, for example, the OpenCV library or the Pillow library in Python to draw a rectangular box on the original image of the scene, whose position and size exactly correspond to the extracted bounding box coordinates. For visual clarity, the color of the bounding box can be preset, for example, uniformly using green to represent detected devices and a line thickness of, for example, 2 pixels.
[0043] Next, the final device type label of the detection result is extracted, for example, gas cooker-embedded, gas water heater-forced exhaust, and an optional corrected confidence level. This text information will be placed near the drawn bounding box, such as in the upper left corner or above the bounding box. To ensure the readability of the text, it is necessary to preset the font type, for example, the system default font, font size, for example, 16-point font, and text color, for example, white or yellow, so that it can be clearly displayed on different backgrounds. To further improve the contrast, a background color block that contrasts with the text color can be added under the text. For example, if the text is white, the background color block can be translucent black.
[0044] This process will be repeated until all gas equipment in the optimized detection result list is labeled one by one on the original image on the spot. Finally, a visualization image with the bounding box and type label of all detected gas equipment is generated. This image can be saved in a common image format such as JPEG or PNG and can be directly displayed on the screen of a mobile device to the security personnel or uploaded to a cloud server for archiving and further analysis.
[0045] In summary, the household security gas equipment classification detection method based on the embodiments of the present application is illustrated, which first acquires the on-site image through the mobile device camera, and uses the attention-enhanced visual feature extraction network to pre-process and extract features from the image, so that it can adaptively focus on the target device area, effectively cope with complex environmental interference such as cluttered background, variable lighting and partial occlusion, and solve the problem that traditional manual detection is easily affected by the environment. Secondly, the hierarchical detection head component is used for device recognition, first detecting the device category, and then performing secondary recognition on the category that needs further subdivision, thereby effectively solving the problem of a large number of gas equipment types, similar or large differences in appearance, and significantly improving the precision of fine-grained classification. Finally, a set of spatial relationship rules of gas equipment is constructed, and based on this, the initial detection result is context-aware and optimized. By using the inherent spatial and functional association between gas equipment, the detection result with low confidence or possible error is corrected, overcoming the defect that the existing general AI model ignores the context information, leading to unstable recognition. This method realizes efficient and accurate automatic classification and detection of gas equipment, and significantly overcomes the limitations of traditional manual detection and the recognition difficulty in complex environments.
[0046] The above has described various embodiments of the present disclosure, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications, or improvements to the technology in the market of the embodiments, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A classification detection method for household gas equipment security inspection, characterized in that: include: Obtaining the original image of the scene captured by the camera of the mobile device; After image preprocessing, the on-site original image is input into a visual feature extraction network comprising a backbone network and an attention module to obtain a visual feature map of on-site equipment focus; Inputting the field device focused visual feature map into a hierarchical detection head component comprising a first-level detection module and a second-level detection module to obtain a list of initial detection results, wherein each initial detection result in the list of initial detection results comprises a final device type label, bounding box coordinates, and a confidence level; Constructing a gas equipment spatial relationship rule set, and performing context perception and result optimization on the list of initial detection results based on the gas equipment spatial relationship rule set to obtain a list of optimized detection results; Based on the list of optimized detection results and the original image of the scene, a visualization image with bounding boxes and device type labels is generated.
2. The classification detection method for household safety inspection gas equipment according to claim 1 is characterized in that: The backbone network is EfficientNet-B3, and the attention module is the CBAM attention module.
3. The classification detection method for household gas equipment for security inspection according to claim 2 is characterized in that: Inputting the focused visual feature map of the field device into a hierarchical detection head component including a first-level detection module and a second-level detection module to obtain a list of initial detection results, including: Inputting the on-site device focused visual feature map into the first-level detection module to obtain a list of device category detection results, wherein each device category detection result in the list of device category detection results includes a bounding box of the device category, a device category label, and a confidence level; In response to the device category label being a specific category that does not require further subdivision, using the device category label as the final device type label; In response to the device category label being a specific category that needs to be further subdivided, extracting a specific category device ROI image from the on-site device focused visual feature map; The specific large-category device ROI image is input into the second-level detection module to obtain a device subdivision type label as the final device type label.
4. The classification detection method for household gas equipment for security inspection according to claim 3 is characterized in that: The first-level detection module is a Yolo detection head structure, and the second-level detection module is a fine-grained classifier.
5. The classification detection method for household safety inspection gas equipment according to claim 4 is characterized in that: Input the focused visual feature map of the on-site equipment into the first-level detection module to obtain a list of equipment category detection results, including: Performing class shift fractal probability boundary calculation on each eigenvalue in the on-site device focused visual feature map to obtain a class probability feature probability boundary value corresponding to each eigenvalue; Based on the category probability feature probability boundary value corresponding to each eigenvalue, constructing the characteristic trajectory phase evolution gradient value corresponding to each eigenvalue; The characteristic trajectory phase evolution gradient value corresponding to each eigenvalue is used as a fractal probability boundary constraint hyperparameter to obtain a multi-scale interactive reconstruction inverse mapping core; Using the multi-scale interactive reconstruction inverse mapping core as a density reference benchmark for high-dimensional multi-granularity coupled oscillation, interactively constraining each eigenvalue in the field device focused visual feature map to obtain an improved field device focused visual feature map; The improved on-site device focused visual feature map is input into the first-level detection module to obtain a list of the device major category detection results.
6. The classification detection method for household gas equipment for security inspection according to claim 1 is characterized in that: Constructing a gas equipment spatial relationship rule set, and performing context perception and result optimization on the list of initial detection results based on the gas equipment spatial relationship rule set to obtain a list of optimized detection results, including: extracting a first initial detection result and a second initial detection result from the list of initial detection results; extracting a gas equipment spatial relationship rule that is adapted to the first initial detection result and the second initial detection result from the gas equipment spatial relationship rule set; Based on the gas equipment spatial relationship rule, the confidence of the initial detection result with a lower confidence among the first initial detection result and the second initial detection result is corrected to obtain a corrected confidence.
7. The classification detection method for household gas equipment for security inspection according to claim 6 is characterized in that: Based on the gas equipment spatial relationship rule, the confidence of the initial detection result with a lower confidence in the first initial detection result and the second initial detection result is corrected to obtain a corrected confidence, including: Setting the confidence level in the first initial detection result to a first confidence level, setting the confidence level in the second initial detection result to a second confidence level, and the first confidence level is greater than the second confidence level; The confidence level in the second initial detection result is corrected using the following formula: ;in, is the first confidence level, is the second confidence level, is the confidence boost factor, To obtain the minimum function, is the modified confidence.
Citation Information
Cited By
Target detection method, system and device based on hierarchical collaborative reasoning
CN121330440A