Computer vision-based intelligent gas station safety supervision method and system

By using computer vision technology, combined with image enhancement, feature construction, and abnormal texture synthesis, the problem of missed detection of small targets and occluded targets by single-stage detectors has been solved, achieving high-precision and high-efficiency detection for gas station safety supervision.

CN122156581APending Publication Date: 2026-06-05XIAN ZHONGZHI IOT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN ZHONGZHI IOT TECH CO LTD
Filing Date
2026-02-09
Publication Date
2026-06-05

Smart Images

  • Figure CN122156581A_ABST
    Figure CN122156581A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision, and solves the technical problem that the single-stage detector in the prior art has a high missing detection rate for small targets and severely occluded targets, the classification result of the detection frame is rough, and accurate detection results cannot be provided, and particularly relates to a smart gas station safety supervision method and system based on computer vision, the steps of the method being: obtaining an original single frame image, performing image enhancement processing on the original single frame image to obtain an enhanced balanced image, the high-precision target detection model and the fine-grained state classification model are combined in cascade in the present application, all potential targets are positioned with a high recall rate, the combination mode of rough positioning combined with fine discrimination is used, the target detection integrity in a complex and crowded scene and the accuracy of state recognition are significantly improved, the comprehensive recognition rate is greatly improved on the basis of the prior art, and the problems of missing detection and false detection of small targets and dense targets are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a smart gas station safety supervision method and system based on computer vision. Background Technology

[0002] Computer vision refers to using cameras and computers to identify, track, and measure targets instead of human eyes, enabling computer processing to produce images that are more suitable for human observation or transmission to instruments for detection. For example, surveillance cameras can be used to detect smoke and flames, dangerous behaviors such as people smoking or using mobile phones, and vehicles that are illegally parked or have unusual intrusions in real time. Existing technologies are based on fire alarm systems using traditional sensors, manual inspections and video surveillance, and data integration with IoT devices. Through edge computing or cloud platforms, video streams are analyzed in real time to automatically trigger alarms and link with emergency systems.

[0003] Existing technologies generally have a high false negative rate for small targets and severely occluded targets, such as cigarette butts or individuals in dense crowds at gas stations, using a single, single-stage detector. Furthermore, the detection box classification results are coarse and cannot provide accurate detection results, such as detailed status information like vehicles not turned off or people not wearing work clothes. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a smart gas station safety supervision method and system based on computer vision. It solves the technical problems of existing technologies, which use a single, single-stage detector, resulting in a high rate of missed detection for small targets and severely occluded targets, coarse detection box classification results, and inability to provide accurate detection results.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a smart gas station safety supervision method based on computer vision, the steps of which are as follows: The original single-frame image is acquired, and image enhancement processing is performed on the original single-frame image to obtain an enhanced equalized image; Based on the enhanced equalization image, a structured list is obtained through feature construction and structuring. Based on structured lists and enhanced equalization images, synthetic anomalous images and anomalous annotations are obtained through anomalous texture synthesis. The synthetic anomalous images and anomalous annotations are then integrated into a dangerous behavior recognition dataset. The dangerous behavior recognition dataset and structured list are converted into a structured list sequence. The structured list sequence is then refined and an early warning detection is performed to obtain behavior labels and early warning signals. Feedback information reports are obtained through retrieval and similarity processing based on warning signals, structured list sequences, and enhanced equalization images.

[0006] Preferably, image enhancement processing is performed on the original single-frame image, including: An enhanced illumination image is obtained by enhancing the original single-frame image through an adversarial network. A denoised image is obtained by blurring and denoising the enhanced illumination image. Histogram equalization is performed on the denoised image to obtain an enhanced equalized image.

[0007] Preferably, based on the enhanced equalization image, feature construction and structuring processing include: A feature pyramid fusion layer is constructed based on the enhanced equalization image to obtain the fused feature pyramid. An attention prediction layer is constructed, and attention processing is performed based on the fused feature pyramid to obtain a set of candidate targets; A structured list is obtained by fine-grained classification and structuring based on the candidate target set and the enhanced equalization image.

[0008] Preferably, based on the candidate target set and the enhanced equalization image, fine-grained classification and structuring processing are performed, including: A cropping and alignment layer is constructed, and cropping and alignment processing is performed based on the candidate target set and the enhancement and equalization image to obtain an aligned image set; Construct a residual classification layer, perform multi-task residual processing based on the aligned image set, and obtain fine-grained state probability distribution; The fine-grained state probability distribution is encapsulated in a structured manner to obtain a structured list.

[0009] Preferably, the process of synthesizing anomalous textures based on a structured list and enhanced equalization images includes: Based on the enhanced equalization image, a semantic segmentation mask is obtained through semantic segmentation processing. Based on semantic segmentation masks and structured lists, local abnormal texture patches are obtained through texture anomaly processing. Consistency adjustment is performed on the enhanced equalized image and local anomalous texture patches to obtain synthetic anomalous images and anomalous annotations, and a dangerous behavior recognition dataset is generated.

[0010] Preferably, the structured list sequence is refined and an early warning detection is performed, including: Based on the structured list sequence, a preliminary set of schemes is obtained through high- and low-frequency information processing; Multi-target correlation processing is performed on the preliminary scheme set and the structured list sequence to obtain a multi-target trajectory feature set; Enhancement, refinement, and early warning processing are performed based on the multi-target trajectory feature set to obtain behavior labels and early warning signals.

[0011] Preferably, enhanced refinement and early warning processing are performed based on multi-target trajectory feature sets, including: Based on the multi-target trajectory feature set, spatiotemporal enhancement processing is used to obtain the context enhancement scheme features; The features of the context enhancement scheme are classified and refined to obtain the refined scheme and classification probability. Based on the dangerous behavior recognition dataset, refinement scheme, and classification probability, early warning recognition processing is performed to obtain behavior labels and early warning signals.

[0012] Preferably, based on the warning signal, the structured list sequence, and the enhanced equalization image, the process includes retrieval and similarity processing, including: A multi-branch deep network layer is constructed based on the early warning signal, the structured list sequence, and the enhanced equalization image to obtain a multi-granularity feature vector group; Based on the multi-granularity feature vector group, a normalized feature vector is obtained through cross-modal feature fusion processing. The normalized feature vectors are processed to construct a hierarchical index, resulting in a hierarchical index structure. Similarity retrieval is performed based on normalized feature vectors and a hierarchical index structure to obtain a list of retrieval results; The search results are processed using dynamic thresholding to generate a feedback report.

[0013] This technical solution also provides a system for applying the aforementioned computer vision-based smart gas station safety supervision method, the system comprising: The equalization module is used to acquire the original single-frame image, perform image enhancement processing on the original single-frame image, and obtain an enhanced equalized image. The structuring module is used to construct and structure the list based on the enhanced equalization image through feature construction and structuring. The anomaly labeling module is used to synthesize anomaly images and anomaly labels based on a structured list and augmented equalization images through anomaly texture synthesis, and to integrate all the synthesized anomaly images and anomaly labels into a dangerous behavior recognition dataset. The early warning module is used to convert the dangerous behavior identification dataset and structured list into a structured list sequence, refine the structured list sequence and perform early warning detection to obtain behavior labels and early warning signals; The classification feedback module is used to obtain feedback information reports based on the warning signal, structured list sequence, and enhanced equalization image through retrieval and similarity processing.

[0014] By employing the above technical solution, the present invention provides a smart gas station safety supervision method and system based on computer vision, which has at least the following beneficial effects: 1. This invention cascades a high-precision target detection model with a fine-grained state classification model. First, a detection model with optimized feature pyramids and attention mechanisms is responsible for locating all potential targets with high recall. Then, for each detection region, a dedicated state classification model performs secondary in-depth analysis and attribute discrimination. This combination of coarse localization and fine discrimination significantly improves the completeness of target detection and the accuracy of state recognition in complex and crowded scenes, greatly improving the overall recognition rate on the existing basis and solving the problems of missed detection and false detection of small and dense targets.

[0015] 2. This invention enhances tracking robustness by fusing appearance and motion information through dual paths and combining trajectory association algorithms based on appearance features. It generates context-enhanced features based on spatiotemporal dependencies and achieves accurate behavior classification and flexible early warning triggering through dynamic thresholds. This effectively solves the shortcomings of traditional methods in feature fusion, trajectory association, context reasoning, and adaptability, and improves the recognition accuracy and response efficiency of dangerous behaviors in complex dynamic scenarios.

[0016] 3. This invention integrates the deep global features of images, the local features of key targets, and the semantic tags of events to form a unified feature vector for joint indexing and similarity measurement. The subsequent process is determined based on the similarity of the search results. High similarity results are directly associated with historical cases to assist decision-making, while low similarity results are archived as new events and trigger database updates. This greatly improves the accuracy of tracing complex events and the system's self-evolution capability, enabling gas station managers to quickly find historical data with similar vehicles, similar personnel behaviors, or similar violations. At the same time, it accumulates rare event samples for the system, significantly improving the retrieval recall rate and processing efficiency. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the computer vision-based smart gas station safety supervision method of the present invention; Figure 2 This is a structural block diagram of the intelligent gas station safety monitoring system based on computer vision according to the present invention. Detailed Implementation

[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0019] Example 1: Due to the high false negative rate of existing technologies using a single, single-stage detector for small targets and severely occluded targets, and the resulting coarse classification of bounding boxes, accurate detection results cannot be provided. Please refer to [the relevant documentation / reference]. Figure 1 This embodiment provides a smart gas station safety supervision method based on computer vision, which can reduce the probability of missed detections, make the detection box classification more detailed, and significantly improve the detection accuracy. The method includes the following steps: S1. Obtain the original single-frame image and perform image enhancement processing on the original single-frame image to obtain an enhanced equalized image. Existing technologies using traditional histogram equalization methods are prone to loss of detail or over-enhancement when dealing with extreme lighting or complex noise, resulting in blurred target features. To solve the above problems, the specific implementation steps are as follows: S11. Based on the original single-frame image, an adversarial network is used to enhance the image to obtain an enhanced illumination image. In this step, a generator and a discriminator are first constructed. The generator learns the transformation rules from the original single-frame image to the enhanced illumination image, where random noise is used to adjust the randomness of the generation process. The discriminator is trained adversarially by comparing the difference between the real normal illumination image and the generated image. Specifically, the adversarial loss calculation consists of two parts: the first part is the expected value of the probability that the discriminator correctly identifies the real normal illumination image; the second part is the expected value of the probability that the discriminator misidentifies the generated image. The final loss is the sum of the expected values ​​of these two parts. The original single-frame image contains pixel brightness values, color channel information, and spatial structure. The data structure, such as a low-light night scene photo, generally has low pixel brightness values. The red channel may appear dark due to the loss of shadow details, the green channel retains some vegetation outlines, and the blue channel shows noise due to insufficient light. When the generator processes the data, it will adjust the pixel brightness values ​​by multiplying them by a brightness enhancement factor, add contrast enhancement values ​​to the color channels, and add edge sharpening factors to the spatial structure data. The final output is a light-enhanced image. For example, the blurry building outlines in the original image become clear after being enhanced by the edge sharpening factor, more textures are visible after the shadow details are enhanced by the brightness enhancement factor, and the color channels are restored to natural tones after the contrast enhancement value is adjusted. The overall effect is close to a clear scene under normal lighting, and the final image is a light-enhanced image.

[0020] S12. Obtain a denoised image by blurring and denoising the enhanced image. In this step, deblurring is first performed in the frequency domain using Wiener filtering. Specifically, this involves multiplying the conjugate of the point spread function by the Fourier transform of the blurred image, then dividing by the sum of the squared magnitude of the point spread function and the noise power ratio constant. This operation is equivalent to frequency compensation in the frequency domain, enhancing high-frequency details while suppressing noise amplification. Subsequently, denoising is performed in the spatial domain using nonlocal mean filtering. This step calculates weights based on the similarity of image patches surrounding a pixel, and weights the pixel values ​​in similar regions. For example, taking a night scene image with enhanced illumination as an example, it may contain motion blur caused by lens shake and sensor noise. After Wiener filtering, the blurred street light outlines will be restored to clear edges through frequency compensation. The graininess originally caused by noise will be smoothed by the weighted average of similar image blocks in non-local mean filtering. The final output clear denoised image not only retains the sharp edges of the street lights, but also eliminates the snowflake-like noise in the dark areas. The overall picture is as clear and natural as a night scene shot with a stable lens under normal lighting, and finally a clear denoised image is obtained.

[0021] S13. Perform histogram equalization on the denoised image to obtain an enhanced equalized image. In this step, the clear denoised image is first divided into multiple sub-regions. When performing histogram equalization on each region, if the number of pixels at a certain gray level exceeds the contrast threshold, the excess is evenly distributed to other gray levels to avoid excessive enhancement of local contrast that could amplify noise. For example, in the original image, dark areas appear hazy due to concentrated pixel values. After histogram cropping, the pixel values ​​in the dark areas are appropriately stretched, making the originally blurry shadow textures clearer. Subsequently, bilinear interpolation is used to weightedly fuse the edge pixels of adjacent sub-regions, eliminating the blocky boundaries caused by the block processing. For example, color abrupt changes at the boundaries of different regions in the original image are smoothed out, making the overall image color transition natural. The final output... The enhanced image retains detail clarity and restores natural color contrast. For example, in a clear, denoised night scene photo, the building exterior texture, which was originally difficult to distinguish due to insufficient contrast, will present clear brick and stone textures and natural color gradations after processing. The overall effect is like a clear image taken under normal lighting. The low-light enhancement stage of this invention uses a conditional generative adversarial network. Through adversarial training between the generator and the discriminator, it can more accurately learn the complex mapping from low light to normal light. In the deblurring and denoising stages, it effectively solves the limitation that a single method cannot handle blur and noise simultaneously. By combining block equalization with contrast threshold clipping, it stretches the details in the dark areas while limiting noise amplification. Bilinear interpolation further eliminates block boundaries, overcoming the defects of block artifacts and excessive noise enhancement.

[0022] S2. Based on the enhanced equalization image, a structured list is obtained through feature construction and structuring. Existing technologies generally use a single, single-stage detector, which has a high rate of missed detection for small targets, such as cigarette butts and severely occluded targets, such as individuals in dense crowds. Furthermore, the classification results of the detection boxes are coarse and cannot provide accurate detection results, such as detailed status information like vehicles not turned off or people not wearing work clothes. To solve the above problems, the specific implementation steps are as follows: S21. Construct a feature pyramid fusion layer based on the enhanced equalization image to obtain the fused feature pyramid. In this step, after inputting the enhanced equalization image, CSPDarknet first extracts initial feature maps at different scales, and then performs bidirectional fusion from top to bottom and bottom to top through the PANet structure. The specific fusion process is as follows: the upper layer feature map is upsampled and enlarged, and then convolved with the current layer feature map to form the fused feature layers. For example, for an enhanced night scene image, the fused feature pyramid contains three levels: P3, P4, and P5. The P3 level retains high-resolution detail data, such as the brick texture of building exteriors and the edge contours of streetlights. The P4 level integrates medium-scale semantic data, such as the relative positions of vehicles and pedestrians in the scene. The P5 level contains low-resolution global data, such as the overall brightness of the scene. The distribution and spatial structure, through upsampling and channel concatenation operations, preserve local details and integrate global semantics in each level of features. The final output fused feature pyramid can simultaneously support detail enhancement and scene understanding, providing richer feature representations for subsequent image processing tasks. Among them, CSPDarknet is a backbone network based on Darknet53. Its core adopts a cross-stage partial connection (CSP) structure. By dividing the feature map into two parts and processing them in parallel, one part is processed by convolution, and the other part is concatenated after residual blocks, which reduces computational redundancy and enhances feature reuse. PANet is a path aggregation network. It innovates a bidirectional feature fusion mechanism based on the feature pyramid network FPN. It adds a bottom-up path, such as downsampling from P2 to P5 through 3×3 convolution and fusing with high-level features to form a bidirectional processing from semantics to details. This will not be elaborated here.

[0023] S22. Construct an attention prediction layer and perform attention processing based on the fused feature pyramid to obtain a candidate target set. In this step, channel attention is first calculated for each level of the fused feature pyramid feature map. Global average pooling and max pooling are used to extract the overall semantic information and local saliency information of the feature map, respectively. After fusion by a multilayer perceptron, channel weights are generated to enhance the feature response of important channels. Then, spatial attention is calculated, and the results of average pooling and max pooling are concatenated in the spatial dimension. The importance weights of spatial location are extracted through convolutional layers to focus on key areas. For example, in the P3 high-resolution level, channel attention will enhance channels related to target edges and textures, and spatial attention will highlight the local area where the target is located. The two are superimposed to form an enhanced feature map. When the detection head predicts bounding boxes on the enhanced feature map, it can more accurately locate the target position, such as the outline of pedestrians in a night scene, and estimate the target size, such as the width-to-height ratio of a vehicle. It outputs a set of candidate boxes containing position coordinates, size, confidence and initial category, such as simultaneously detecting vehicles, streetlights and pedestrians on the street, and the bounding boxes are close to the target edge to avoid missed detections or false detections.

[0024] S23. Based on the candidate target set and the enhanced equalization image, a structured list is obtained through fine-grained classification and structuring. S231. Construct a cropping and alignment layer. Based on the candidate target set and the enhancement-equalization image, perform cropping and alignment processing to obtain an aligned image set. In this step, candidate boxes are first filtered based on confidence and overlap, retaining only those with a confidence level higher than a preset threshold and an overlap with other boxes lower than the threshold. For example, in night scene detection, a vehicle candidate box with a confidence level of 0.8 and an IoU with a pedestrian box less than 0.3 is retained. Then, the corresponding region of the retained box is cropped from the enhancement-equalization image, and the region size and position are adjusted using an affine transformation matrix through translation, etc. Operations such as rotation and scaling align targets of different sizes to a standard size. For example, the vehicle area is scaled from the original 200 x 150 pixels to 128 x 128 pixels while maintaining the aspect ratio. In the final output set of aligned images, each target area retains its original detailed features, such as license plate texture and pedestrian clothing texture. The size alignment also eliminates the processing deviation caused by the difference in target size. For example, all vehicle targets are standardized to a uniform size, which facilitates subsequent fine recognition or classification, and finally yields the set of aligned images.

[0025] S232. Construct a residual classification layer and perform multi-task residual processing based on the aligned image set to obtain a fine-grained state probability distribution. In this step, the skip connection characteristic of the residual block is first utilized to add the input features to the features processed by the convolutional layer. For example, when the aligned vehicle image is processed by the residual block, the original edge information is superimposed with the texture features extracted by convolution, which preserves details and enhances semantic expression. Then, a multi-task classification head is set at the top of the network. Each classification head maps the final feature vector to the corresponding task space through a linear transformation, and then applies the Softmax function to transform it into a probability distribution. For example, after alignment... The image of the gas station area is processed by a residual network to extract features. The gas station status classification head outputs the probability distribution of normal, suspicious (such as dangerous actions by personnel), and abnormal (such as making phone calls or smoke) states. For example, the probability of normal is 0.7, the probability of suspicious state is 0.2, and the probability of abnormal state is 0.1. At the same time, the personnel clothing classification head outputs the probability distribution of uniforms, casual clothes, and protective clothing for personnel areas in the same image. This design avoids information decay in deep networks through residual connections, while the multi-task classification head achieves accurate identification of different state dimensions. The final fine-grained state probability distribution provides a reliable basis for subsequent decision-making.

[0026] S233. The fine-grained state probability distribution is encapsulated in a structured manner to obtain a structured list. In this step, for each high-confidence target bounding box, its initial category is semantically validated against the corresponding fine-grained state with the highest probability. If the initial category, such as "vehicles in a gas station," semantically matches the state with the highest probability, such as "driving" (0.9 probability), it is directly merged into the final category and state label. If there is a conflict, such as "personnel" in the initial category but "faulty vehicle" in the state with the highest probability, a manual validation process is triggered. For example, when a vehicle bounding box is detected, if the initial category "vehicles" matches the state with the highest probability "driving," the final label is "vehicle in motion," with a confidence level of 0.9. Simultaneously, the structured list summarizes the bounding box coordinates, final category, state, and confidence level by target ID, such as "driving" (0.9) and "stationary" (0.1), and supports... Multiple states coexist, such as personnel simultaneously labeled with uniform state 0.8 and casual wear state 0.2. The final output structured target list includes unique target identifiers, precise bounding boxes, final fusion categories, multi-dimensional states and their confidence levels, forming a standardized data structure that can be directly used for decision-making. This invention cascades a high-precision target detection model with a fine-grained state classification model. First, a detection model with optimized feature pyramids and attention mechanisms is responsible for locating all potential targets with high recall. Then, for each detection region, a dedicated state classification model performs secondary deep analysis and attribute discrimination. This combination of coarse localization and fine discrimination significantly improves the completeness of target detection and the accuracy of state recognition in complex and crowded scenes, greatly improving the overall recognition rate on the existing basis and solving the problems of missed detection and false detection of small and dense targets.

[0027] S3. Based on the structured list and enhanced equalization image, anomaly texture synthesis is performed to obtain synthetic anomaly images and anomaly annotations. These synthetic anomaly images and anomaly annotations are then integrated into a dangerous behavior recognition dataset. Existing technologies in image processing generally suffer from problems such as unclear segmentation boundaries, lack of realism in texture generation, and obvious splicing marks in the fusion region, making it difficult to meet the requirements for generating highly realistic anomaly scenes. To solve these problems, the specific implementation steps are as follows: S31. Based on the enhanced equalization image, semantic segmentation is performed to obtain a semantic segmentation mask. In this step, after inputting the enhanced equalization image, DeepLabv3+ first uses dilated convolution to expand the receptive field of the convolution kernel, capturing contextual information at different distances in the image. Simultaneously, it integrates multi-scale features through a spatial pyramid pooling module, such as parallel processing of dilated convolution branches with different dilation rates. Finally, features are fused through 1-by-1 convolution and the Softmax function is applied to assign a class probability to each pixel. For example, in the enhanced gas station fire image, the segmentation mask accurately distinguishes areas such as burning vehicles, pedestrians, and the gas station building. Pedestrian pixels are assigned the highest pedestrian class probability, such as 0.95. Pixels in the vehicle region highlight the vehicle category (e.g., 0.88), while pixels in the gas station building region enhance the building category (e.g., 0.92). This process ultimately generates a semantic segmentation mask containing the category label for each pixel, achieving complete semantic understanding from the pixel level to the scene level. DeepLabv3+ is a commonly used advanced model in semantic segmentation, employing an encoder-decoder architecture. Its core is to expand the receptive field and preserve spatial details through dilated convolutions. Spatial pyramid pooling modules, such as ASPP, extract features in parallel through multiple dilated convolution branches with different dilation rates, fusing contextual information at different scales, such as details of small objects and large scene structures. Finally, the decoder refines the boundaries. Further details are omitted here.

[0028] S32. Based on semantic segmentation masks and structured lists, local anomalous texture patches are obtained through texture anomaly processing. In this step, an initial anomalous texture, such as a flame or oil leak, is generated at the specified anomalous location in the target list using a physical model, such as a fuel tank opening. Subsequently, a generative adversarial network, such as a GAN, is used to make the initial texture more realistic. The generator generates candidate textures based on physical texture features and random noise, and the discriminator optimizes the generation effect by comparing it with the texture characteristics of real scenes, such as light and shadow reflection and edge blurring. Finally, a local anomalous texture image patch that is seamlessly integrated with the background is output. For example, at the location of a vehicle's fuel tank opening, the physical model generates an initial flame texture. The generative adversarial network learns the dynamic light and shadow changes of real flames, such as orange-red gradients and smoke diffusion patterns, and makes it more realistic, so that the flame texture is highly matched with the high gloss reflection characteristics of the surrounding metallic paint, forming a highly realistic local flame anomalous texture image patch. This achieves a dual improvement in the physical realism and visual naturalness of the anomalous effect. The generator and discriminator are commonly used processing methods in adversarial networks, which will not be elaborated here.

[0029] S33. Perform consistency adjustments on the enhanced equalization image and local anomalous texture patches to obtain a synthesized anomalous image and anomalous annotations, and generate a hazardous behavior recognition dataset. In this step, within a specified region of the original enhanced image, the Poisson equation is used to solve for the optimal fusion solution under gradient field constraints. The Poisson equation is a commonly used second-order partial differential equation describing the relationship between the gradient field and the source term, which will not be elaborated here. Boundary gradient matching maintains the continuity of region edges, while adjusting lighting parameters to ensure a natural transition between local anomalous texture patches, such as flames and backgrounds, or the light and shadow reflections and shadow distributions on the metal surface of a vehicle. For example, when fusing anomalous flame textures at the fuel tank opening, the Poisson equation uses the fuel tank opening edge as a boundary constraint. Gradient field matching seamlessly connects the orange-red gradient of the flames with the high-gloss reflectivity of the vehicle's metallic paint, while simultaneously adjusting the flame region... The brightness distribution is consistent with the surrounding ambient lighting, ultimately generating a visually natural, seamless synthetic anomalous image, along with anomalous region annotation information, which is the synthetic anomalous image and the anomalous annotation. These synthetic anomalous images and anomalous annotations are integrated together to generate a dangerous behavior recognition dataset. In the scene analysis stage of this invention, dilated convolution and spatial pyramid pooling are used to achieve accurate fusion of multi-scale features, solving the problems of blurred boundaries and missed detection of small targets. In the anomalous texture generation stage, physical models and realism are combined. When generating anomalous textures such as flame leaks at specified locations, the physical characteristics are ensured to be realistic, and the matching degree of light and shadow reflection with the background is optimized through the discriminator. This breaks through the limitations of traditional texture synthesis, which is stiff and has obvious fusion traces. Through boundary gradient constraints and lighting consistency adjustment, seamless fusion of anomalous texture blocks with the original image is achieved, avoiding the problems of edge breakage and lighting inconsistency.

[0030] S4. Convert the dangerous behavior recognition dataset and structured list into a structured list sequence, refine the structured list sequence and perform early warning detection to obtain behavior labels and early warning signals. Existing technologies often suffer from fragmented spatiotemporal feature extraction, such as focusing on only a single dimension, relying on a single indicator for multi-target tracking, being susceptible to occlusion interference when using only IoU, lacking contextual association leading to misclassification, and being unable to dynamically adapt to changing environments. To solve these problems, the specific implementation steps are as follows: S41. Based on the structured list sequence, a preliminary scheme set is obtained through high- and low-frequency information processing. In this step, after inputting the structured list sequence, the image sequence is processed at a low frame rate through the Slow path in the SlowFast network to capture the appearance features of targets such as vehicles and pedestrians, such as color and shape. The Fast path processes the same sequence at a high frame rate to capture the motion trajectory of the targets, such as vehicle speed and the direction of movement of gas station staff. The two paths are connected laterally to deeply fuse appearance information and motion information. Specifically, this is achieved through dynamic weighted fusion. The dynamic weights can be generated based on the amount of data in the Slow and Fast paths to form a feature map that combines detail and dynamics. Finally, temporal convolution is used to scan and generate the fused features. The initial approach combines an attention mechanism to automatically focus on key behavioral periods, such as sudden braking of vehicles or sudden stops by staff, generating more accurate preliminary plans. For example, in surveillance video, the system can identify behaviors such as a vehicle traveling at a medium speed from time t1 to t2 or a staff member crossing a gas station from time t3 to t4. Each plan includes a start time, an end time, and the corresponding behavior category confidence score, providing accurate spatiotemporal localization and category basis for subsequent behavior recognition. Among these, the SlowFast network is a commonly used two-stream architecture for video analysis, temporal convolution is a commonly used method for fusing features in the time dimension, and the attention mechanism is a commonly used method for dynamically adjusting feature weights and focusing on key behavioral periods, which will not be elaborated upon here.

[0031] S42. Perform multi-target association processing on the preliminary scheme set and the structured list sequence to obtain a multi-target trajectory feature set. In this step, within the spatiotemporal range of the preliminary scheme set, calculate the intersection-union ratio (IUGR) for targets in each frame. This reflects the degree of positional overlap and the cosine similarity with appearance features, reflecting visual consistency such as color and texture. After weighted fusion, an association index is formed. The Hungarian algorithm is used to match associated targets across frames to form a coherent trajectory, such as the continuous movement path of a vehicle from t1 to t5. Finally, spatiotemporal features such as vehicle color change sequences and speed curves are converged along the trajectory to generate a multi-target trajectory feature set containing trajectory features. For example, in a monitoring scenario, the system can identify the positional association of the same vehicle in different frames. By fusing the IUGR with appearance features, the trajectory coherence is ensured, and spatiotemporal features such as vehicle color and speed are converged to form a complete target trajectory feature set, providing accurate trajectory-level feature support for subsequent behavior analysis. The Hungarian algorithm is a classic algorithm commonly used to solve bipartite graph matching problems. It gradually expands the matching by finding augmenting paths and finally achieves the optimal matching. It will not be elaborated here.

[0032] S43. Based on the multi-target trajectory feature set, perform enhancement, refinement, and early warning processing to obtain behavior labels and early warning signals; S431. Based on the multi-target trajectory feature set, spatiotemporal enhancement processing is used to obtain the context enhancement scheme features. In this step, the trajectory features of targets such as gas station attendants, vehicles, and gas pumps are transformed into a feature token sequence according to the time series using a spatiotemporal Transformer encoder-decoder structure. The spatiotemporal dependency weights between each token are calculated using a self-attention mechanism. Through linear projection and dot product operations of query, key, and value, the interaction relationships between targets are dynamically captured, such as the relationship between the gas station attendant approaching the fuel tank opening and the vehicle starting. Then, the context information is aggregated through multi-layer encoding to finally generate a context enhancement scheme containing the spatiotemporal associations between targets. Features, such as those used in gas station safety management, can be enhanced when a gas station attendant approaches the fuel tank opening at time t1. The self-attention mechanism strengthens the association weight between the attendant's trajectory and the fuel tank opening feature. If the vehicle starts at time t2, the system uses trajectory feature fusion to identify the contextually enhanced features of the attendant approaching the fuel tank opening and the vehicle starting during the t1-t2 period. This provides timely warnings of potential safety risks, such as refueling without turning off the engine, and offers accurate spatiotemporal correlation for gas station safety monitoring. Linear projection and dot product operations are commonly used feature data processing methods. Multi-layer coding aggregation of contextual information is a commonly used method in coding aggregation, which will not be elaborated here.

[0033] S432. The context enhancement scheme features are classified and refined to obtain the refined scheme and classification probability. In this step, the input context enhancement scheme features are first linearly transformed by a fully connected layer. Specifically, the feature data is multiplied by the weight matrix and then a bias is added for transformation. The Softmax function is used to map the feature data to the probability distribution of various behaviors to complete the behavior classification. At the same time, another fully connected layer performs a linear transformation on the same feature to generate the adjustment amount of the proposal start and end time. For example, the start and end time adjustment amount is equal to the context enhancement scheme feature multiplied by the weight and the bias, thus completing the fine-tuning of the time boundary. Finally, the refined scheme and classification probability are output. For example, in the safety management of gas stations, after the system recognizes the context features of the gas station attendant approaching the fuel tank opening and the vehicle starting, it is classified as a refueling behavior with a probability of 95% and the proposal time is adjusted to the t1-t3 period. A refined scheme containing precise time boundaries and high confidence categories is generated to provide a reliable basis for safety warnings.

[0034] S433. Based on the hazardous behavior identification dataset, refined schemes, and classification probabilities, perform early warning identification processing to obtain behavior labels and early warning signals. In this step, for each refined scheme, if the highest value of the classification probability exceeds the preset high-risk threshold and the behavior category belongs to the hazardous behavior identification dataset, an early warning is immediately triggered, generating information including behavior type, risk level, location, time range, and involved targets. Simultaneously, based on specific scenario rules, such as continuous smoking exceeding five seconds, a comprehensive judgment is made to ensure the accuracy of the early warning. For example, in gas station safety management, if the system identifies that the classification probability of refueling without turning off the engine reaches 98%, exceeding the high-risk threshold and belonging to the preset high-risk behavior set, a high-risk level early warning is immediately triggered, with the location labeled as the No. 3 fuel dispenser area, the time range as 10:00-10:05, and the involved target as a fuel dispenser. After confirming that there are no other safe actions such as engine shutdown, the system outputs a complete behavior label and warning signal, providing a clear basis for real-time safety intervention. The high-risk threshold can be set by comprehensively considering historical data, including VOCs concentration at the gas station, non-methane total hydrocarbon density, and behavior probability thresholds. It is a multi-dimensional comprehensive threshold that makes judgments based on feature types. This invention enhances tracking robustness by fusing appearance and motion information through dual paths and combining trajectory association algorithms based on appearance features. It generates context-enhanced features based on spatiotemporal dependencies and achieves accurate behavior classification and flexible warning triggering through dynamic thresholds. This effectively solves the shortcomings of traditional methods in feature fusion, trajectory association, context reasoning, and adaptability, and improves the recognition accuracy and response efficiency of dangerous behaviors in complex dynamic scenarios.

[0035] S5. Based on the warning signal, structured list sequence, and enhanced equalization image, feedback information is reported through retrieval and similarity processing. Existing technologies only archive by time and text tags, making the query unintuitive. Image retrieval based on global features has low discrimination and precision in similar scenarios at gas stations, and lacks a linkage feedback mechanism with real-time warnings. To solve the above problems, the specific implementation steps are as follows: S51. Construct a multi-branch deep network layer based on the warning signal, structured list sequence, and enhanced equalization image to obtain a multi-granularity feature vector group. In this step, firstly, global scene features are extracted from the enhanced equalization image. A pre-trained deep convolutional network, such as ResNet, is used to capture the overall semantic information of the image. After global average pooling, it is compressed into a fixed-dimensional global description vector. Secondly, for key targets in the structured list, such as license plates, faces, and fire extinguishers, their regional images are extracted and processed one by one through a local feature network. After max pooling, they are aggregated into local target feature vectors. Finally, the behavior type, risk level, and main target category in the structured list are embedded and encoded. The event semantic information is fused by vector concatenation. The final output is a multi-granularity feature vector group containing global scene, local target, and event semantics, providing rich feature support for subsequent analysis through hierarchical encoding. ResNet is a commonly used deep convolutional network, which will not be elaborated here.

[0036] S52. Based on the multi-granularity feature vector group, a normalized feature vector is obtained through cross-modal feature fusion processing. In this step, the global scene, local target, and event semantic features in the multi-granularity feature vector group are mapped to the same dimension through a fully connected layer. Then, the shared attention weight matrix is ​​multiplied with each feature and the fusion weight is calculated by Softmax. Through weighted summation, such as multiplying the global scene feature by 0.3, adding 0.5 multiplied by the local target feature, and adding 0.2 multiplied by the event semantic feature, a fused feature is generated. Then, a bottleneck layer performs a linear transformation, specifically multiplying the fused feature by the weight matrix and adding bias, ReLU activation, and L2 normalization to output a normalized feature vector. For example, in the safety management of gas stations, the system integrates the global scene features of the refueling area, the local target features of the license plate, and the semantic features of the un-extinguished refueling event. Through weighted fusion and dimensionality reduction normalization, a 128-dimensional joint feature is generated, which improves the accuracy of encoding multimodal information association and provides unified feature support for risk assessment.

[0037] S53. A hierarchical indexing process is constructed on the normalized feature vectors to obtain a hierarchical index structure. In this step, an HNSW graph is first built in the joint feature vector space as a coarse index. Each feature vector is treated as a graph node, and a navigable network structure is formed through similarity connections. When inserting a new node, a greedy search is used to locate the nearest neighbor and establish a connection, ensuring rapid global similarity. Secondly, a product quantization (PQ) fine indexing process is performed on the high-dimensional feature vectors. The vectors are divided into multiple sub-segments according to their dimensions, and each sub-segment is independently clustered to generate a codebook. The original vector is represented by the cluster center ID sequence corresponding to the sub-segment, achieving data compression and efficient storage. For example, in the gas station safety event database, the system stores the 128-dimensional joint feature vectors of historical events in the HNSW graph, supporting millisecond-level similar event retrieval, such as finding similar refueling events where the engine is not turned off. Simultaneously, the storage space is compressed to one-tenth of its original size through the PQ codebook, ensuring rapid response and scalable storage for large-scale event databases, providing efficient indexing support for risk trend analysis.

[0038] S54. Similarity retrieval is performed based on normalized feature vectors and a hierarchical index structure to obtain a list of retrieval results. In this step, a coarse search is first performed on the HNSW graph based on the normalized feature vectors. Starting from the entry point, a greedy traversal is used to find the L nearest neighbor candidate nodes. Euclidean distance is calculated to locate the initial candidate set. Then, the candidate vectors are refined using PQ (Problem-Quickness) processing. The quantized encoding of the query vector and the database vector is used to calculate the asymmetric distance in subspaces. The sum of the squared differences between each sub-segment query vector and its corresponding cluster center is weighted to obtain the precise distance value. Combining this with a time decay factor (where the weight of recent events decays exponentially with time difference), the time weight is multiplied by the distance term, incremented by 1, and the reciprocal is taken to generate a similarity score. Finally, the similarity scores are sorted in descending order. The system sorts and outputs a Top-K search results list containing event IDs, scores, and metadata. The metadata is generated automatically when the event is triggered or by the system actively extracting key information. In the gas station safety management scenario, when events such as refueling with the engine still running are detected, the system will simultaneously collect the timestamp of the event, the unique identifier that triggered the alarm (such as the alarm ID), the list of associated targets (such as license plate numbers, face IDs), and structured information such as device sensor data (such as camera numbers and temperature values), and align them spatiotemporally with the corresponding enhanced equalized images and warning signals. For example, in gas station event retrieval, the system prioritizes returning recent refueling events with high similarity, supporting rapid risk case retrospective and trend analysis.

[0039] S55. Based on the search results list, dynamic threshold processing is applied to obtain a feedback report. In this step, the dynamic threshold is first determined based on the historical positive and negative sample score distribution. Specifically, it is the average historical negative sample score minus the product of the sensitivity coefficient and the difference between the average positive and negative sample scores. For example, if the average negative sample score is 0.6, subtracting 0.3 and multiplying by 0.2 gives 0.54. Then, the highest score in the search results list is compared with this threshold. If the highest score exceeds the threshold, it is determined to be a highly similar event, generating an association report containing similarity point comparisons and historical processing records, and pushing it to the management console. If it does not exceed the threshold, it is considered a new type of event, storing the complete data of the current event in the database and updating the index structure. Simultaneously, a new event alert is triggered. If necessary, the process returns to the dangerous behavior identification workflow for manual review, forming a detection... This invention employs a complete process of searching, judging, providing feedback, and updating to continuously improve the system's ability to identify and respond to new types of events. It integrates deep global features of images, local features of key targets, and semantic tags of events to form a unified feature vector for joint indexing and similarity measurement. The subsequent process is determined based on the similarity of the search results. High-similarity results are directly linked to historical cases to aid decision-making, while low-similarity results are archived as new events and trigger database updates. This significantly improves the accuracy of tracing complex events and the system's self-evolutionary capabilities, enabling gas station managers to quickly find historical data with similar vehicles, similar personnel behaviors, or similar violations. Simultaneously, it accumulates rare event samples for the system, greatly improving retrieval recall and processing efficiency.

[0040] Example 2: Due to the high false negative rate of existing technologies using a single, single-stage detector for small targets and severely occluded targets, and the resulting coarse bounding box classification, accurate detection results cannot be provided. For further information, please refer to [link to relevant documentation]. Figure 2 The diagram shown is a structural block diagram of the smart gas station safety supervision system based on computer vision provided in this embodiment. The system includes a balance module, a structure module, an anomaly labeling module, an early warning module, and a classification feedback module. The equalization module is used to acquire the original single-frame image, perform image enhancement processing on the original single-frame image, and obtain an enhanced equalized image. The structuring module is used to construct and structure the list based on the enhanced equalization image through feature construction and structuring. The anomaly labeling module is used to synthesize anomaly images and anomaly labels based on a structured list and augmented equalization images through anomaly texture synthesis, and to integrate all the synthesized anomaly images and anomaly labels into a dangerous behavior recognition dataset. The early warning module is used to convert the dangerous behavior identification dataset and structured list into a structured list sequence, refine the structured list sequence and perform early warning detection to obtain behavior labels and early warning signals; The classification feedback module is used to obtain feedback information reports based on the warning signal, structured list sequence, and enhanced equalization image through retrieval and similarity processing.

[0041] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc.

[0042] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A smart gas station safety supervision method based on computer vision, characterized in that, The steps of this method are as follows: acquire the original single-frame image, perform image enhancement processing on the original single-frame image, and obtain an enhanced equalization image; Based on the enhanced equalization image, a structured list is obtained through feature construction and structuring. Based on structured lists and enhanced equalization images, synthetic anomalous images and anomalous annotations are obtained through anomalous texture synthesis. The synthetic anomalous images and anomalous annotations are then integrated into a dangerous behavior recognition dataset. The dangerous behavior recognition dataset and structured list are converted into a structured list sequence. The structured list sequence is then refined and an early warning detection is performed to obtain behavior labels and early warning signals. Feedback information reports are obtained through retrieval and similarity processing based on warning signals, structured list sequences, and enhanced equalization images.

2. The intelligent gas station safety supervision method based on computer vision according to claim 1, characterized in that, Image enhancement processing is performed on the original single-frame image, including: enhancing the original single-frame image through an adversarial network to obtain an illumination-enhanced image; A denoised image is obtained by blurring and denoising the enhanced illumination image. Histogram equalization is performed on the denoised image to obtain an enhanced equalized image.

3. The intelligent gas station safety supervision method based on computer vision according to claim 1, characterized in that, Based on the enhanced equalization image, feature construction and structuring processing are performed, including: constructing a feature pyramid fusion layer based on the enhanced equalization image to obtain a fused feature pyramid; An attention prediction layer is constructed, and attention processing is performed based on the fused feature pyramid to obtain a set of candidate targets; A structured list is obtained by fine-grained classification and structuring based on the candidate target set and the enhanced equalization image.

4. The intelligent gas station safety supervision method based on computer vision according to claim 3, characterized in that, Based on the candidate target set and the enhanced equalization image, fine-grained classification and structured processing are performed, including: constructing a cropping and alignment layer, and performing cropping and alignment processing based on the candidate target set and the enhanced equalization image to obtain an aligned image set; Construct a residual classification layer, perform multi-task residual processing based on the aligned image set, and obtain fine-grained state probability distribution; The fine-grained state probability distribution is encapsulated in a structured manner to obtain a structured list.

5. The intelligent gas station safety supervision method based on computer vision according to claim 1, characterized in that, Based on structured lists and enhanced equalization images, anomalous texture synthesis is performed, including: Based on the enhanced equalization image, a semantic segmentation mask is obtained through semantic segmentation processing. Based on semantic segmentation masks and structured lists, local abnormal texture patches are obtained through texture anomaly processing. Consistency adjustment is performed on the enhanced equalized image and local anomalous texture patches to obtain synthetic anomalous images and anomalous annotations, and a dangerous behavior recognition dataset is generated.

6. The intelligent gas station safety supervision method based on computer vision according to claim 1, characterized in that, The structured list sequence is refined and an early warning detection is performed, including: processing high and low frequency information based on the structured list sequence to obtain a preliminary set of solutions; Multi-target correlation processing is performed on the preliminary scheme set and the structured list sequence to obtain a multi-target trajectory feature set; Enhancement, refinement, and early warning processing are performed based on the multi-target trajectory feature set to obtain behavior labels and early warning signals.

7. The intelligent gas station safety supervision method based on computer vision according to claim 6, characterized in that, Enhancement, refinement, and early warning processing are performed based on the multi-target trajectory feature set, including: obtaining context enhancement scheme features through spatiotemporal enhancement processing based on the multi-target trajectory feature set; The features of the context enhancement scheme are classified and refined to obtain the refined scheme and classification probability. Based on the dangerous behavior recognition dataset, refinement scheme, and classification probability, early warning recognition processing is performed to obtain behavior labels and early warning signals.

8. The intelligent gas station safety supervision method based on computer vision according to claim 1, characterized in that, Based on the warning signal, structured list sequence, and enhanced equalization image, the process involves retrieval and similarity processing, including: constructing a multi-branch deep network layer based on the warning signal, structured list sequence, and enhanced equalization image to obtain a multi-granularity feature vector group; Based on the multi-granularity feature vector group, a normalized feature vector is obtained through cross-modal feature fusion processing. The normalized feature vectors are processed to construct a hierarchical index, resulting in a hierarchical index structure. Similarity retrieval is performed based on normalized feature vectors and a hierarchical index structure to obtain a list of retrieval results; The search results are processed using dynamic thresholding to generate a feedback report.

9. A system applied to the computer vision-based smart gas station safety supervision method according to any one of claims 1-8, characterized in that, The system includes: The equalization module is used to acquire the original single-frame image, perform image enhancement processing on the original single-frame image, and obtain an enhanced equalized image. The structuring module is used to construct and structure the list based on the enhanced equalization image through feature construction and structuring. The anomaly labeling module is used to synthesize anomaly images and anomaly labels based on a structured list and augmented equalization images through anomaly texture synthesis processing, and then integrate the synthesized anomaly images and anomaly labels into a dangerous behavior recognition dataset. The early warning module is used to convert the dangerous behavior identification dataset and structured list into a structured list sequence, refine the structured list sequence and perform early warning detection to obtain behavior labels and early warning signals; The classification feedback module is used to obtain feedback information reports based on the warning signal, structured list sequence, and enhanced equalization image through retrieval and similarity processing.