Method and system for detecting illegal operation of cigarette station based on image recognition
By obtaining and analyzing image sequences in the tobacco station operation scenarios and using the violation discrimination model to detect illegal operations, the time limitations and subjectivity of manual inspections are solved, and the accurate detection and timely disposal of illegal operations are achieved, and the operation efficiency and management level of the tobacco station are improved.
Patent Information
- Application Number
- CN202510965611.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Violated operation inspections in tobacco station operations rely on manual inspections, which have problems such as time limitations, subjectivity and low efficiency, making it difficult to achieve all-round and real-time supervision and timely disposal.
By obtaining the continuous image sequence of the smoke station operation scene, performing operation feature analysis, extracting the dynamic behavior characteristics of the operator and the standard constraint characteristics of the operation scene, and using the pre-trained violation discrimination model for compliance matching, generating violation type identification and positioning information, and generating warning instructions to trigger the violation handling process.
Accurate detection and timely handling of illegal operations has been achieved, the operation efficiency and management level of tobacco stations have been improved, and the standardization and safety of operations have been ensured.
Smart Images

Figure CN120452070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for detecting illegal operations at a tobacco station based on image recognition. Background Art
[0002] In the actual operational management of tobacco stations, ensuring that operators strictly adhere to standardized operating procedures is crucial for ensuring tobacco leaf quality, maintaining operational safety, and improving overall operational efficiency. However, tobacco stations currently rely primarily on manual inspections and on-site supervision to detect illegal operations. Due to the complex operational scenarios of tobacco stations, involving multiple operational links and numerous operators, manual inspections cannot provide comprehensive, real-time supervision. On the one hand, manual inspections are time-limited and cannot continuously monitor the operational process, making it easy to miss some momentary illegal operations. On the other hand, manual judgment is easily influenced by subjective factors, and different supervisors may have different understandings of the standards for illegal operations, resulting in inconsistent and inaccurate judgment results. In addition, manual inspections require a large amount of manpower costs and are relatively inefficient. They cannot locate and deal with illegal operations in a timely manner, which makes it difficult to meet the needs of efficient management of modern tobacco stations. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for detecting illegal operations of a cigarette station based on image recognition, the method comprising: Acquire a continuous image sequence of the smoke station operation scene, wherein the continuous image sequence is composed of multiple frames of images acquired in time sequence, and each frame of the image contains visual information of the operator's body movements, material contact status, and tool use position; Performing operation feature analysis on the continuous image sequence to obtain dynamic behavior features of the operator and standard constraint features of the operation scene; Inputting the dynamic behavior features and the regulatory constraint features into a pre-trained violation discrimination model for compliance matching, and generating a discrimination result including a violation type identifier; Locating the time starting point, time ending point and spatial effect area of the illegal operation in the continuous image sequence according to the discrimination result; Based on the time starting point, time ending point and spatial action area, a warning instruction containing violation location information is generated, and the warning instruction is sent to the tobacco station supervision terminal to trigger the violation handling process.
[0004] On the other hand, an embodiment of the present invention also provides a tobacco station illegal operation detection system based on image recognition, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0005] Based on the above aspects, by acquiring continuous image sequences of the tobacco station operation scene, key visual information such as the operator's body movements, material contact status, and tool usage position during the operation process is captured. Operational feature analysis processing is performed on the continuous image sequences, which can accurately extract the dynamic behavioral characteristics of the operators and the standard constraint characteristics of the operation scene, making the description of the operation process more accurate and comprehensive. The dynamic behavioral characteristics and standard constraint characteristics are input into the pre-trained violation discrimination model for compliance matching. Leveraging the model's powerful learning and judgment capabilities, it can quickly and accurately generate discrimination results containing violation type identification, greatly improving the efficiency and accuracy of violation operation detection. Based on the discrimination results, the time start and end points of the violation operation in the continuous image sequence and the spatial impact area in the image are located, achieving precise positioning of the violation operation. Based on this positioning information, a warning instruction containing the violation location information is generated and sent to the tobacco station supervision terminal to trigger the violation handling process. It can timely and effectively intervene and handle the violation operation, prevent the further expansion and impact of the violation operation, thereby ensuring the standardization and safety of tobacco station operations and improving the overall operational efficiency and management level of the tobacco station. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 It is a schematic diagram of the execution flow of the method for detecting illegal operations of a cigarette station based on image recognition provided by an embodiment of the present invention.
[0007] Figure 2 Schematic diagram of exemplary hardware and software components of a cigarette station illegal operation detection system based on image recognition provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0008] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a method for detecting illegal operations at a cigarette station based on image recognition provided by an embodiment of the present invention. The method for detecting illegal operations at a cigarette station based on image recognition is introduced in detail below.
[0009] Step S110: Acquire a continuous image sequence of the tobacco station operation scene, wherein the continuous image sequence is composed of multiple frames of images collected in chronological order, and each frame of the image contains visual information of the operator's body movements, material contact status, and tool use position.
[0010] To capture the continuous image sequence at the smoke station operation site, image acquisition equipment must be deployed in appropriate locations. This equipment must possess high resolution, high frame rate, and excellent low-light performance to ensure clear and accurate image capture under varying lighting conditions and operating environments. For example, cameras should be strategically placed above or around material storage areas, tool operation areas, and primary operator activity areas to ensure that all key areas are captured.
[0011] Image acquisition equipment captures images at pre-set time intervals. The time intervals must be set based on the speed of change in the work action and the burden of data processing. If the time interval is too long, key details of the action may be missed; if the time interval is too short, a large amount of data will be generated, increasing the difficulty of subsequent processing. Each captured image contains a wealth of visual information. For example, the images may show the worker's body movements, such as arms raised, lowered, stretched, and bent, and legs moving, standing, and squatting. In terms of material contact status, it can be seen whether the worker is contacting the material, picking up, carrying, or placing the material. The tool usage position can be clearly seen from the image to determine whether the tool is being used on the material or on a specific workbench. The multiple frames of images arranged in chronological order above constitute a continuous image sequence.
[0012] Step S120: performing operation feature analysis on the continuous image sequence to obtain the dynamic behavior features of the operator and the standard constraint features of the operation scene.
[0013] After obtaining a continuous image sequence of a smoke station operation scene, in order to determine whether the operator's operations are in compliance, it is necessary to conduct in-depth analysis of the operation characteristics of the continuous image sequence, thereby obtaining the operator's dynamic behavior characteristics and the standard constraint characteristics of the operation scene. The dynamic behavior characteristics can reflect the operator's movement characteristics and behavior patterns during the operation process, while the standard constraint characteristics reflect the standard requirements that the smoke station operation scene should comply with.
[0014] Step S121: performing personnel key point detection processing on each frame of the continuous image sequence, and extracting the pixel coordinates of the operator's shoulders, elbows, wrists and knees as basic action points through a human posture estimation algorithm.
[0015] For each frame in a continuous image sequence, the human pose estimation algorithm analyzes each frame at multiple levels to extract basic motion points. First, the algorithm performs a color space analysis of the image. Different body parts may have certain characteristic differences in color. For example, the distribution of skin color in the image can serve as a basis for initially locating the approximate position of the human body. Next, edge detection techniques are used to identify the outline of the human body. This edge information helps further determine the shape and approximate dimensions of the human body.
[0016] After determining the approximate position and outline of the human body, the human pose estimation algorithm uses a pre-trained model to accurately locate the shoulders, elbows, wrists, and knees. This model, trained on a large amount of sample data on human poses, can identify these key areas based on features such as texture and shape in the image. For example, the shoulder typically has distinct skeletal structural features, and the algorithm uses these features to locate the specific location of the shoulder in the image and extract its corresponding pixel coordinates. As for the wrist, due to its relatively small size and flexible movements, the algorithm combines the overall movement and shape of the hand to more accurately locate its pixel coordinates. In this way, the pixel coordinates of the operator's shoulders, elbows, wrists, and knees are extracted from each frame of the image as basic action points.
[0017] Step S122: performing key point tracking processing on adjacent image frames in the continuous image sequence, recording the coordinate changes of each basic action point in the continuous frames, and generating an action point motion trajectory sequence.
[0018] After obtaining the basic action points of each frame, it is necessary to track the basic action points in adjacent image frames to record their coordinate changes.
[0019] Step S1221: selecting the basic motion point coordinates of the operator's right wrist in the previous image frame as the historical coordinates, and selecting the basic motion point coordinates of the right wrist in the current image frame as the current coordinates.
[0020] In order to accurately record the movement of the basic action point of the right wrist, in the two adjacent frames of images, the coordinates of the basic action point of the right wrist in the previous image frame are selected as the historical coordinates, which represent the position of the right wrist at the previous moment; the coordinates of the basic action point of the right wrist in the current image frame are selected as the current coordinates, which reflect the position of the right wrist at the current moment.
[0021] Step S1222: Calculate the horizontal displacement difference and the vertical displacement difference between the historical coordinates and the current coordinates in the image to obtain the single-frame displacement vector of the action point.
[0022] After obtaining the historical coordinates and current coordinates of the right wrist, their horizontal and vertical displacement differences in the image are calculated respectively. The horizontal displacement difference is obtained by subtracting the horizontal coordinate of the historical coordinate from the horizontal coordinate of the current coordinate, which reflects the horizontal movement distance and direction of the right wrist. The vertical displacement difference is obtained by subtracting the vertical coordinate of the historical coordinate from the vertical coordinate of the current coordinate, which reflects the vertical movement of the right wrist. Combining these two displacement differences forms the single-frame displacement vector of the action point. This single-frame displacement vector of the action point not only contains the displacement size of the right wrist in this frame of image, but also contains the direction information of the displacement, which can comprehensively describe the movement of the right wrist in this frame of image.
[0023] Step S1223: Count the displacement vector sequences of the basic action points of the right wrist, left elbow and knee in the continuous image frames, and connect them in chronological order to generate a motion trajectory sequence of the action points.
[0024] For the fundamental action points of the right wrist, left elbow, and knee in a continuous image sequence, their single-frame displacement vectors are calculated according to the above steps. These displacement vectors are then concatenated in chronological order. For example, the displacement vector from the first frame to the second frame is recorded first, followed by the displacement vector from the second frame to the third frame, and so on. This forms a sequence of action point motion trajectories. This sequence can be viewed as a quantitative representation of the motion paths of these fundamental action points in the continuous image frames. By analyzing this sequence of action point motion trajectories, we can intuitively understand the motion trajectories and changes of these key parts of the operator during the operation.
[0025] Step S1224: Smoothing the motion trajectory sequence of the action points, eliminating abnormal fluctuations of the displacement vector caused by image acquisition errors by using a sliding average algorithm, and generating a smoothed motion trajectory sequence.
[0026] Because image acquisition can be affected by a variety of factors, such as lighting changes and device jitter, the sequence of action point motion trajectories may exhibit abnormal fluctuations in displacement vectors. To eliminate these fluctuations, a sliding average algorithm is used to smooth the sequence.
[0027] The sliding average algorithm sets a fixed-size window and slides it across the sequence of action point motion trajectories. For each displacement vector within the window, its average is calculated. For example, if the window size is 5 and the window slides to a certain position in the sequence, the average of the five displacement vectors before and after that position is calculated. This average is then used to replace the displacement vector at the center of the window. As the window slides across the sequence, the above processing is performed on each position in turn, ultimately resulting in a smoothed motion trajectory sequence. This processed sequence more accurately reflects the operator's actual movements and reduces interference caused by image acquisition errors.
[0028] Step S1225: performing directional consistency analysis on the smoothed motion trajectory sequence, calculating the changing trend of the direction of the continuous displacement vector, and generating a trajectory directional stability parameter.
[0029] When analyzing the directional consistency of a smoothed trajectory sequence, we can focus on the changes in the directions of consecutive displacement vectors. First, we can calculate the angle between adjacent displacement vectors, which reflects the degree of change in displacement direction. To more accurately analyze directional trends, we can statistically analyze the angles between multiple adjacent displacement vectors.
[0030] For example, calculate the angles between 10 consecutive displacement vectors and observe the changes in their values. If the fluctuations in the values of these angles are small, it indicates that the directional changes of the consecutive displacement vectors are small, the operator's movements have good directional consistency, and the motion trajectory is relatively stable. Conversely, if the angles fluctuate greatly, the movement stability is poor. Through comprehensive analysis of these angles, a trajectory directional stability parameter can be generated. This trajectory directional stability parameter can be obtained by performing certain mathematical operations on multiple angles, such as taking the standard deviation of these angles. A smaller standard deviation indicates better directional stability; a larger standard deviation indicates worse directional stability.
[0031] Step S1226: performing velocity continuity analysis on the smoothed motion trajectory sequence, calculating the gradient of the continuous displacement vector modulus, and generating a trajectory velocity stability parameter.
[0032] In addition to directional stability, the velocity continuity of the motion trajectory also needs to be analyzed. The modulus of the displacement vector represents the velocity of the underlying action point. When performing velocity continuity analysis on a smoothed motion trajectory sequence, the gradient of the modulus of the continuous displacement vector can be calculated.
[0033] First, calculate the difference between the moduli of adjacent displacement vectors, which reflects the change in speed. Then, divide the difference by the time interval to obtain the speed change gradient. For example, calculate the difference between the displacement vector modulus from the second frame to the third frame and the displacement vector modulus from the first frame to the second frame, and divide it by the time interval between the two frames to obtain the speed change gradient within this time period. If the change gradient of the continuous displacement vector modulus is small, it means that the change in motion speed is relatively gentle and the speed continuity is good; conversely, if the change gradient is large, it means that the speed changes drastically and the speed continuity is poor. Based on these analysis results, a trajectory speed stability parameter can be generated. The trajectory speed stability parameter can be a value that comprehensively considers multiple speed change gradients, such as taking the average of the speed change gradients in multiple consecutive time periods. The smaller the average value, the better the speed stability.
[0034] Step S1227: combining the smoothed motion trajectory sequence, trajectory direction stability parameter, and trajectory speed stability parameter to form a complete action point motion trajectory sequence.
[0035] Combining the smoothed trajectory sequence, the trajectory directional stability parameters, and the trajectory velocity stability parameters creates a complete action point trajectory sequence. This action point trajectory sequence not only includes the basic action point trajectory information, but also the trajectory directional stability and velocity stability information. This provides a more comprehensive and accurate description of the operator's motion characteristics at key points during the operation. For example, when determining whether an operator's operation complies with regulations, one can not only examine the trajectory itself, but also combine directional stability and velocity stability to make a comprehensive judgment.
[0036] Step S123: Based on the motion trajectory sequence of the action points, the operator's motion amplitude parameter is calculated. The motion amplitude parameter is the maximum displacement of the basic action points in consecutive frames; the motion frequency parameter is calculated. The motion frequency parameter is the number of displacements of the basic action points per unit time; and the motion continuity parameter is calculated. The motion continuity parameter is the minimum angle between consecutive displacement direction changes. The motion amplitude parameter, motion frequency parameter, and motion continuity parameter are combined to form a dynamic behavior feature.
[0037] Based on the complete sequence of action point motion trajectories, the operator's dynamic behavioral characteristics can be further calculated. For the motion amplitude parameter, the maximum displacement of the basic action points is found in consecutive frames. In the action point motion trajectory sequence, each basic action point has a series of displacement values. By comparing these displacement values, the maximum value is found as the motion amplitude parameter. For example, the displacement values of the right wrist in 10 consecutive frames are different. By comparing these 10 values, the maximum displacement value is determined as the motion amplitude parameter of the right wrist. This motion amplitude parameter reflects the maximum amplitude of the operator's movement.
[0038] The motion frequency parameter is calculated by counting the number of displacements of the basic motion points per unit time. In the motion point trajectory sequence, the moment each basic motion point shifts is recorded, and then the number of displacements within a fixed time unit is counted. For example, using a 1-second time unit as the time unit, the number of displacements of the right wrist within that 1-second time unit is counted. This number is used as the motion frequency parameter for the right wrist, reflecting the frequency of the operator's movements.
[0039] The motion continuity parameter calculates the minimum angle between successive changes in displacement direction. In the motion trajectory sequence of an action point, adjacent displacement vectors have certain angles between them, which reflect the changes in displacement direction. By traversing the angles between all adjacent displacement vectors, the minimum value is found and used as the motion continuity parameter. The smaller the angle, the better the operator's motion continuity and the smoother the transitions between movements. The combination of the motion amplitude parameter, the motion frequency parameter, and the motion continuity parameter forms the operator's dynamic behavior characteristics, which can comprehensively describe the operator's movement characteristics during the operation.
[0040] Step S124: performing scene semantic segmentation processing on the single-frame image in the continuous image sequence, and dividing the image into a material storage area, a tool operation area, and a safety isolation area through a semantic segmentation model.
[0041] The purpose of performing scene semantic segmentation on single-frame images in a continuous image sequence is to accurately classify different areas in the image in order to understand the layout and structure of the work scene.
[0042] Step S1241: pre-processing the single-frame image, reducing image noise through a Gaussian filtering algorithm, and enhancing edge contrast of materials, tools, and safety isolation areas to obtain a pre-processed image.
[0043] Before performing scene semantic segmentation, single-frame image preprocessing is required. The Gaussian filter algorithm is a linear smoothing filter method that performs a weighted average of each pixel in the image and its neighboring pixels based on a Gaussian function. The characteristic of the Gaussian function is that pixels closer to the center pixel have a greater weight, while pixels farther from the center pixel have a smaller weight. This weighted average effectively reduces noise interference in the image. For example, an image may contain some random salt-and-pepper noise points, but after Gaussian filtering, the impact of these noise points is reduced.
[0044] At the same time, the Gaussian filter algorithm also enhances the edge contrast of materials, tools, and safety isolation zones, highlighting their boundaries and making them more clearly distinguishable from the surrounding environment. This is because in the subsequent semantic segmentation, clear edge information helps to more accurately delineate different areas.
[0045] Step S1242: Input the preprocessed image into the feature extraction layer of the semantic segmentation model, and extract the local texture features and global context features of the image through a convolutional neural network.
[0046] The preprocessed image is fed into the feature extraction layer of the semantic segmentation model. Here, a convolutional neural network (CNN) extracts multi-scale features from the image. A CNN consists of multiple convolutional layers, pooling layers, and activation function layers.
[0047] The convolutional layer is the core layer of a CNN, using convolution kernels to perform sliding convolution operations on the image. Different convolution kernels can extract different features. For example, some convolution kernels can extract edge features of an image, while others can extract texture features. When extracting local texture features, the convolution kernel performs convolution on a local area of the image, capturing information such as the surface texture of the material and the detailed texture of the tool. For global context features, the network analyzes the overall information of the image through multiple layers of convolution and pooling operations. The pooling layer downsamples the output of the convolution layer, reducing the size of the feature map while retaining important feature information. Activation function layers, such as the ReLU function, introduce nonlinear factors to enhance the network's expressive power. Through the combination of these layers, convolutional neural networks can extract local texture features and global context features of an image.
[0048] Step S1243: Input the local texture features and global context features into the upsampling layer of the semantic segmentation model, restore the feature map size to be consistent with the input image through deconvolution operation, and generate a segmentation probability map containing pixel-level classification information.
[0049] Local texture features and global context features are input to the upsampling layer of the semantic segmentation model. Deconvolution is the core operation of the upsampling layer, which restores the size of the feature map to the same as the input image.
[0050] During the feature extraction process, convolution and pooling operations reduce the size of the feature map, while deconvolution gradually increases the size of the feature map through a series of convolution and interpolation operations. Specifically, the deconvolution operation uses a learnable deconvolution kernel to convolve the input feature map with the deconvolution kernel, while performing interpolation during the convolution process to increase the size of the feature map. During the deconvolution process, based on the feature information in the feature map, a probability value is generated for each pixel as belonging to different areas (material storage area, tool operation area, and safety isolation area). These probability values constitute the segmentation probability map, which contains pixel-level classification information and accurately reflects the probability of each pixel belonging to a different area.
[0051] Step S1244: performing threshold processing on the segmentation probability map, classifying pixels with probability values greater than a preset threshold into corresponding areas, and generating initial segmentation results of the material storage area, tool operation area, and safety isolation area.
[0052] When performing threshold processing on the segmentation probability map, a threshold can be set in advance. For each pixel in the segmentation probability map, the probability value of its belonging to a certain area can be compared with the preset threshold. If the probability value is greater than the preset threshold, the pixel is classified as the corresponding area. For example, if the probability value of a pixel belonging to the material storage area is greater than the preset threshold, the pixel is marked as part of the material storage area. By performing the above processing on all pixels in the segmentation probability map, the initial segmentation results of the material storage area, tool operation area and safety isolation area are generated. This initial segmentation result may have some small flaws, such as holes in the area, isolated noise points outside the area, etc., which require further post-processing.
[0053] Step S1245: Post-process the initial segmentation results, fill the holes in the area through morphological closing operations, and eliminate isolated noise points outside the area through morphological opening operations to generate the final material storage area, tool operation area and safety isolation area.
[0054] When post-processing the initial segmentation results, you can use morphological closing and opening operations. Morphological closing first dilates the initial segmentation results, followed by an erosion operation. Dilation expands the region outward, filling any small holes within it. For example, a material storage area might have small internal holes, which would be filled after dilation. Erosion shrinks the region inward, restoring its general shape. The combination of these two operations effectively fills holes within the region, making it more complete.
[0055] The morphological opening operation first performs an erosion operation, followed by a dilation operation. The erosion operation removes isolated noise points outside the region, narrowing the area. For example, there may be isolated noise points around the tool operation area; the erosion operation removes these noise points. The dilation operation restores the area to its original size. The morphological opening operation eliminates isolated noise points outside the region, making the segmentation result more accurate. After these two post-processing steps, the final material storage area, tool operation area, and safety isolation area are generated.
[0056] Step S125: Perform spatial overlap detection on the motion amplitude parameter in the dynamic behavior feature and the material storage area boundary closure in the specification constraint feature to generate a material area cross-border contact feature. The material area cross-border contact feature indicates whether the motion amplitude parameter causes the basic action point to enter the non-permitted area outside the material storage area boundary.
[0057] After determining the dynamic behavior characteristics of the operator and the standard constraints of the work scenario, we can perform spatial overlap detection on the motion amplitude parameter and the material storage area boundary closure. The motion amplitude parameter reflects the maximum amplitude of the operator's motion, while the material storage area boundary closure indicates the degree of enclosure of the material storage area boundary.
[0058] First, based on the boundary information of the material storage area, a non-permitted area outside the material storage area can be determined. This non-permitted area is defined based on operational specifications and safety requirements and is typically a certain distance outside the material storage area boundary. Then, combining the motion amplitude parameters and the motion trajectory of the basic motion point, it is determined whether the basic motion point is likely to enter the non-permitted area. If the motion amplitude is large and the motion trajectory of the basic motion point is close to the material storage area boundary, it is possible that the basic motion point has entered the non-permitted area.
[0059] The specific detection process is to draw a circular area with the motion point as the center and the motion amplitude as the radius for the motion trajectory of each basic action point (assuming that the motion is a movement on a two-dimensional plane). Then, it is determined whether the circular area intersects with the non-permitted area outside the boundary of the material storage area. If there is an intersection, it means that the motion amplitude parameter may cause the basic action point to enter the non-permitted area. Through the above-mentioned spatial overlap detection, the material area cross-border contact feature is generated. The material area cross-border contact feature is a Boolean value or binary value used to indicate whether the motion amplitude parameter causes the basic action point to enter the non-permitted area outside the boundary of the material storage area. If it enters the non-permitted area, the feature value is true or 1; if it does not enter, the feature value is false or 0.
[0060] Step S126: Perform a time-series correlation analysis on the action frequency parameter in the dynamic behavior feature and the tool operation area area ratio in the standard constraint feature to generate a tool area operation intensity matching feature, which indicates whether the action frequency per unit time matches the tool operation area area ratio.
[0061] A temporal correlation analysis was conducted on the action frequency parameter and the tool operation area ratio. The action frequency parameter reflects the frequency of the operator's actions, while the tool operation area ratio reflects the relative size of the tool operation area in the entire image.
[0062] In this embodiment, the changes in the action frequency parameter and the area ratio of the tool operation area can be counted over a period of time. First, the period is divided into multiple small time windows, and the action frequency parameter and the area ratio of the tool operation area are counted in each time window. For example, with a time window of 1 minute, the operator's action frequency and the area ratio of the tool operation area are counted in each 1-minute period.
[0063] Next, we analyze the relationship between the action frequency parameter and the area ratio of the tool operation area. If the tool operation area ratio is large, it indicates that there is more room for tool operation, and theoretically, the operator can have a higher action frequency. If the tool operation area ratio is small, the action frequency may be relatively low. By comparing the numerical changes of these two parameters within different time windows, we can determine whether the action frequency per unit time matches the area ratio of the tool operation area.
[0064] For example, if the tool operation area percentage suddenly increases within a certain time window without a corresponding increase in the action frequency, this indicates a mismatch between the action frequency and the tool operation area percentage. If the two trends are consistent, this indicates a match between the action frequency and the tool operation area percentage. Based on the above analysis results, a tool area operation intensity matching feature is generated. This tool area operation intensity matching feature can also be a Boolean or binary value, indicating whether the action frequency per unit time matches the tool operation area percentage. If so, the feature value is true or 1; if not, the feature value is false or 0.
[0065] Step S130: Input the dynamic behavior feature and the standard constraint feature into a pre-trained violation discrimination model for compliance matching, and generate a discrimination result including a violation type identifier.
[0066] In the tobacco station operation scenario, after obtaining the dynamic behavioral characteristics of the operators and the standard constraint characteristics of the operation scenario, these characteristics must be input into the pre-trained violation discrimination model for compliance matching. This violation discrimination model is the core component of the entire detection process. Its construction purpose is to accurately determine whether the operator's operation complies with the standard requirements of the tobacco station operation. This model is composed of multiple sub-modules, namely the behavior analysis sub-module, the scene analysis sub-module, the fusion analysis sub-module and the classification output sub-module. Each sub-module works together to gradually process and analyze the input features, and finally outputs the discrimination result including the violation type identification.
[0067] Step S131: inputting the motion amplitude parameter, motion frequency parameter and motion continuity parameter in the dynamic behavior feature into the behavior analysis submodule of the violation discrimination model, and extracting the time dependency information of the motion feature through the long short-term memory network.
[0068] Among the acquired dynamic behavior features, the motion amplitude parameter reflects the maximum range of the operator's movements, the motion frequency parameter reflects the frequency of movements per unit time, and the motion coherence parameter measures the smoothness of the transitions between movements. These parameters are input into the behavior analysis submodule of the violation discrimination model. This submodule uses a long short-term memory network (LSTM) to extract the temporal dependency of the motion features.
[0069] LSTM is a recurrent neural network specifically designed for processing sequential data. It features memory cells and a gating mechanism, effectively capturing long-term dependencies in sequential data. In this embodiment, the LSTM processes motion amplitude, frequency, and continuity parameters in chronological order, frame by frame. At each time step, the LSTM's memory cells are updated based on the current input and the output and memory state from the previous time step.
[0070] Specifically, the LSTM's forget gate determines which information in the memory cell should be forgotten based on the output of the previous time step and the current input. For example, if the action amplitude was large in the previous time step, but the current time step shows a significantly smaller action amplitude, the forget gate may choose to forget the information related to the previous large action amplitude. The input gate, on the other hand, determines which new information should be added to the memory cell based on the current input. For example, if the action frequency suddenly increases in the current time step, the input gate will add information related to this increase in frequency to the memory cell. The output gate determines the output content of the current time step based on the current state of the memory cell and the current input.
[0071] Through this processing, LSTM can extract temporal dependencies of motion features. For example, it can detect whether motion amplitude exhibits periodic changes over time, whether motion frequency exhibits sudden fluctuations, and whether motion coherence varies across time periods. This temporal dependency information is crucial for determining operational compliance, as some illegal operations may be difficult to detect in the short term but may reveal abnormal characteristics when analyzed from a time series perspective.
[0072] Step S132: Input the boundary closure, area proportion and inter-region adjacency relationship in the standard constraint features into the scene analysis submodule of the violation discrimination model, and extract the spatial correlation information of the scene features through the spatial attention network.
[0073] The boundary closure within the regulatory constraint features describes the degree of enclosure within the boundaries of areas such as material storage areas, tool operation areas, and safety isolation zones. The area ratio reflects the proportion of each area within the overall work scene, and the inter-region adjacency relationship reflects the adjacent positional relationships between different areas. These features are input into the scene analysis submodule of the violation discrimination model, which utilizes a spatial attention network to extract spatial correlation information from scene features.
[0074] The Spatial Attention Network performs a multi-scale analysis of input features. First, it performs feature mapping on each feature, converting it into a feature representation suitable for network processing. Then, through the attention mechanism, the network focuses on feature information at different spatial locations.
[0075] When analyzing boundary closure, the attention mechanism assigns different weights to different locations along the boundary. For example, if a region's boundary has a gap, the gap will be given a higher weight because it may affect the safety and compliance of the region. Regarding area proportion, the network considers the proportional relationship between the various regions and whether this proportion complies with the requirements of the operating specifications. If the area proportion of the tool operation zone is too small, it may affect work efficiency or pose a safety hazard. In this case, the network will focus on the relevant features of this area.
[0076] When processing inter-region adjacency relationships, the spatial attention network analyzes the proximity of different areas. For example, it determines whether the distance between the material storage area and the safety isolation zone is appropriate, or whether the tool operation area is adjacent to a flammable area. Through this approach, the spatial attention network can extract spatial correlation information about scene features, helping to determine whether the layout of the work scene complies with regulations.
[0077] Step S133: inputting the time dependency information and the spatial correlation information into the fusion analysis submodule of the violation discrimination model, performing fusion processing through the feature splicing layer, and generating a fusion feature vector.
[0078] Step S1331: Acquire the characteristic dimension information of the time dependency information and the characteristic dimension information of the spatial correlation information.
[0079] After inputting the temporal dependency information and spatial correlation information into the fusion analysis submodule, the first step is to obtain their respective feature dimension information. Feature dimension information reflects the length of the feature vector, that is, the number of elements contained in the feature.
[0080] For temporal dependency information, its characteristic dimension may be related to the hidden layer dimension of the LSTM output. For example, the hidden layer of an LSTM may have multiple neurons, and the output of each neuron constitutes an element of temporal dependency information. The number of these elements is the characteristic dimension of temporal dependency information. The characteristic dimension of spatial correlation information depends on the output dimension of the spatial attention network, which is also related to the number and structure of neurons in the network.
[0081] Acquiring feature dimension information is to prepare for the subsequent dimension alignment operation. Only by ensuring that the dimensions of the two pieces of information are consistent can effective feature splicing be performed.
[0082] Step S1332: If the feature dimension of the time dependency information is smaller than the feature dimension of the spatial correlation information, a zero-value feature is added to the end of the time dependency information through a zero-padding operation to make its dimension consistent with the spatial correlation information.
[0083] If it is found through comparison that the feature dimension of the time-dependent information is smaller than the feature dimension of the spatial-correlation information, a zero-padding operation may be used to enable the two to be concatenated in the same dimension.
[0084] Specifically, this involves adding a certain number of zero-valued features to the end of the temporal dependency information. The number of zero-valued features added is equal to the feature dimension of the spatial correlation information minus the feature dimension of the temporal dependency information. For example, if the feature dimension of the temporal dependency information is m and the feature dimension of the spatial correlation information is n (where n > m), then nm zero-valued features need to be added to the end of the temporal dependency information, bringing the dimension of the temporal dependency information to n, the same as the dimension of the spatial correlation information.
[0085] The purpose of this is to ensure that each corresponding position of the two information has a corresponding value during the splicing process, avoiding information loss or splicing errors due to dimensional inconsistency. The zero padding operation does not change the original characteristics of the time-dependent information, but only meets the dimensional alignment requirements.
[0086] Step S1333: If the feature dimension of the time dependency information is greater than the feature dimension of the spatial correlation information, redundant features at the end of the time dependency information are removed by truncation to make its dimension consistent with the spatial correlation information.
[0087] When the feature dimension of the temporal dependency information is larger than that of the spatial correlation information, truncation can be performed. This removes a portion of the features at the end of the temporal dependency information. The number of features removed is equal to the feature dimension of the temporal dependency information minus the feature dimension of the spatial correlation information.
[0088] For example, if the feature dimension of the time-dependent information is m and the feature dimension of the spatial correlation information is n (m>n), the mn features at the end of the time-dependent information will be removed, and the dimension of the time-dependent information will be adjusted to n to make it consistent with the dimension of the spatial correlation information.
[0089] By performing truncation, we can remove some redundant information that may not be helpful for subsequent classification tasks, while ensuring that the two information dimensions match, facilitating effective feature concatenation. When performing truncation, we need to ensure that the removed features do not lose important information. We usually sort the features by their importance, prioritizing the important ones.
[0090] Step S1334: If the characteristic dimension of the time-dependent information is equal to the characteristic dimension of the spatial-correlation information, directly retain the original dimensions of both.
[0091] If the comparison finds that the characteristic dimension of the time-dependent information is equal to the characteristic dimension of the spatial correlation information, it means that the dimensions of the two information have matched, and no additional processing is required, and their original dimensions can be directly retained.
[0092] Doing so can avoid unnecessary information modification and ensure the integrity and accuracy of the information. In the above case, the two information can be directly spliced without problems caused by inconsistent dimensions.
[0093] Step S1335: splicing the dimensionally aligned temporal dependency information and spatial correlation information along the feature channel dimension to generate an initial fusion feature.
[0094] After dimensional alignment, the temporal dependency information and spatial correlation information are concatenated along the feature channel dimension. The feature channel dimension is a specific dimension of the feature vector that is used to distinguish different types of features.
[0095] For example, time dependency information can be regarded as a set of vectors about time features, and spatial correlation information can be regarded as a set of vectors about spatial features. Splicing them in the feature channel dimension is equivalent to combining time features and spatial features together.
[0096] The specific concatenation process involves sequentially arranging each element of the temporal dependency information with the corresponding element of the spatial correlation information. Assuming the dimension of the temporal dependency information is n and the dimension of the spatial correlation information is also n, the dimension of the initial fused feature after concatenation becomes 2n. This concatenation method integrates both temporal and spatial information to form a new feature representation.
[0097] Step S1336: Normalize the initial fusion features, map the feature values to the range of 0 to 1 through the layer normalization algorithm, perform linear transformation on the normalized initial fusion features, map the feature values to the input dimensions required by the model classification output submodule through the fully connected layer, and generate a fusion feature vector.
[0098] Normalize the initial fusion features using the layer normalization algorithm. The layer normalization algorithm processes each feature vector of the initial fusion features and calculates its mean and variance.
[0099] Specifically, for each eigenvector, we first calculate the sum of all its elements and divide it by the number of elements to obtain the mean. Next, we calculate the sum of the squares of the differences between each element and the mean and divide it by the number of elements to obtain the variance. By subtracting the mean from each element and dividing it by the standard deviation (which is the square root of the variance), we can map the eigenvalues to the range of 0 to 1.
[0100] The purpose of this is to eliminate scale differences between different features and make feature values comparable. For example, if one feature has a large range of values and another has a smaller range, normalization can unify their ranges, helping to improve model training efficiency and stability.
[0101] The normalized initial fused features are then linearly transformed through a fully connected layer. A fully connected layer consists of multiple neurons, each connected to all input feature elements. Through a linear combination of a weight matrix and a bias term, the normalized feature values are mapped to the input dimensions required by the model's classification output submodule.
[0102] Assuming the model's classification output submodule requires an input feature vector of dimension k, the fully connected layer converts the normalized initial fused features into a fused feature vector of dimension k by setting appropriate weight matrices and bias terms. The feature representation of this fused feature vector meets the input requirements of the model's classification output submodule and can be better classified by the classification output submodule.
[0103] Step S134: inputting the fused feature vector into the classification output submodule of the violation discrimination model, and calculating the confidence value of each violation type through a fully connected network.
[0104] The fused feature vector is input into the classification output submodule of the violation discrimination model. This submodule uses a fully connected network to calculate the confidence value of each violation type.
[0105] A fully connected network consists of multiple fully connected layers, each containing multiple neurons. The input fused feature vector first enters the first fully connected layer, where each neuron performs a linear combination of the input feature vector, multiplying it by the corresponding weight and adding a bias term.
[0106] For example, assuming that the input fused feature vector has m elements and the first fully connected layer has n neurons, then each neuron will have m weight values, which are multiplied by the m elements of the fused feature vector respectively, and then the products are added and the bias term is added to obtain the output value of the neuron.
[0107] After being processed by the first fully connected layer, the output result will be used as the input of the next fully connected layer, and so on. After being processed by multiple fully connected layers, a vector with the same dimension as the number of violation types is finally obtained.
[0108] During this process, each fully connected layer introduces a nonlinear activation function, such as the ReLU function, to enhance the network's expressive power. The ReLU function sets input values less than 0 to 0 and leaves input values greater than 0 unchanged. This introduces nonlinear factors, enabling the network to learn more complex patterns.
[0109] Each element in the resulting vector represents an unnormalized confidence score for the corresponding violation type. For example, if there are three violation types, the final output vector will have a dimension of 3, with each element corresponding to the confidence score for each violation type. These scores require further processing to obtain the final confidence score.
[0110] Step S135: Filter out violation types whose confidence values exceed a preset threshold, and generate a discrimination result including a violation type identifier and a corresponding confidence value; if all confidence values do not exceed the preset threshold, generate a discrimination result including a compliance identifier.
[0111] After obtaining the confidence values for each violation type, these values need to be compared with the preset threshold. The preset threshold is a critical value set based on actual conditions and experience to determine whether the confidence level of a violation type is high enough to be considered a violation.
[0112] If the confidence value of a certain violation type exceeds the preset threshold, it means that the job operation is likely to belong to this violation type. Therefore, the identification and corresponding confidence of the violation type that exceeds the preset threshold can be filtered out to generate a judgment result including the violation type identification and corresponding confidence.
[0113] For example, suppose there are violation types A, B, and C, the preset threshold is 0.5, and the calculated confidence value for violation type A is 0.6, the confidence value for violation type B is 0.3, and the confidence value for violation type C is 0.7. Violation types A and C will be screened out, and the generated discrimination result will record the identifiers of violation types A and C and their corresponding confidence values of 0.6 and 0.7.
[0114] If the confidence values for all violation types do not exceed the preset threshold, this indicates that the operation complies with regulatory requirements based on the current model's judgment, with no obvious violations detected. In this case, a judgment result with a compliance indicator can be generated, informing supervisors that the operation is compliant. This approach allows for accurate determination of the presence of violations and the specific types of violations.
[0115] Step S140: locating the time starting point, time ending point and spatial effect area of the illegal operation in the continuous image sequence according to the determination result.
[0116] After obtaining the identification result including the violation type identifier, it is necessary to locate the time starting point, time ending point and spatial effect area of the violation operation in the continuous image sequence according to the identification result.
[0117] Step S141: extracting the image frame timestamp containing the violation type identifier in the discrimination result and marking it as the violation-related time point.
[0118] Extract the image frame timestamps containing the violation type identifier from the discrimination results. These timestamps record the approximate time when the violation occurred. Mark these timestamps as the violation-related time points, which are important for subsequently locating the time range of the violation.
[0119] Step S142: traverse the continuous image sequence forward and find the first image frame timestamp with the violation type identifier as the time starting point.
[0120] After obtaining the violation-associated time points, we can traverse the continuous image sequence forward. Starting from the last violation-associated time point, we examine the preceding image frames one by one to find the timestamp of the first image frame with the violation type identifier. This timestamp is the starting point of the violation operation, marking the moment when the violation operation began.
[0121] Step S143: traverse the continuous image sequence backwards to find the timestamp of the last image frame where the violation type identifier appears as the time end point.
[0122] Similarly, we can traverse the continuous image sequence backwards. Starting from the first violation-related time point, we check the subsequent image frames in sequence to find the timestamp of the last image frame with the violation type identifier. This image frame timestamp is the time end point of the violation operation, marking the moment when the violation operation ended.
[0123] Step S144: extracting all image frames between the time start point and the time end point, and obtaining a coordinate set of the operator's basic action points in each image frame.
[0124] After determining the start and end time points, all image frames between these two time points can be extracted. For these image frames, the coordinate set of the operator's basic action points in each image frame can be obtained. The coordinate set records the position information of the operator's key parts during the illegal operation.
[0125] Step S145: performing spatial density analysis on the coordinate set, counting the number of times each image coordinate point is covered by the basic action point, and generating a trajectory coverage density map.
[0126] The spatial density analysis of the coordinate set is performed to understand the distribution of basic action points of the operators during the illegal operation.
[0127] Step S1451: Initialize a two-dimensional array with the same size as a single-frame image as a density recording matrix, where each element of the density recording matrix corresponds to a coordinate point in the image.
[0128] In this embodiment, a two-dimensional array of the same size as a single frame image can be initialized as a density record matrix. Each element of the density record matrix corresponds to a coordinate point in the image. Initially, all elements in the matrix are set to 0, indicating that each coordinate point is covered by the basic action point 0 times.
[0129] Step S1452: traverse each basic action point coordinate in the coordinate set to obtain its row and column index in the density record matrix.
[0130] For each basic action point coordinate in the coordinate set, its row and column index in the density record matrix can be found. This row and column index represents the position of the coordinate point in the matrix. Using this row and column index, the value of the corresponding element in the density record matrix can be accurately accessed and updated.
[0131] Step S1453: perform an addition operation on the density record matrix element corresponding to the row and column index to update the number of times the coordinate point is covered.
[0132] After obtaining the row and column indices of the basic action point coordinates in the density record matrix, you can increment the matrix element corresponding to that index by 1. Each time a basic action point falls on that coordinate point, the value of that element is incremented by 1. In this way, the number of times that the coordinate point is covered by a basic action point is updated.
[0133] Step S1454: After completing the traversal of all basic action point coordinates, the density record matrix is converted into a grayscale image, in which the pixel grayscale value is positively correlated with the number of coverage times, to generate an initial density map.
[0134] After traversing the coordinates of all basic action points, the density record matrix can be converted into a grayscale image. In this grayscale image, the pixel grayscale value is positively correlated with the number of coverages. That is, the greater the number of coverages, the higher the corresponding pixel grayscale value; the fewer the number of coverages, the lower the pixel grayscale value. Through this conversion, the data in the density record matrix is presented as an image, generating an initial density map.
[0135] Step S1455: performing Gaussian blur processing on the initial density map, and performing histogram equalization processing on the blurred initial density map, and using the processed initial density map as the final trajectory coverage density map.
[0136] The initial density map is Gaussian blurred. This smoothes the pixels in the image, reducing noise and detail, making the image more blurry and continuous. The blurred initial density map is then subjected to histogram equalization. Histogram equalization adjusts the image's grayscale distribution, enhancing contrast and highlighting important information. After these two processing steps, the processed initial density map is used as the final track coverage density map. This track coverage density map visually demonstrates the distribution density of the operator's basic action points during the illegal operation.
[0137] Step S146: performing threshold binarization processing on the trajectory coverage density map, setting pixel points with coverage times lower than a preset density threshold as background, and setting pixel points with coverage times higher than or equal to the preset density threshold as foreground, to generate a binary density map.
[0138] The trajectory coverage density map is subjected to threshold binarization processing, and a preset density threshold can be set. For each pixel in the trajectory coverage density map, its coverage times can be compared with the preset density threshold. If the coverage times are lower than the preset density threshold, it means that the pixel is covered less times by the basic action point, and it is likely to be the background area. The pixel is set as the background, and its pixel value is set to 0. If the coverage times are higher than or equal to the preset density threshold, it means that the pixel is covered more times by the basic action point, and it is likely to be the main area of illegal operation. The pixel is set as the foreground, and its pixel value is set to 1. Through the above-mentioned threshold binarization processing, a binary density map is generated. The binarized density map can clearly demarcate the main area of illegal operation and the background area.
[0139] Step S147: performing connected region analysis on the binary density map to identify the largest connected foreground region as the core region of the spatial action area.
[0140] Connected region analysis is performed on the binary density map to identify all connected foreground regions within the map. Connected foreground regions are defined as areas consisting of adjacent foreground pixels. By analyzing the size and shape of these connected regions, the largest connected foreground region is identified. This largest connected foreground region is likely the core area of the illegal operation and is considered the core area of the spatial action region, reflecting the primary activity area of the operator during the illegal operation.
[0141] Step S148: performing convex hull calculation on all foreground areas in the binary density map to generate a minimum convex polygon containing all foreground areas as an expansion area of the spatial action area.
[0142] Calculate the convex hull of all foreground regions in the binary density map. The convex hull is the smallest convex polygon that encompasses all foreground regions. By calculating the convex hull, a minimum region that encompasses all foreground regions is obtained. This minimum region serves as an extension of the spatial action area, more comprehensively covering the operator's movements during the illegal operation, including some marginal and scattered areas.
[0143] Step S149: combining the time start point, time end point, core area and extended area to form location information of the illegal operation.
[0144] Combining the time start point, time end point, core area, and extended area together forms the location information of the illegal operation. This location information can accurately indicate the time range of the illegal operation and the spatial area of action in the image.
[0145] Step S150: Generate a warning instruction containing violation location information based on the time start point, time end point and spatial action area, and send the warning instruction to the tobacco station supervision terminal to trigger the violation handling process.
[0146] After obtaining the location information of the illegal operation, a warning instruction containing the illegal location information can be generated based on the location information and sent to the tobacco station supervision terminal.
[0147] Step S151: Convert the time start point and the time end point into a time interval character string in a standard time format.
[0148] In this embodiment, the time start point and the time end point can be converted into a time interval string in a standard time format. The standard time format can be a common date and time representation, such as year-month-day hour:minute:second. The time start point and the time end point are converted according to this format and combined into a time interval string, such as "20XX-XX-XXXX:XX:XX-20XX-XX-XXXX:XX:XX". This time interval string can clearly represent the time range in which the illegal operation occurred.
[0149] Step S152: Convert the pixel coordinate point sets of the core area and the extended area of the spatial action area into the coordinate point sets of the actual geographic coordinate system of the tobacco station, and map the image pixel coordinates to the actual geographic coordinates through the pre-calibrated camera intrinsic and extrinsic parameter matrices.
[0150] In this embodiment, the pixel coordinate point sets of the core area and the extended area of the spatial action area can be converted into coordinate point sets of the actual geographic coordinate system of the tobacco station. First, the camera calibration parameters of the tobacco station operation scene can be obtained. These camera calibration parameters include information such as the camera's focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient.
[0151] Distortion correction is performed on the pixel coordinate points in the core and extended areas. The radial and tangential distortion coefficients in the camera calibration parameters are used to correct the pixel coordinates to eliminate the effects of image distortion. The corrected pixel coordinate points are then converted to the camera coordinate system. The pixel coordinates are then converted to 3D coordinates in the camera coordinate system using the camera intrinsic parameter matrix.
[0152] Finally, the coordinate point set in the camera coordinate system is converted to the coordinate point set of the actual geographic coordinate system of the tobacco station. The three-dimensional coordinates of the camera coordinate system are converted to the three-dimensional coordinates of the geographic coordinate system using the camera extrinsic parameter matrix. The three-dimensional coordinates of the geographic coordinate system are projected onto the two-dimensional plane coordinate system of the tobacco station operation scene to generate the final geographic coordinate point set. This geographic coordinate point set accurately represents the location of the illegal operation in the actual geographic space of the tobacco station.
[0153] Step S153: According to the violation type identifier in the discrimination result, a predefined violation type and disposal measure comparison table is queried to obtain the corresponding disposal measure code and supervision priority level.
[0154] Based on the violation type identifier in the identification result, a predefined violation type and action comparison table can be queried. This table records the action code and regulatory priority level corresponding to each violation type. By querying this table, the action code and regulatory priority level corresponding to the violation type identifier can be obtained. The action code indicates the specific action to be taken for the violation type, and the regulatory priority level indicates the severity of the violation type and the urgency of the regulatory oversight required.
[0155] Step S154: The time interval character string, the coordinate point set of the geographic coordinate system, the disposal measure code and the supervision priority level are subjected to data integration processing to generate a comprehensive data set including time positioning, space positioning, disposal measures and supervision priority.
[0156] By integrating the time interval string, geographic coordinate point set, action code, and regulatory priority level, this information can be combined to form a comprehensive dataset that includes time location, spatial location, action, and regulatory priority. This comprehensive dataset can fully describe the relevant information of the illegal operation.
[0157] Step S155: compress the comprehensive data set, digitally sign the compressed comprehensive data set, encode the signed comprehensive data set according to the communication protocol of the tobacco station supervision terminal, and generate a warning instruction containing violation location information.
[0158] Data compression is performed on the integrated dataset, reducing its size and alleviating the burden of data transmission. The compressed dataset can then be digitally signed. Digital signatures encrypt the dataset using an encryption algorithm, ensuring data integrity and authenticity and preventing data tampering during transmission.
[0159] Finally, the signed, comprehensive dataset is encoded according to the communication protocol of the tobacco station's supervisory terminal. This protocol specifies the data transmission format and rules. This encoded dataset generates a warning instruction containing the location of the violation. This warning instruction is sent to the tobacco station's supervisory terminal, triggering the violation handling process, allowing supervisors to promptly identify the violation and take appropriate action.
[0160] Figure 2 A schematic diagram illustrates exemplary hardware and software components of an image recognition-based system 100 for detecting illegal cigarette operation at a cigarette stand, according to some embodiments of the present application. For example, a processor 120 may be used in the image recognition-based system 100 for detecting illegal cigarette operation at a cigarette stand, and may be used to perform the functions described herein.
[0161] The image recognition-based cigarette station illegal operation detection system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the image recognition-based cigarette station illegal operation detection method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0162] For example, the image recognition-based cigarette station illegal operation detection system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the image recognition-based cigarette station illegal operation detection system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The image recognition-based cigarette station illegal operation detection system 100 also includes an I / O interface 150 between the computer and other input and output devices.
[0163] For ease of explanation, only one processor is described in the image recognition-based cigarette station illegal operation detection system 100. However, it should be noted that the image recognition-based cigarette station illegal operation detection system 100 in this application can also include multiple processors, so the steps performed by one processor described in this application can also be performed jointly or individually by multiple processors. For example, if the processor of the image recognition-based cigarette station illegal operation detection system 100 executes step A and step B, it should be understood that step A and step B can also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0164] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned method for detecting illegal operations of a cigarette station based on image recognition is implemented.
[0165] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A method for detecting illegal operations at a cigarette station based on image recognition, characterized in that: The method comprises: Acquire a continuous image sequence of the smoke station operation scene, wherein the continuous image sequence is composed of multiple frames of images acquired in time sequence, and each frame of the image contains visual information of the operator's body movements, material contact status, and tool use position; Performing operation feature analysis on the continuous image sequence to obtain dynamic behavior features of the operator and standard constraint features of the operation scene; Inputting the dynamic behavior features and the regulatory constraint features into a pre-trained violation discrimination model for compliance matching, and generating a discrimination result including a violation type identifier; Locating the time starting point, time ending point and spatial effect area of the illegal operation in the continuous image sequence according to the discrimination result; Based on the time starting point, time ending point and spatial action area, a warning instruction containing violation location information is generated, and the warning instruction is sent to the tobacco station supervision terminal to trigger the violation handling process.
2. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 1 is characterized in that: The operation feature analysis processing of the continuous image sequence to obtain the dynamic behavior features of the operator and the standard constraint features of the operation scene includes: Performing personnel key point detection processing on each frame of the continuous image sequence, and extracting the pixel coordinates of the operator's shoulder, elbow, wrist and knee as basic action points through a human posture estimation algorithm; Performing key point tracking processing on adjacent image frames in the continuous image sequence, recording the coordinate changes of each basic action point in the continuous frames, and generating an action point motion trajectory sequence; Based on the motion trajectory sequence of the action points, the action amplitude parameter of the operator is calculated, where the action amplitude parameter is the maximum displacement of the basic action point in consecutive frames; the action frequency parameter is calculated, where the action frequency parameter is the number of times the basic action point is displaced per unit time; and the action continuity parameter is calculated, where the action continuity parameter is the minimum angle between changes in the direction of consecutive displacements; and the action amplitude parameter, the action frequency parameter, and the action continuity parameter are combined to form a dynamic behavior feature. Performing scene semantic segmentation processing on a single-frame image in the continuous image sequence, dividing the image into a material storage area, a tool operation area, and a safety isolation area through a semantic segmentation model; extracting the boundary closure of the material storage area, wherein the boundary closure is the ratio of the distance between the beginning and end points of the material storage area boundary to the total length of the boundary; extracting the area ratio of the tool operation area, wherein the area ratio is the ratio of the number of pixels in the tool operation area to the total number of pixels in the image; extracting the inter-regional adjacency relationship between the safety isolation area and the tool operation area, wherein the inter-regional adjacency relationship is the shared boundary length between the safety isolation area and the tool operation area; and combining the boundary closure, area ratio, and inter-regional adjacency relationship to form a standard constraint feature; Performing spatial overlap detection on the motion amplitude parameter in the dynamic behavior feature and the material storage area boundary closure degree in the specification constraint feature to generate a material area cross-border contact feature, wherein the material area cross-border contact feature indicates whether the motion amplitude parameter causes the basic motion point to enter a non-permitted area outside the material storage area boundary; A time series correlation analysis is performed on the action frequency parameter in the dynamic behavior feature and the tool operation area area ratio in the standard constraint feature to generate a tool area operation intensity matching feature. The tool area operation intensity matching feature indicates whether the action frequency per unit time matches the tool operation area area ratio.
3. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 2 is characterized in that: The step of performing key point tracking processing on adjacent image frames in the continuous image sequence, recording the coordinate changes of each basic action point in the continuous frames, and generating an action point motion trajectory sequence includes: Select the basic motion point coordinates of the operator's right wrist in the previous image frame as the historical coordinates, and select the basic motion point coordinates of the right wrist in the current image frame as the current coordinates; Calculating the horizontal and vertical displacement differences between the historical coordinates and the current coordinates in the image to obtain a single-frame displacement vector of the action point; Count the displacement vector sequences of the right wrist, left elbow, and knee basic action points in consecutive image frames, and connect them in chronological order to generate a sequence of action point motion trajectories; Smoothing the motion trajectory sequence of the action points, eliminating abnormal fluctuations of the displacement vector caused by image acquisition errors by using a sliding average algorithm, and generating a smoothed motion trajectory sequence; Performing directional consistency analysis on the smoothed motion trajectory sequence, calculating the changing trend of the continuous displacement vector direction, and generating a trajectory directional stability parameter; Performing velocity continuity analysis on the smoothed motion trajectory sequence, calculating the change gradient of the continuous displacement vector modulus, and generating a trajectory velocity stability parameter; The smoothed motion trajectory sequence, the trajectory direction stability parameter, and the trajectory speed stability parameter are combined to form a complete action point motion trajectory sequence.
4. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 2 is characterized in that: The performing scene semantic segmentation processing on the single-frame image in the continuous image sequence, and dividing the image into a material storage area, a tool operation area, and a safety isolation area by using a semantic segmentation model, includes: Preprocessing the single-frame image, reducing image noise using a Gaussian filtering algorithm, and enhancing edge contrast of materials, tools, and safety isolation areas to obtain a preprocessed image; The preprocessed image is input into the feature extraction layer of the semantic segmentation model, and the local texture features and global context features of the image are extracted through the convolutional neural network; Inputting the local texture features and global context features into the upsampling layer of the semantic segmentation model, restoring the feature map size to be consistent with the input image through a deconvolution operation, and generating a segmentation probability map containing pixel-level classification information; Threshold processing is performed on the segmentation probability map, and pixels with probability values greater than a preset threshold are classified into corresponding areas to generate initial segmentation results for the material storage area, tool operation area, and safety isolation area; The initial segmentation results are post-processed to fill the holes in the area through morphological closing operations and to eliminate isolated noise points outside the area through morphological opening operations, thereby generating the final material storage area, tool operation area and safety isolation area.
5. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 1 is characterized in that: The step of inputting the dynamic behavior features and the regulatory constraint features into a pre-trained violation discrimination model for compliance matching to generate a discrimination result including a violation type identifier includes: Inputting the motion amplitude parameter, motion frequency parameter, and motion coherence parameter in the dynamic behavior feature into the behavior analysis submodule of the violation discrimination model, and extracting the time dependency information of the motion feature through a long short-term memory network; Input the boundary closure, area proportion and inter-region adjacency relationship in the regulatory constraint features into the scene analysis submodule of the violation discrimination model, and extract the spatial correlation information of the scene features through the spatial attention network; Inputting the time dependency information and the spatial correlation information into the fusion analysis submodule of the violation discrimination model, performing fusion processing through the feature splicing layer, and generating a fusion feature vector; Inputting the fused feature vector into the classification output submodule of the violation discrimination model, and calculating the confidence value of each violation type through a fully connected network; Filter out violation types whose confidence values exceed the preset threshold and generate a judgment result including the violation type identification and the corresponding confidence; if all confidence values do not exceed the preset threshold, generate a judgment result including the compliance identification.
6. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 5 is characterized in that: The step of inputting the time dependency information and the spatial correlation information into the fusion analysis submodule of the violation discrimination model, performing fusion processing through a feature splicing layer, and generating a fusion feature vector includes: Acquiring characteristic dimension information of the time-dependent information and characteristic dimension information of the spatial correlation information; If the feature dimension of the time dependency information is smaller than the feature dimension of the spatial correlation information, a zero-value feature is added to the end of the time dependency information through a zero-padding operation to make its dimension consistent with the spatial correlation information; If the feature dimension of the time dependency information is greater than the feature dimension of the spatial correlation information, removing redundant features at the end of the time dependency information by truncation so that its dimension is consistent with the spatial correlation information; If the characteristic dimension of the time-dependent information is equal to the characteristic dimension of the spatial correlation information, directly retain the original dimensions of both; The time dependency information and spatial correlation information after dimension alignment are spliced along the feature channel dimension to generate the initial fusion feature; The initial fusion features are normalized, the feature values are mapped to the range of 0 to 1 through the layer normalization algorithm, the normalized initial fusion features are linearly transformed, and the feature values are mapped to the input dimensions required by the model classification output submodule through the fully connected layer to generate a fusion feature vector.
7. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 1 is characterized in that: The step of locating the time starting point, the time ending point, and the spatial effect area of the illegal operation in the continuous image sequence according to the discrimination result includes: Extracting the timestamp of the image frame containing the violation type identifier in the discrimination result and marking it as the violation-related time point; Traverse the continuous image sequence forward and find the first image frame timestamp with the violation type identifier as the time starting point; Traverse the continuous image sequence backward and find the timestamp of the last image frame with the violation type identifier as the time end point; Extracting all image frames between the time start point and the time end point, and obtaining a coordinate set of the basic action points of the operator in each image frame; Performing spatial density analysis on the coordinate set, counting the number of times each image coordinate point is covered by the basic action point, and generating a trajectory coverage density map; Performing threshold binarization processing on the trajectory coverage density map, setting pixel points with coverage times lower than a preset density threshold as background, and setting pixel points with coverage times higher than or equal to the preset density threshold as foreground, to generate a binary density map; Performing connected region analysis on the binary density map to identify the largest connected foreground region as the core region of the spatial action area; Performing convex hull calculation on all foreground areas in the binary density map to generate a minimum convex polygon containing all foreground areas as an extension area of the spatial action area; The time starting point, time ending point, core area and extended area are combined to form location information of the illegal operation.
8. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 7 is characterized in that: The performing of spatial density analysis on the coordinate set, counting the number of times each image coordinate point is covered by a basic action point, and generating a trajectory coverage density map includes: Initialize a two-dimensional array with the same size as the single-frame image as a density recording matrix, where each element of the density recording matrix corresponds to a coordinate point in the image; Traversing each basic action point coordinate in the coordinate set to obtain its row and column index in the density record matrix; Add 1 to the density record matrix element corresponding to the row and column index to update the coverage times of the coordinate point; After completing the traversal of all basic action point coordinates, the density record matrix is converted into a grayscale image, in which the pixel grayscale value is positively correlated with the number of coverages, to generate an initial density map; Gaussian blur processing is performed on the initial density map, and histogram equalization processing is performed on the blurred initial density map, and the processed initial density map is used as the final trajectory coverage density map.
9. The method for detecting illegal operation of a cigarette station based on image recognition according to claim 1, characterized in that: The generating of a warning instruction including illegal location information based on the time starting point, the time ending point, and the spatial action area includes: Convert the time start point and the time end point into a time interval character string in a standard time format; The pixel coordinate point sets of the core area and the extended area of the spatial action area are converted into the coordinate point sets of the actual geographic coordinate system of the tobacco station, and the image pixel coordinates are mapped to the actual geographic coordinates through the pre-calibrated camera intrinsic and extrinsic parameter matrices; According to the violation type identifier in the judgment result, query the predefined violation type and disposal measure comparison table to obtain the corresponding disposal measure code and supervision priority level; Performing data integration processing on the time interval character string, the coordinate point set of the geographic coordinate system, the disposal measure code, and the regulatory priority level to generate a comprehensive data set including time positioning, spatial positioning, disposal measures, and regulatory priority; The comprehensive data set is compressed, the compressed comprehensive data set is digitally signed, and the signed comprehensive data set is encoded according to the communication protocol of the tobacco station supervision terminal to generate a warning instruction containing violation location information.
10. A smoke station illegal operation detection system based on image recognition, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the image recognition-based smoking station illegal operation detection method described in any one of claims 1 to 9.
Citation Information
Cited By
Method, device, system and equipment for detecting violation state of individual protection equipment
CN121033770A
Equipment operation compliance detection method and related equipment
CN122369125A
A method for testing equipment operation compliance and related equipment
CN122369125B