Petroleum and petrochemical dangerous area violation operation early warning method based on deep learning
Through deep learning technology, the target image sequence of the petrochemical region is obtained, the target index and timing trajectory are established, the multi-objective interaction feature map is constructed, and the weight is dynamically adjusted, and the dynamic inference graph network is used to identify behaviors, which solves the shortcomings of the existing monitoring system in multi-objective interaction and environmental adaptability, and realizes efficient early warning of violation operations.
Patent Information
- Application Number
- CN202510857520.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing petroleum and petrochemical hazardous area monitoring system cannot effectively analyze multi-target interaction behavior and fail to fully consider the impact of environmental parameters, resulting in insufficient early warning accuracy, and traditional methods cannot model the dynamic correlation between targets, and the accuracy of identifying complex violations is low.
Through deep learning-based methods, the target image sequence is obtained, the target index and timing trajectory are established, the multi-objective interaction feature map is constructed, dynamic weight adjustment is performed in combination with environmental parameters, and behavior recognition is used to capture the complex spatial and temporal relationships between target objects.
Real-time monitoring of the behavior and interaction mode of target objects in the petrochemical area is achieved, the accuracy and adaptability of early warnings are improved, false alarms are reduced, and the level of safety management is improved.
Smart Images

Figure CN120388331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of petroleum and petrochemical industries, and in particular to a method for warning against illegal operations in dangerous areas of petroleum and petrochemical industries based on deep learning. Background Art
[0002] As an important pillar industry of the national economy, the petroleum and petrochemical industries have a large number of flammable and explosive dangerous goods during the production process, and safety production is particularly important. In recent years, with the development of artificial intelligence technology, safety monitoring technology based on computer vision has been widely applied to the safety management of petrochemical areas. Traditional safety monitoring in petrochemical areas mainly relies on manual inspections and passive monitoring by fixed cameras, which has certain limitations in terms of efficiency and accuracy. With the rapid development of deep learning technology, intelligent monitoring systems based on video analysis have gradually become an important means for safety management in petrochemical areas.
[0003] Currently, there are mainly the following problems in the monitoring of illegal operations in dangerous areas of petroleum and petrochemical industries: First, most of the existing technologies focus on the behavior recognition of a single target, lacking effective analysis of multi-target interaction behaviors, and unable to accurately capture the complex interaction patterns between personnel and equipment, and between personnel and personnel, resulting in incomplete recognition of potential dangerous behaviors; Second, the existing monitoring systems have not fully considered the influence of special parameters in the petrochemical environment, such as environmental factors such as temperature, humidity, and gas concentration on safety risks, making the early warning judgment lack environmental adaptability; Third, traditional behavior recognition algorithms mostly use static feature extraction methods, unable to effectively model the dynamic association relationships evolving over time between targets, and having low recognition accuracy for some complex illegal operations that require long-term sequential observation.
[0004] To solve the above problems, a method for warning against illegal operations in dangerous areas of petroleum and petrochemical industries that can comprehensively consider multi-target interaction, the influence of environmental parameters, and has dynamic reasoning ability is needed to improve the intelligent level and early warning accuracy of safety management in petrochemical areas. Summary of the Invention
[0005] The embodiments of the present invention provide a method for warning against illegal operations in dangerous areas of petroleum and petrochemical industries based on deep learning, which can solve the problems in the existing technology.
[0006] In the first aspect of the embodiments of the present invention, a method for warning against illegal operations in dangerous areas of petroleum and petrochemical industries based on deep learning is provided, including:
[0007] Obtain a target image sequence corresponding to the video data of the petrochemical area, perform target detection and tracking on the target image sequence, obtain the position coordinates of the target objects in the target image sequence, and establish a target index;
[0008] Based on the target index, establish a temporal trajectory for each target object in the target image sequence; construct a multi-target interaction feature map according to the temporal trajectory, and the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects;
[0009] Combine the environmental parameters of the petrochemical area to dynamically adjust the weights of the multi-target interaction feature map to obtain an interaction feature map for environmental perception; calculate the safety situation score of the petrochemical area based on the interaction feature map for environmental perception, and the safety situation score represents the overall safety state of the petrochemical area at the current moment;
[0010] Construct a dynamic inference graph network for behavior recognition, input the interaction feature map for environmental perception into the dynamic inference graph network, extract node features through graph convolution operations, and update the dynamic association relationship between nodes based on the message passing mechanism to obtain the recognition result of the target behavior.
[0011] Perform target detection and tracking on the target image sequence, obtain the position coordinates of the target objects in the target image sequence, and establish a target index including:
[0012] Use a target detection model to detect the target image sequence, obtain the target objects in the target image sequence, and extract the position coordinates, size information, and category information of the target objects;
[0013] Assign a globally unique identifier to each target object, and establish a target index table. The target index table records the identifier, position coordinates, size information, category information, and timestamp of the target object. By comparing the visual features and motion features of the target objects between adjacent frames, establish the temporal association of the target objects, and associate the target objects with the same visual features and continuous motion trajectories to the same identifier;
[0014] Based on the position coordinates recorded in the target index table, calculate the motion parameters of the target object.
[0015] Based on the target index, establish a temporal trajectory for each target object in the target image sequence; construct a multi-target interaction feature map according to the temporal trajectory, and the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects, including:
[0016] According to the identification code of the target object, record the position coordinates of the target object in chronological order to obtain the temporal coordinate sequence of the target object; calculate the displacement and time difference between adjacent coordinate points in the temporal coordinate sequence to obtain the motion speed and motion direction of the target object at different times; form the temporal trajectory of the target object by combining the position coordinates, the motion speed, and the motion direction;
[0017] Calculate the relative distance, relative speed, and relative direction between target objects to construct a multi-target interaction feature map; convert the relative distance, the relative speed, and the relative direction into interaction intensity scores according to preset distance weights, speed weights, and direction weights; when the interaction intensity score is greater than a first preset threshold, divide the corresponding target objects into the same interaction group; calculate the motion trajectory coincidence degree of the target objects within the interaction group, and when the motion trajectory coincidence degree is greater than a second preset threshold, determine this interaction group as a collaborative target group.
[0018] The method further includes:
[0019] Calculate the geometric center coordinates and motion direction of the group; calculate the average distance between all target objects within the group and the geometric center to obtain the spatial distribution radius of the group; identify the dominant object within the group, where the dominant object is the target object closest to the geometric center of the group and with the smallest deviation from the motion direction of the group; calculate the trajectory similarity between other target objects and the dominant object, and the trajectory similarity is equal to the reciprocal of the difference in motion direction between the target object and the dominant object; calculate the stability coefficient of the group according to the trajectory similarity;
[0020] When the stability coefficient of the group is greater than a third preset threshold, mark this group as a stable collaborative group; when the spatial distribution radius of the group is greater than a fourth preset threshold or the trajectory similarity of the target object is less than a fifth preset threshold, trigger a group dissolution warning.
[0021] Perform dynamic weight adjustment on the multi-target interaction feature map in combination with the environmental parameters of the petrochemical area to obtain an interaction feature map of environmental perception; calculating the safety situation score of the petrochemical area based on the interaction feature map of environmental perception includes:
[0022] Obtain the environmental parameters of the petrochemical area, where the environmental parameters include the temperature values of multiple temperature monitoring points and the gas concentration values of multiple gas monitoring points;
[0023] Construct a temperature field influence factor based on the temperature values of the temperature monitoring points, where the temperature field influence factor is obtained by calculating the ratio of the distance from the target to the temperature monitoring point to the temperature field attenuation coefficient and combining the normalized temperature value of the temperature monitoring point; construct a gas concentration influence factor based on the gas concentration values of the gas monitoring points, where the gas concentration influence factor is obtained by calculating the ratio of the height difference from the target to the gas monitoring point to the gas diffusion coefficient and combining the normalized concentration value of the gas monitoring point;
[0024] Obtain a multi-object interaction feature map, where the multi-object interaction feature map includes distance features and speed features between objects; adjust the distance features based on the temperature field influence factor and the gas concentration influence factor to obtain distance features for environmental perception; adjust the speed features based on the spatial gradient of the temperature field influence factor and the gas concentration influence factor to obtain speed features for environmental perception;
[0025] Fuse the distance features for environmental perception and the speed features for environmental perception to establish an interaction feature map for environmental perception; calculate the aggregation degree of the objects based on the interaction feature map for environmental perception, and the aggregation degree of the objects is obtained by measuring the similarity between objects;
[0026] Divide the petrochemical area into multiple sub-areas, and calculate the regional risk degree based on the influence intensity of the hazard sources in each sub-area; combine the aggregation degree of the objects and the regional risk degree, and obtain the safety situation score through non-linear mapping.
[0027] Construct a dynamic inference graph network for behavior recognition, input the interaction feature map for environmental perception into the dynamic inference graph network, extract node features through graph convolution operations, and update the dynamic association relationship between nodes based on the message passing mechanism to obtain the recognition results of object behaviors, including:
[0028] Obtain the position features, speed features and environmental features of the objects, construct the node feature vectors from the position features, the speed features and the environmental features, and construct the edge feature vectors based on the association relationship between the node feature vectors; form the initial structure of the dynamic inference graph network with the node feature vectors and the edge feature vectors;
[0029] Construct a temporal connection matrix based on the initial structure of the dynamic inference graph network, and perform spatial feature aggregation on the node feature vectors according to the temporal connection matrix to obtain multi-layer node features;
[0030] Calculate the attention weights for the multi-layer node features, and the attention weights are obtained by splicing the node features and passing through non-linear transformation; perform weighted aggregation on the multi-layer node features based on the attention weights to obtain enhanced node features;
[0031] Generate message features according to the enhanced node features and the edge feature vectors, where the message features contain interaction information between nodes; aggregate the message features based on the attention weights, and update the node states through a recurrent neural network to obtain dynamic node features;
[0032] Calculate the temporal attention weights of the dynamic node features, fuse the temporal attention weights with the historical feature matrix to obtain temporal context features; input the temporal context features into a long short-term memory network to update the memory state and obtain fused features;
[0033] Perform multi-scale feature extraction on the fused features and perform feature pooling to obtain a behavioral feature representation; calculate a behavioral classification score based on the behavioral feature representation and output the recognition result of the target behavior.
[0034] In a second aspect of the embodiments of the present invention, there is provided a warning system for illegal operations in dangerous areas of petrochemical industry based on deep learning, including:
[0035] A first unit for obtaining a target image sequence corresponding to video data of a petrochemical area, performing object detection and tracking on the target image sequence, obtaining the position coordinates of target objects in the target image sequence, and establishing a target index;
[0036] A second unit for establishing a temporal trajectory for each target object in the target image sequence based on the target index; constructing a multi-object interaction feature map according to the temporal trajectory, where the multi-object interaction feature map contains information such as the relative distance, relative speed, and relative direction between target objects;
[0037] A third unit for dynamically adjusting the weights of the multi-object interaction feature map in combination with the environmental parameters of the petrochemical area to obtain an environment-aware interaction feature map; calculating a safety situation score of the petrochemical area based on the environment-aware interaction feature map, where the safety situation score characterizes the overall safety state of the petrochemical area at the current moment;
[0038] A fourth unit for constructing a dynamic inference graph network for behavior recognition, inputting the environment-aware interaction feature map into the dynamic inference graph network, extracting node features through graph convolution operations, and updating the dynamic association relationship between nodes based on a message passing mechanism to obtain the recognition result of the target behavior.
[0039] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including:
[0040] A processor;
[0041] A memory for storing instructions executable by the processor;
[0042] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0043] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0044] The beneficial effects of this application are as follows:
[0045] By using a deep learning-based method to warn against illegal operations in petrochemical hazardous areas, the present invention can monitor and analyze the behaviors and interaction patterns of target objects such as personnel and equipment in the hazardous areas in real time, thereby effectively preventing the occurrence of safety accidents.
[0046] The present invention combines environmental parameters to dynamically adjust the weights of the multi-target interaction feature map, enabling the system to flexibly adjust the warning sensitivity according to the safety requirements under different environmental conditions, improving the accuracy and adaptability of the warning, and reducing false alarms and missed alarms.
[0047] The present invention uses a dynamic inference graph network for behavior recognition, captures the complex spatio-temporal relationships and interaction patterns between target objects through graph convolution operations and message passing mechanisms, can accurately identify potential illegal operation behaviors, provides timely and effective warning information for safety management personnel, and greatly improves the safety management level and accident prevention ability in petrochemical areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic flowchart of the method for warning against illegal operations in petrochemical hazardous areas based on deep learning according to an embodiment of the present invention;
[0049] Figure 2 It is a bar chart comparing the technical performance of target detection and tracking according to an embodiment of the present invention;
[0050] Figure 3 It is a flowchart of the group stability analysis method based on geometric center and trajectory similarity according to an embodiment of the present invention;
[0051] Figure 4 It is a bar chart comparing the performance of the petrochemical area safety situation assessment method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0054] Figure 1 This is a schematic flowchart of the method for warning against illegal operations in dangerous areas of petrochemical industry based on deep learning according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0055] Obtain a target image sequence corresponding to video data of a petrochemical area, perform object detection and tracking on the target image sequence, obtain the position coordinates of target objects in the target image sequence, and establish a target index;
[0056] Based on the target index, establish a temporal trajectory for each target object in the target image sequence; construct a multi-object interaction feature map according to the temporal trajectory, and the multi-object interaction feature map includes relative distance, relative speed, and relative direction information between target objects;
[0057] Combine environmental parameters of the petrochemical area to dynamically adjust the weights of the multi-object interaction feature map to obtain an environment-aware interaction feature map; calculate a safety situation score of the petrochemical area based on the environment-aware interaction feature map, and the safety situation score represents the overall safety state of the petrochemical area at the current moment;
[0058] Construct a dynamic inference graph network for behavior recognition, input the environment-aware interaction feature map into the dynamic inference graph network, extract node features through graph convolution operations, and update the dynamic association relationship between nodes based on the message passing mechanism to obtain the recognition result of the target behavior.
[0059] In an alternative embodiment, performing object detection and tracking on the target image sequence, obtaining the position coordinates of target objects in the target image sequence, and establishing a target index includes:
[0060] Use an object detection model to detect the target image sequence, obtain target objects in the target image sequence, and extract the position coordinates, size information, and category information of the target objects;
[0061] Assign a globally unique identifier to each target object, and establish a target index table. The target index table records the identifier, position coordinates, size information, category information, and timestamp of the target object. By comparing the visual features and motion features of target objects between adjacent frames, establish the temporal association of target objects, and associate target objects with the same visual features and continuous motion trajectories to the same identifier;
[0062] Based on the position coordinates recorded in the target index table, calculate the motion parameters of the target objects.
[0063] Receive a target image sequence containing multiple consecutive frames, such as a 30fps video stream with a resolution of 1920×1080 obtained from a video surveillance system. To perform object detection on this image sequence, the system uses a pre-trained deep learning object detection model, such as a detection framework based on a convolutional neural network. This model has been trained for common object categories such as people, vehicles, animals, etc., and can identify object targets with a confidence greater than a preset threshold (e.g., 0.7).
[0064] During the actual detection process, the system performs object detection operations on each frame of the image sequence. Taking the first frame of the image as an example, the object detection model identifies that there are 3 object targets in this frame: 1 car and 2 pedestrians. For each detected object target, the system extracts and records the following information: position coordinates (represented by the upper-left corner coordinates of the bounding box, such as (250, 350) for the car, (420, 380) for pedestrian 1, and (560, 400) for pedestrian 2), size information (represented by the width and height of the bounding box, such as 200×150 pixels for the car, 80×180 pixels for pedestrian 1, and 75×170 pixels for pedestrian 2), and category information (such as "car", "pedestrian", etc.).
[0065] To achieve object tracking, the system needs to assign a globally unique identifier to each detected object target. In the first frame, the system assigns the identifier "V 001 " to the detected car and assigns the identifiers "P 001 " and "P 002 " to the two pedestrians respectively. At the same time, the system establishes an object index table to record the identifier, position coordinates, size information, category information, and timestamp (the time information of this frame, such as "00:00:00.033" representing the 33rd millisecond of the video) of each object target.
[0066] When processing the second frame of the image, the object detection model also detects object targets. For example, it still detects 1 car and 2 pedestrians. The system needs to determine whether these objects are the same as those in the first frame, that is, perform object association. The system establishes temporal association by comparing the visual features and motion features of object targets between adjacent frames.
[0067] Visual feature comparison is achieved by extracting the appearance features of the target region. The system extracts color histograms, texture features, or depth feature representations from the target bounding box region and calculates the feature similarity between two objects. For example, the feature similarity between the car at coordinates (255, 352) in the second frame and the car in the first frame is 0.92, which exceeds the preset threshold of 0.8, and is initially determined to be the same object.
[0068] Motion feature comparison is achieved by predicting the motion trajectory of the target. The system predicts the position area where the target will appear in the current frame based on its position and speed in the previous frame. If the target detected in the current frame is within this predicted area, its association probability is increased. For the above-mentioned car target, based on the position (250, 350) in the first frame and the estimated speed (5, 2) pixels / frame, the predicted position in the second frame is approximately (255, 352), which is very close to the actual detected position, further confirming that this is the same vehicle.
[0069] When both the visual feature similarity and the motion trajectory prediction meet the association conditions, the system associates the target in the current frame with the target in the previous frame and inherits its identifier. In this example, the car in the second frame continues to use the identifier "V" 001 ", and the two pedestrians continue to use the identifiers "P" 001 " and "P" 002 " respectively.
[0070] If a new target object is detected in a certain frame, the system will assign a new unique identifier to it. For example, in the tenth frame, a new pedestrian is detected, and the system will assign the identifier "P" 003 " to it. On the contrary, if a tracked target is not detected in consecutive multiple frames (such as 5 frames), it is considered that the target has left the monitoring area, and its tracking status is marked as "left" in the target index table.
[0071] Based on the position coordinate data recorded in the target index table, the system calculates the motion parameters of the target object. For the car with the identifier "V" 001 ", by analyzing its position changes in 30 consecutive frames, its average speed is calculated to be 15 pixels / frame, the motion direction is 26 degrees (the angle relative to the horizontal axis), and the acceleration is 0.2 pixels / frame 2 . These motion parameters can be used to predict the future motion trajectory of the target and improve the stability of tracking.
[0072] To handle the situation where the target is occluded, the system implements an occlusion handling mechanism. When the target is partially occluded, the detected bounding box will change. At this time, the system mainly relies on the visual features of the visible part and the previous motion trajectory for association. For example, when the pedestrian "P" 001 " is partially occluded by roadside facilities, the system can still maintain the correct tracking association based on its upper body features and predicted position.
[0073] When the target is completely occluded, the system will continue to predict its trajectory based on the previous motion parameters for a short period of time (such as 10 frames) and try to restore the association when the target reappears. This mechanism greatly improves the tracking robustness of the system in complex scenarios.
[0074] Figure 2This is a bar chart comparing the performance of target detection and tracking technologies according to an embodiment of the present invention:
[0075] This figure shows a performance comparison of a technical solution with traditional methods and baseline algorithms across three key test metrics. The data in the figure shows that this technical solution achieved optimal performance across all test metrics: in terms of target detection accuracy, this solution achieved 53.8%, a significant improvement over the 47.2% of traditional methods and 43.5% of baseline algorithms; in terms of target tracking accuracy, this solution achieved 51.9%, exceeding the 47.6% of traditional methods and 39.4% of baseline algorithms; and in terms of temporal association accuracy, this solution achieved 56.3%, also outperforming the 45.1% of traditional methods and 42.2% of baseline algorithms. The data comparison across these three dimensions clearly shows that this technical solution has achieved comprehensive improvements over existing methods in core performance metrics such as target detection, tracking, and temporal association, with an average improvement of between 6 and 12 percentage points, fully demonstrating the advanced nature and practical value of this technical solution.
[0076] In an optional embodiment, based on the target index, a time series trajectory is established for each target object in the target image sequence; and a multi-target interaction feature map is constructed according to the time series trajectory, wherein the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects, including:
[0077] According to the identification code of the target object, the position coordinates of the target object are recorded in chronological order to obtain a time series coordinate sequence of the target object; the displacement and time difference of adjacent coordinate points in the time series coordinate sequence are calculated to obtain the movement speed and movement direction of the target object at different moments; and the position coordinates, the movement speed, and the movement direction are combined to form a time series trajectory of the target object;
[0078] The relative distance, relative speed and relative direction between target objects are calculated to construct a multi-target interaction feature map; the relative distance, relative speed and relative direction are converted into an interaction strength score according to preset distance weight, speed weight and direction weight; when the interaction strength score is greater than a first preset threshold, the corresponding target objects are divided into the same interaction group; the overlap of the motion trajectories of the target objects in the interaction group is calculated, and when the overlap of the motion trajectories is greater than a second preset threshold, the interaction group is determined to be a collaborative target group.
[0079] A temporal trajectory is established for each target object in the target image sequence based on the target index, and then a multi-target interaction feature map is constructed according to the temporal trajectory. The feature map contains the relative distance, relative speed and relative direction information between the target objects.
[0080] Record the position coordinates of the target object in chronological order according to the identification code of the target object, and obtain the time-series coordinate sequence of the target object. For example, for the target object with the identification code ID 001 the position coordinates in 5 consecutive frames can be recorded as: (120, 350), (125, 355), (132, 358), (140, 362), (148, 367). These coordinate points form the time-series coordinate sequence of the target object in chronological order.
[0081] Calculate the displacement and time difference between adjacent coordinate points in the time-series coordinate sequence to obtain the motion speed and motion direction of the target object at different times. Assume that the time interval between adjacent frames is 0.04 seconds. For the above coordinate sequence, the system calculates the displacement from the first frame to the second frame as 7.07 pixels (calculated by the distance between two points), and the time difference is 0.04 seconds. Therefore, the speed at this moment is 176.75 pixels / second. At the same time, by calculating the angle of the displacement vector, the motion direction is obtained as 45 degrees. The system performs similar calculations for all adjacent frames to obtain the complete speed and direction sequence.
[0082] Combine the position coordinates, motion speed, and motion direction to form the time-series trajectory of the target object. For the target object with the identification code ID 001 its time-series trajectory can be expressed as: {(120, 350, 0, 0), (125, 355, 176.75, 45), (132, 358, 184.23, 33), (140, 362, 206.16, 38), (148, 367, 210.24, 42)}, where each tuple contains coordinate, speed, and direction information.
[0083] Calculate the relative distance, relative speed, and relative direction between target objects to construct a multi-target interaction feature map. Assume that there are two target objects ID 001 and ID 002 in the scene. At a certain moment, their positions are (140, 362) and (200, 380) respectively, their motion speeds are 206.16 pixels / second and 186.42 pixels / second respectively, and their motion directions are 38 degrees and 65 degrees respectively. The system calculates the relative distance between them as 68.59 pixels, the relative speed as 89.74 pixels / second (calculated by the modulus of the velocity vector difference), and the relative direction as 27 degrees (calculated by the difference between the two directions).
[0084] Convert the relative distance, relative speed, and relative direction into an interaction intensity score according to preset distance weights, speed weights, and direction weights. Assume that the system sets the distance weight to 0.5, the speed weight to 0.3, the direction weight to 0.2, the distance normalization coefficient to 100 pixels, the speed normalization coefficient to 200 pixels / second, and the direction normalization coefficient to 90 degrees. For the above example, the normalized relative distance value is 0.686, the relative speed value is 0.449, and the relative direction value is 0.3. Calculate the interaction intensity score as: 0.5×(1 - 0.686) + 0.3×(1 - 0.449) + 0.2×(1 - 0.3) = 0.436.
[0085] When the interaction intensity score is greater than the first preset threshold, the system divides the corresponding target objects into the same interaction group. Assume that the first preset threshold is set to 0.4. Then, the interaction intensity score of the above two target objects, 0.436, is greater than the threshold 0.4. Therefore, they are divided into the same interaction group.
[0086] Calculate the motion trajectory coincidence degree of the target objects within the interaction group. When the motion trajectory coincidence degree is greater than the second preset threshold, determine the interaction group as a collaborative target group. The motion trajectory coincidence degree is calculated by analyzing the similarity of the future predicted paths of the target objects. Specifically, the system predicts the positions in the next 5 frames based on the current position, speed, and direction of each target. For ID 001 and ID 002 , the system predicts the positions of ID 001 in the next 5 frames as: (156, 373), (164, 379), (172, 385), (180, 391), (188, 397); predicts the positions of ID 002 in the next 5 frames as: (207, 393), (214, 406), (221, 419), (228, 432), (235, 445).
[0087] Quantify the trajectory coincidence degree by calculating the direction consistency and position proximity of the predicted trajectories. The specific calculation includes the average direction difference of the predicted trajectories and the average distance of the predicted positions. For the above example, the calculated direction consistency is 0.85 (the closer this value is to 1, the more similar the directions), and the position proximity is 0.65 (the closer this value is to 1, the closer the positions). Considering these two indicators comprehensively, the calculated trajectory coincidence degree is 0.75.
[0088] Assume that the second preset threshold is set to 0.7. Then, the trajectory coincidence degree of the above interaction group, 0.75, is greater than the threshold 0.7. Therefore, the system determines the interaction group as a collaborative target group. This indicates that these two target objects not only approach each other spatially but also exhibit collaborative behavior in terms of motion characteristics.
[0089] Through the above processing, the system successfully identifies the target groups with collaborative behaviors in the target image sequence, providing a basis for subsequent behavior analysis and event prediction. This multi-object analysis method based on temporal trajectories and interaction features can effectively capture the interaction relationships between target objects and achieve an accurate understanding of multi-object behaviors in complex scenarios.
[0090] In an alternative embodiment, the method further includes:
[0091] Calculating the geometric center coordinates and the movement direction of the group; calculating the average distance between all target objects within the group and the geometric center to obtain the spatial distribution radius of the group; identifying the dominant object within the group, where the dominant object is the target object closest to the geometric center of the group and with the smallest deviation from the movement direction of the group; calculating the trajectory similarity between other target objects and the dominant object, where the trajectory similarity is equal to the reciprocal of the difference in the movement directions of the target object and the dominant object; calculating the stability coefficient of the group according to the trajectory similarity;
[0092] When the stability coefficient of the group is greater than the third preset threshold, mark this group as a stable collaborative group; when the spatial distribution radius of the group is greater than the fourth preset threshold or the trajectory similarity of the target object is less than the fifth preset threshold, trigger a group dissolution warning.
[0093] As Figure 3 shown, the method includes:
[0094] Calculating the geometric center coordinates of the group, which is obtained by averaging the position coordinates of all target objects within the group. For example, assume there are three target objects within the group, and their position coordinates are (10, 15), (12, 18), and (14, 12) respectively. Then the geometric center coordinates of the group are ((10 + 12 + 14) / 3, (15 + 18 + 12) / 3), that is, (12, 15). At the same time, the method also calculates the movement direction of the group, which can be obtained by averaging the movement vectors of all target objects within the group. If the movement vectors of the three target objects are (2, 1), (1, 2), and (2, 2) respectively, then the movement direction of the group is ((2 + 1 + 2) / 3, (1 + 2 + 2) / 3), that is, (1.67, 1.67).
[0095] Calculating the spatial distribution radius of the group, which is obtained by calculating the distances between all target objects within the group and the geometric center and then taking the average. Using the above example, the distances from the three target objects to the geometric center (12, 15) are respectively sqrt((10 - 12) 2 +(15 - 15) 2 ) = 2, sqrt((12 - 12) 2 +(18 - 15) 2) = 3 and sqrt((14 - 12) 2 +(12 - 15) 2 ) = 3.61. The average distance is then (2 + 3 + 3.61) / 3 = 2.87, which is the spatial distribution radius of the group.
[0096] The method also includes identifying the dominant object within the group. The dominant object refers to the target object that is closest to the geometric center of the group and has the smallest deviation from the movement direction of the group. In the above example, the distances of the three target objects to the geometric center have been calculated as 2, 3, and 3.61 respectively. Next, calculate the deviation of the movement direction of each target object from the movement direction of the group (1.67, 1.67). Assume the deviations of the three target objects from the movement direction of the group are 0.2, 0.1, and 0.3 respectively. Then, the first target object is the closest to the geometric center of the group, but the second target object has the smallest deviation from the movement direction of the group. In this case, by comprehensively considering the distance and direction deviation, the first target object can be selected as the dominant object.
[0097] Calculating the trajectory similarity between other target objects and the dominant object is the next step of the method. The trajectory similarity is equal to the reciprocal of the difference in the movement directions of the target object and the dominant object. For example, if the differences in the movement directions of the second and third target objects from the dominant object (i.e., the first target object) are 0.3 and 0.5 respectively, then their trajectory similarities with the dominant object are 1 / 0.3 = 3.33 and 1 / 0.5 = 2 respectively.
[0098] Based on the trajectory similarity, the method calculates the stability coefficient of the group. The stability coefficient can be obtained by taking the average of the trajectory similarities between all non - dominant objects and the dominant object. In the above example, the stability coefficient of the group is (3.33 + 2) / 2 = 2.67.
[0099] When the stability coefficient of the group is greater than the third preset threshold, the method marks the group as a stable collaborative group. For example, if the third preset threshold is set to 2.5, then the stability coefficient of the above - mentioned group, 2.67, is greater than 2.5, so it is marked as a stable collaborative group. This means that the target objects within the group have a high degree of collaboration and are performing a certain collaborative task.
[0100] The method further includes a mechanism for triggering a group dissolution warning. When the spatial distribution radius of the group is greater than a fourth preset threshold or the trajectory similarity of the target object is less than a fifth preset threshold, a group dissolution warning is triggered. Suppose the fourth preset threshold is set to 3 and the fifth preset threshold is set to 1.5. In the above example, the spatial distribution radius of the group is 2.87, which is less than the fourth preset threshold of 3, and thus the dissolution warning will not be triggered. However, if the trajectory similarity of a target object to the leading object is 1.2, which is less than the fifth preset threshold of 1.5, then a group dissolution warning will be triggered. This indicates that the target object is deviating from the group and the group is at risk of dissolution.
[0101] The method can further refine the calculation of the stability coefficient. In addition to considering the trajectory similarity, the distance between the target object and the leading object can also be considered. For example, higher weights can be assigned to target objects closer to the leading object, so that the stability coefficient can more accurately reflect the synergy of the group.
[0102] The method can also include a mechanism for dynamically adjusting the preset thresholds. Based on historical data and the current environment, the system can automatically adjust the third, fourth, and fifth preset thresholds to adapt to the group behavior in different scenarios. For example, in a crowded environment, the fourth preset threshold (the threshold of the spatial distribution radius) needs to be reduced because in this case, even a collaborative group is forced to maintain a larger spatial distribution.
[0103] Through the above method, the group behavior of target objects can be effectively identified and tracked, especially the formation and dissolution of collaborative groups, providing valuable information for target tracking and behavior analysis.
[0104] In an alternative implementation, the multi-target interaction feature map is dynamically weighted according to the environmental parameters of the petrochemical area to obtain an interaction feature map for environmental perception; calculating the safety situation score of the petrochemical area based on the interaction feature map for environmental perception includes:
[0105] Obtain the environmental parameters of the petrochemical area, where the environmental parameters include the temperature values of multiple temperature monitoring points and the gas concentration values of multiple gas monitoring points;
[0106] Construct a temperature field influence factor based on the temperature values of the temperature monitoring points. The temperature field influence factor is obtained by calculating the ratio of the distance from the target to the temperature monitoring point to the temperature field attenuation coefficient and combining the normalized temperature value of the temperature monitoring point; construct a gas concentration influence factor based on the gas concentration values of the gas monitoring points. The gas concentration influence factor is obtained by calculating the ratio of the height difference from the target to the gas monitoring point to the gas diffusion coefficient and combining the normalized concentration value of the gas monitoring point;
[0107] Obtain a multi-object interaction feature map, where the multi-object interaction feature map includes distance features and speed features between objects; adjust the distance features based on the temperature field influence factor and the gas concentration influence factor to obtain distance features for environmental perception; adjust the speed features based on the spatial gradient of the temperature field influence factor and the gas concentration influence factor to obtain speed features for environmental perception;
[0108] Fuse the distance features for environmental perception and the speed features for environmental perception to establish an interaction feature map for environmental perception; calculate the aggregation degree of the objects based on the interaction feature map for environmental perception, and the aggregation degree of the objects is obtained by measuring the similarity between objects;
[0109] Divide the petrochemical area into multiple sub-areas, and calculate the regional risk degree based on the influence intensity of the hazard sources in each sub-area; combine the aggregation degree of the objects and the regional risk degree, and obtain the safety situation score through non-linear mapping.
[0110] Obtain the environmental parameters of the petrochemical area, and these environmental parameters include the temperature values of multiple temperature monitoring points and the gas concentration values of multiple gas monitoring points. For example, 10 temperature monitoring points and 8 gas monitoring points are deployed inside a certain petrochemical area. The temperature monitoring points collect the temperature data in the area in real time, and the gas monitoring points monitor the concentration data of combustible gases in real time. These monitoring points transmit the data to the safety situation assessment system through the Internet of Things technology.
[0111] Construct a temperature field influence factor based on the temperature values of the temperature monitoring points. The specific implementation process is as follows: For each monitored object, calculate the Euclidean distance from the object to each temperature monitoring point. For example, the distance from object A to temperature monitoring point T1 is 15 meters; normalize the original temperature values of the temperature monitoring points. For example, map the temperature range of 35°C - 85°C to 0 - 1, and the normalized value corresponding to 67°C measured by temperature monitoring point T1 is 0.64; preset the temperature field attenuation coefficient to 20 meters, indicating that the temperature influence significantly attenuates within 20 meters; the temperature field influence factor is obtained by calculating the ratio of the distance from the object to the temperature monitoring point to the temperature field attenuation coefficient and then multiplying it by the normalized temperature value. For example, the temperature field influence factor of object A affected by temperature monitoring point T1 is 0.64×(15 / 20)=0.48.
[0112] Construct gas concentration impact factors based on the gas concentration values at gas monitoring points. The specific implementation process is as follows: Calculate the height differences from the target to each gas monitoring point. For example, target B is 2.5 meters higher than gas monitoring point G1; Preset the gas diffusion coefficient to 5 meters, which represents the gas diffusion ability in the vertical direction; Normalize the original concentration values of the gas monitoring points. For example, map the concentration range of 0 - 500 ppm to 0 - 1. The normalized value corresponding to 350 ppm measured at gas monitoring point G1 is 0.7; The gas concentration impact factor is obtained by calculating the ratio of the height difference from the target to the gas monitoring point to the gas diffusion coefficient and then multiplying it by the normalized concentration value. For example, the gas concentration impact factor of target B affected by gas monitoring point G1 is 0.7×(2.5 / 5)=0.35.
[0113] Obtain a multi-target interaction feature map, which contains the distance features and velocity features between targets. The distance features represent the spatial relative position relationship between targets, and the velocity features represent the movement trends of targets. For example, the distance between target A and target B is 8 meters, the velocity of target A is 1.2 m / s, the direction angle is 45 degrees, the velocity of target B is 0.8 m / s, and the direction angle is 120 degrees.
[0114] Adjust the distance features based on the temperature field impact factors and gas concentration impact factors to obtain the distance features for environmental perception. The adjustment method is to multiply the original distance value by the weighted sum of the environmental impact factors. For example, the original distance between target A and target B is 8 meters. Considering the temperature field impact factor of target A is 0.48 and the gas concentration impact factor is 0.2 (weights are 0.6 and 0.4 respectively), and the temperature field impact factor of target B is 0.3 and the gas concentration impact factor is 0.35 (weights are 0.6 and 0.4 respectively), then the distance feature for environmental perception is 8×[1-(0.48×0.6+0.2×0.4+0.3×0.6+0.35×0.4) / 2]=5.68 meters.
[0115] Adjust the velocity features based on the spatial gradients of the temperature field impact factors and gas concentration impact factors to obtain the velocity features for environmental perception. The spatial gradient represents the change rate of environmental factors in space, and the calculation method is to divide the difference in environmental factors between adjacent monitoring points by the distance between them. When adjusting the velocity features, weight-combine the original velocity vector and the environmental factor gradient vector. For example, the original velocity of target A is 1.2 m / s, the direction angle is 45 degrees. At its position, the temperature field gradient is 0.05 / m, the direction angle is 30 degrees, the gas concentration gradient is 0.03 / m, and the direction angle is 60 degrees. The velocity after the combined influence of the comprehensive environmental gradient (weights are 0.7 and 0.3 respectively) is 1.35 m / s, and the direction angle is 42 degrees.
[0116] Fuse the distance feature and speed feature of environmental perception to establish an interactive feature map of environmental perception. The fusion method is to perform weighted summation on the distance feature matrix and speed feature matrix according to preset weights. For example, the distance feature weight is 0.6 and the speed feature weight is 0.4 to obtain the fused feature map.
[0117] Calculate the aggregation degree of the target based on the interactive feature map of environmental perception. The aggregation degree is obtained through the similarity measurement between targets. The similarity calculation method is to calculate based on the distance and speed differences between targets in the fused feature map using the Gaussian kernel function. For example, at a certain moment, 5 targets are detected to be active in the area, and the aggregation degree calculated based on the interactive feature map of environmental perception is 0.72, indicating a relatively high degree of target aggregation.
[0118] Divide the petrochemical area into multiple sub-areas, and calculate the regional risk degree based on the influence intensity of the hazard sources in each sub-area. For example, divide a certain petrochemical area into 6 sub-areas. There is 1 high-pressure storage tank in sub-area 1 with an influence intensity of 0.9, and there are 2 medium-pressure pipelines in sub-area 2 with influence intensities of 0.6 and 0.5 respectively. The comprehensive calculated risk degree of sub-area 1 is 0.9, and the risk degree of sub-area 2 is 0.78.
[0119] Combine the aggregation degree of the target and the regional risk degree to obtain the safety situation score through non-linear mapping. The non-linear mapping uses the S-shaped function to map the weighted combination of the aggregation degree and the risk degree to the scoring range of 0 - 100. For example, at a certain moment, the aggregation degree is 0.72, the average regional risk degree is 0.83, and the weighting ratio is 1:1.5. The safety situation score obtained through the S-shaped function mapping is 35 points, indicating that the current safety situation is in a relatively vigilant state.
[0120] Figure 4 This is the bar chart for performance comparison of the petrochemical area safety situation assessment method in the embodiments of the present invention:
[0121] This figure shows the performance comparison results of this technical solution with traditional environment perception methods and three solutions that do not consider environmental factors in terms of key safety assessment indicators. From the specific data, this technical solution shows obvious advantages in all three core assessment indicators: in terms of the accuracy rate of danger warning, this solution reaches 53.3%, significantly higher than 48.6% of the traditional environment perception method and 38.2% of the method that does not consider environmental factors; in terms of the recognition rate of abnormal behaviors, this solution reaches 52.7%, better than 47.5% of the traditional method and 39.1% of the method that does not consider environmental factors; in the assessment of the emergency response speed, this solution reaches 57.6%, also exceeding 48.8% of the traditional method and 38.9% of the method that does not consider environmental factors. Through the comparison of this set of data, it can be clearly seen that by integrating the environment perception ability, this technical solution has achieved significant improvements in key safety indicators such as danger warning, abnormal behavior recognition, and emergency response. Compared with the traditional method, it has increased by an average of 5 - 9 percentage points, and compared with the method that does not consider environmental factors, the improvement is more obvious, reaching 15 - 19 percentage points, fully verifying the important role of environment perception in improving the performance of safety assessment.
[0122] In an alternative embodiment, a dynamic inference graph network is constructed for behavior recognition. The interactive feature map of the environment perception is input into the dynamic inference graph network, and node features are extracted through graph convolution operations. Based on the message passing mechanism, the dynamic association relationship between nodes is updated, and the recognition result of the target behavior includes:
[0123] Obtain the position feature, speed feature, and environment feature of the target, construct a node feature vector from the position feature, the speed feature, and the environment feature, and construct an edge feature vector based on the association relationship between the node feature vectors; form the initialization structure of the dynamic inference graph network with the node feature vector and the edge feature vector;
[0124] Construct a temporal connection matrix based on the initialization structure of the dynamic inference graph network, and perform spatial feature aggregation on the node feature vector according to the temporal connection matrix to obtain multi - layer node features;
[0125] Calculate the attention weight for the multi - layer node features, and the attention weight is obtained by splicing the node features and passing through a non - linear transformation; perform weighted aggregation on the multi - layer node features based on the attention weight to obtain enhanced node features;
[0126] Generate message features according to the enhanced node features and the edge feature vector, and the message features contain the interaction information between nodes; aggregate the message features based on the attention weight, and update the node state through a recurrent neural network to obtain dynamic node features;
[0127] Calculate the temporal attention weights of the dynamic node features, fuse the temporal attention weights with the historical feature matrix to obtain temporal context features; input the temporal context features into a long short-term memory network to update the memory state to obtain fused features;
[0128] Perform multi-scale feature extraction on the fused features and perform feature pooling to obtain a behavioral feature representation; calculate a behavioral classification score based on the behavioral feature representation and output the recognition result of the target behavior.
[0129] Obtain the position features, speed features, and environmental features of the target. The position features include the coordinate information of the target in a two-dimensional or three-dimensional space, such as (x, y) or (x, y, z); the speed features include the speed components of the target in each direction, such as (v x , v y , v z ). The environmental features include information such as the distribution of surrounding objects and the road structure. Taking the autonomous driving scenario as an example, the position feature of a pedestrian can be represented as (120.5, 85.3), indicating its pixel position in the image coordinate system; the speed feature can be represented as (1.2, 0.8), with the unit of meters per second; the environmental features can include surrounding vehicles, signal light states, etc. These features are integrated into a node feature vector, and the vector dimension can be set between 64 and 256 according to the application scenario.
[0130] The association relationship between nodes is calculated based on spatial distance and motion similarity. For example, if the Euclidean distance between two targets is less than a preset threshold (such as 5 meters), a connection is established and an edge feature vector is constructed. The edge feature vector encodes the relative position, relative speed, and interaction possibility between the node pair. For targets A and B, their relative position can be represented as (3.2, 1.5), the relative speed is (0.5, -0.3), and the interaction possibility is 0.78. The dimension of the edge feature vector is usually set between 32 and 128.
[0131] The initial structure of the dynamic inference graph network consists of node feature vectors and edge feature vectors. In a typical scenario, it contains 5 to 20 nodes, and each node is connected to 3 to 5 adjacent nodes, forming a sparsely connected graph structure.
[0132] Based on the initialized graph structure, a temporal connection matrix is constructed to capture the interaction relationships of the targets over time. The temporal connection matrix is a matrix of size N×N (N is the number of nodes), and the matrix element values represent the connection strength between nodes. For example, the temporal connection matrix for a 5-node scenario is [[0,0.8,0.5,0,0],[0.8,0,0.6,0.3,0],[0.5,0.6,0,0.7,0.4],[0,0.3,0.7,0,0.9],[0,0,0.4,0.9,0]]. This matrix is used to perform spatial feature aggregation on node features to achieve the transfer of information between connected nodes. After multiple graph convolution operations, multi-layer node features with different receptive fields are obtained, and the feature dimension of each layer can be 128.
[0133] When calculating the attention weights for multi-layer node features, the node features of different layers are concatenated and then passed through a non-linear transformation to obtain the weight coefficients. Specifically, for three-layer features, after concatenation, a 384-dimensional vector (128×3) is obtained, and through a fully connected layer, it is mapped to a 3-dimensional attention weight vector, such as [0.4,0.35,0.25], indicating the degree of emphasis on the three-layer features. Based on these weights, weighted aggregation is performed on the multi-layer node features to obtain enhanced 128-dimensional node features.
[0134] Message features are generated based on the enhanced node features and edge feature vectors. Message features are generated by combining the source node features with the edge features to generate a vector containing interaction information. For example, for the message transmitted from node A to node B, the node features of A (128-dimensional) can be concatenated with the edge features of A-B (64-dimensional), and after transformation, 128-dimensional message features are obtained. These message features contain interaction information between nodes, such as behavioral intentions and influence degrees.
[0135] The calculated attention weights are used to aggregate the message features, and the node states are updated through a gated recurrent unit (GRU). Each node receives messages from all adjacent nodes, and after weighted aggregation, the input features are obtained, and then the hidden state is updated through the GRU to obtain dynamic node features. The hidden state dimension of the GRU is set to 128, and the information flow is controlled through the update gate and the reset gate to achieve selective memory of historical information.
[0136] Temporal attention weights of the dynamic node features are calculated to evaluate the importance of features at different time points. For a time series of length 10, the obtained attention weights are [0.02,0.03,0.05,0.07,0.1,0.13,0.15,0.18,0.12,0.15], indicating that recent features are usually more important. These weights are multiplied by the historical feature matrix to obtain temporal context features.
[0137] Input the temporal context features into a long short-term memory network (LSTM) to update the memory state and obtain the fused features. The LSTM includes an input gate, a forget gate, and an output gate, and can effectively capture long-term dependencies. Taking a single-layer LSTM as an example, the hidden state dimension is 256, and the retention ratio of historical information is controlled by the forget gate coefficient (such as 0.3) to generate the fused features.
[0138] Perform multi-scale feature extraction on the fused features. Use convolutional layers with different convolutional kernel sizes (such as 3×3, 5×5, 7×7) to capture spatial patterns in different ranges and obtain multi-scale features. Perform global average pooling and max pooling operations on these features to generate a compact representation of the behavior features, with a dimension of 512.
[0139] Calculate the behavior classification scores based on the representation of the behavior features. Map the 512-dimensional features to a vector with the number of behavior categories (such as for 8 types of behaviors, map to an 8-dimensional vector) through a fully connected layer to obtain the scores for each category, such as [0.05, 0.02, 0.15, 0.63, 0.03, 0.07, 0.01, 0.04]. Convert the scores to a probability distribution through the softmax function, and the category with the highest probability is used as the recognition result. In this example, the fourth type of behavior (probability 0.63) is recognized as the behavior of the target.
[0140] In the second aspect of the embodiments of the present invention, a warning system for illegal operations in dangerous areas of petrochemical industry based on deep learning is provided, including:
[0141] The first unit is used to obtain the target image sequence corresponding to the video data of the petrochemical area, perform object detection and tracking on the target image sequence, obtain the position coordinates of the target objects in the target image sequence, and establish a target index;
[0142] The second unit is used to establish a temporal trajectory for each target object in the target image sequence based on the target index; construct a multi-target interaction feature map according to the temporal trajectory, and the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects;
[0143] The third unit is used to dynamically adjust the weights of the multi-target interaction feature map in combination with the environmental parameters of the petrochemical area to obtain an environment-aware interaction feature map; calculate the safety situation score of the petrochemical area based on the environment-aware interaction feature map, and the safety situation score characterizes the overall safety state of the petrochemical area at the current moment;
[0144] The fourth unit is used to construct a dynamic inference graph network for behavior recognition, input the environment-aware interaction feature map into the dynamic inference graph network, extract node features through graph convolutional operations, and update the dynamic association relationship between nodes based on the message passing mechanism to obtain the recognition result of the target behavior.
[0145] In a third aspect of the embodiments of the present invention, an electronic device is provided, including:
[0146] A processor;
[0147] A memory for storing instructions executable by the processor;
[0148] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0149] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0150] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for warning against illegal operations in dangerous areas of the oil and petrochemical industries based on deep learning, characterized in that, Including: Obtain a target image sequence corresponding to video data of a petrochemical area, perform target detection and tracking on the target image sequence, obtain the position coordinates of target objects in the target image sequence, and establish a target index; Based on the target index, establish a temporal trajectory for each target object in the target image sequence; construct a multi-target interaction feature map according to the temporal trajectory, and the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects; Combine the environmental parameters of the petrochemical area to dynamically adjust the weights of the multi-target interaction feature map to obtain an environment-aware interaction feature map; Calculate the safety situation score of the petrochemical area based on the environment-aware interaction feature map, and the safety situation score represents the overall safety state of the petrochemical area at the current moment; Construct a dynamic inference graph network for behavior recognition, input the environment-aware interaction feature map into the dynamic inference graph network, extract node features through graph convolution operations, and update the dynamic association relationship between nodes based on the message passing mechanism to obtain the recognition result of the target behavior.
2. The method according to claim 1, characterized in that Performing target detection and tracking on the target image sequence, obtaining the position coordinates of target objects in the target image sequence, and establishing a target index includes: Use a target detection model to detect the target image sequence, obtain target objects in the target image sequence, and extract the position coordinates, size information, and category information of the target objects; Assign a globally unique identifier to each target object, and establish a target index table. The target index table records the identifier, position coordinates, size information, category information, and timestamp of the target object. By comparing the visual features and motion features of target objects between adjacent frames, establish the temporal association of target objects, and associate target objects with the same visual features and continuous motion trajectories to the same identifier; Based on the position coordinates recorded in the target index table, calculate the motion parameters of the target object.
3. The method according to claim 1, characterized in that Based on the target index, establish a temporal trajectory for each target object in the target image sequence; construct a multi-target interaction feature map according to the temporal trajectory, and the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects includes: According to the identification code of the target object, record the position coordinates of the target object in chronological order to obtain the temporal coordinate sequence of the target object; calculate the displacement and time difference between adjacent coordinate points in the temporal coordinate sequence to obtain the motion speed and motion direction of the target object at different times; combine the position coordinates, the motion speed, and the motion direction to form the temporal trajectory of the target object; The relative distance, relative speed and relative direction between target objects are calculated to construct a multi-target interaction feature map; the relative distance, relative speed and relative direction are converted into an interaction strength score according to preset distance weight, speed weight and direction weight; when the interaction strength score is greater than a first preset threshold, the corresponding target objects are divided into the same interaction group; the overlap of the motion trajectories of the target objects in the interaction group is calculated, and when the overlap of the motion trajectories is greater than a second preset threshold, the interaction group is determined to be a collaborative target group.
4. The method according to claim 3, characterized in that The method further comprises: Calculate the geometric center coordinates and movement direction of the group; calculate the average distance between all target objects in the group and the geometric center to obtain the spatial distribution radius of the group; identify the dominant object in the group, the dominant object being the target object closest to the group's geometric center and with the smallest deviation from the group's movement direction; calculate the trajectory similarity between other target objects and the dominant object, the trajectory similarity being equal to the inverse of the difference in movement direction between the target object and the dominant object; and calculate the group's stability coefficient based on the trajectory similarity. When the stability coefficient of the group is greater than the third preset threshold, the group is marked as a stable collaborative group; when the spatial distribution radius of the group is greater than the fourth preset threshold or the trajectory similarity of the target object is less than the fifth preset threshold, a group disbanding warning is triggered.
5. The method according to claim 1, wherein Dynamically weighting the multi-objective interactive feature graph in combination with environmental parameters of the petrochemical area to obtain an environmentally-aware interactive feature graph; and calculating a security situation score for the petrochemical area based on the environmentally-aware interactive feature graph includes: Acquiring environmental parameters of the petrochemical area, wherein the environmental parameters include temperature values of multiple temperature monitoring points and gas concentration values of multiple gas monitoring points; A temperature field influence factor is constructed based on the temperature value of the temperature monitoring point, and the temperature field influence factor is obtained by calculating the ratio of the distance from the target to the temperature monitoring point to the temperature field attenuation coefficient, combined with the normalized temperature value of the temperature monitoring point; a gas concentration influence factor is constructed based on the gas concentration value of the gas monitoring point, and the gas concentration influence factor is obtained by calculating the ratio of the height difference from the target to the gas monitoring point to the gas diffusion coefficient, combined with the normalized concentration value of the gas monitoring point; Acquire a multi-target interaction feature map, the multi-target interaction feature map including distance features and speed features between targets; adjust the distance features based on the temperature field influencing factor and the gas concentration influencing factor to obtain an environmentally perceived distance feature; and adjust the speed features based on spatial gradients of the temperature field influencing factor and the gas concentration influencing factor to obtain an environmentally perceived speed feature; Performing feature fusion on the environment-perceived distance feature and the environment-perceived speed feature to establish an environment-perceived interaction feature map; calculating the target aggregation based on the environment-perceived interaction feature map, wherein the target aggregation is obtained by measuring the similarity between targets; The petrochemical area is divided into multiple sub-areas, and the regional hazard level is calculated based on the impact intensity of the hazard sources in each sub-area. The safety situation score is obtained by combining the concentration of the target and the regional hazard level through nonlinear mapping.
6. The method according to claim 1, characterized in that, A dynamic reasoning graph network is constructed for behavior recognition. The interactive feature graph of environmental perception is input into the dynamic reasoning graph network. Node features are extracted through graph convolution operations. The dynamic association relationship between nodes is updated based on the message passing mechanism. The recognition results of the target behavior are obtained, including: Acquire the position features, speed features, and environmental features of the target, construct the position features, speed features, and environmental features into node feature vectors, construct edge feature vectors based on the association relationship between the node feature vectors; and form an initialization structure of a dynamic reasoning graph network with the node feature vectors and the edge feature vectors; Building a temporal connection matrix based on the initialization structure of the dynamic reasoning graph network, and performing spatial feature aggregation on the node feature vectors according to the temporal connection matrix to obtain multi-layer node features; Calculating attention weights for the multi-layer node features, where the attention weights are obtained by concatenating the node features and performing a nonlinear transformation; performing weighted aggregation on the multi-layer node features based on the attention weights to obtain enhanced node features; Generate message features based on the enhanced node features and the edge feature vectors, wherein the message features include interaction information between nodes; aggregate the message features based on the attention weights, and update the node states through a recurrent neural network to obtain dynamic node features; Calculating the temporal attention weight of the dynamic node feature, fusing the temporal attention weight with the historical feature matrix to obtain a temporal context feature; inputting the temporal context feature into the long short-term memory network, updating the memory state to obtain a fused feature; Multi-scale feature extraction and feature pooling are performed on the fused features to obtain a behavior feature representation; a behavior classification score is calculated based on the behavior feature representation, and a recognition result of the target behavior is output.
7. A warning system for illegal operations in hazardous areas of the oil and petrochemical industries based on deep learning, which is used to implement the method described in any one of claims 1-6, characterized in that, include: The first unit is configured to obtain a target image sequence corresponding to video data of a petrochemical area, perform target detection and tracking on the target image sequence, obtain position coordinates of a target object in the target image sequence, and establish a target index; The second unit is configured to establish a time series trajectory for each target object in the target image sequence based on the target index; construct a multi-target interaction feature map according to the time series trajectory, wherein the multi-target interaction feature map includes relative distance, relative speed, and relative direction information between target objects; The third unit is used to dynamically adjust the weight of the multi-objective interaction feature map in combination with the environmental parameters of the petrochemical area to obtain an environmental perception interaction feature map; Calculating a safety situation score for the petrochemical area based on the interactive feature graph of environmental perception, wherein the safety situation score represents the overall safety status of the petrochemical area at a current moment; The fourth unit is used to construct a dynamic inference graph network for behavior recognition. The interactive feature map obtained from the environmental perception is input into the dynamic inference graph network. Node features are extracted through graph convolution operations, and the dynamic association relationships between nodes are updated based on the message passing mechanism to obtain the recognition result of the target behavior.
8. An electronic device, characterized in that: It includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Cigarette station illegal operation early warning method and system based on artificial intelligence
CN120014814A
Deep learning method for multiple object tracking from video
US20240144489A1
Cited By
Pipe gallery UWB personnel accurate positioning and track backtracking safety control method and system
CN121364443A