Full-station intelligent video control method and device

By using direction gradient feature recognition and behavior pattern coding in coal transportation corridors, accurate identification of equipment, personnel and environment and real-time early warning of abnormal behaviors are achieved, the problem of low intelligence in the existing technology is solved, and the automation and safety management capabilities of the monitoring system are improved.

CN120510563APending Publication Date: 2025-08-19YIXIAN JIAYU PHOTOVOLTAIC POWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510548419.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing monitoring system cannot realize real-time accurate identification and abnormal warning of dynamic operation behavior in coal transportation corridors. It is low in intelligence and relies on manual intervention to fully automate and real-time early warning.

Method used

The direction gradient feature recognition equipment, staff and environmental areas in the video frame are marked, the postures of staff in continuous video frames are analyzed, the operation behavior is identified using behavior pattern encoding, and warning information is sent when abnormal behavior is recognized.

Benefits of technology

Real-time accurate identification and abnormal warning of dynamic operation behaviors in industrial scenarios are realized, monitoring efficiency and safety management capabilities are improved, manual intervention is reduced, and security risks are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510563A_ABST
    Figure CN120510563A_ABST
Patent Text Reader

Abstract

The invention discloses a whole-station intelligent video management and control method and device, and relates to the technical field of intelligent monitoring, and the method comprises the steps: recognizing equipment, workers and environment regions in a video frame according to the direction gradient features of all regions in the video frame; marking the equipment, the working personnel and the environment area in the video frame; determining an operation behavior of the worker according to the posture of the worker in the continuous video frame under the condition that the labeling is completed; identifying the operation behavior through a behavior mode code to obtain an identification result; and when the identification result is that the operation behavior belongs to the abnormal behavior, sending warning information to a preset device, and marking the operation behavior. According to the invention, real-time accurate identification and abnormal early warning of the dynamic operation behavior in the industrial scene are realized, the monitoring efficiency and the safety management capability are obviously improved, the manual intervention demand is reduced, and the safety risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent monitoring technology, and in particular to a method and device for controlling and managing intelligent video on a whole site. Background Art

[0002] In modern industrial production, coal conveyor corridors, as a crucial material transportation link, face complex operating environments and diverse safety challenges. Their densely packed interiors, high dust concentrations, and frequent human activity make them prone to safety hazards such as belt tearing, deviation, and improper operation.

[0003] Currently, the industry's most commonly used monitoring method uses a system combining high-definition cameras and sensors to monitor coal transportation corridors. However, the intelligence level of existing systems still needs to be improved, relying on manual intervention and failing to achieve full automation and real-time early warning. Therefore, achieving real-time, accurate identification of dynamic operational behaviors and anomaly warnings in industrial scenarios has become a pressing issue.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The purpose of this application is to provide a full-station intelligent video control method and device, aiming to solve the technical problem of how to achieve real-time and accurate identification of dynamic operation behaviors and abnormal warning in industrial scenarios.

[0006] To achieve the above objectives, this application proposes a site-wide intelligent video control method, which includes:

[0007] Identifying equipment, staff, and environmental areas in the video frame based on directional gradient features of each area in the video frame;

[0008] marking the device, the staff, and the environmental area in the video frame;

[0009] When the labeling is completed, determining the operator's operation behavior according to the operator's posture in the continuous video frames;

[0010] Identify the operation behavior through behavior pattern coding to obtain an identification result;

[0011] When the recognition result is that the operation is abnormal, a warning message is sent to a preset device and the operation behavior is marked.

[0012] In one embodiment, after the labeling is completed, the step of determining the operator's operating behavior based on the operator's posture in the continuous video frames includes:

[0013] When the annotation is completed, the posture of the image area corresponding to the worker is estimated to obtain a preliminary heat map of the key points of the human body;

[0014] Performing non-maximum suppression on the preliminary heat map to obtain key points and connection information;

[0015] Connecting the key points according to the connection information to obtain a posture estimation result;

[0016] Tracking the worker's motion trajectory according to the posture estimation results of continuous video frames to obtain a motion trajectory map;

[0017] The motion trajectory diagram is input into a behavior recognition model to obtain the operating behavior of the staff member, and the behavior recognition model is obtained by training a three-dimensional convolutional network.

[0018] In one embodiment, the step of performing non-maximum suppression on the preliminary heat map to obtain key points and connection information includes:

[0019] Performing a local maximum search on the preliminary heat map to obtain a local maximum point;

[0020] Filtering the points to be deleted from the local maximum points to obtain key points, wherein the confidence of the points to be deleted is less than a preset confidence threshold;

[0021] The connection information of the key points is determined according to the part affinity field.

[0022] In one embodiment, the step of connecting the key points according to the connection information to obtain a posture estimation result includes:

[0023] Obtain a key point-connection score table according to the connection information and the key points;

[0024] Selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table;

[0025] Adding the key point connection to a human body posture set;

[0026] Returning to the step of selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table until the key point-connection score table is empty;

[0027] The posture estimation result is obtained by connecting and constructing the key points in the human posture set.

[0028] In one embodiment, after the labeling is completed, the step of performing posture estimation on the image area corresponding to the worker to obtain a preliminary heat map of key points of the human body includes:

[0029] After the annotation is completed, the image area corresponding to the worker is input into the backbone network to obtain a feature map. The backbone network is composed of multiple convolutional networks and pooling networks.

[0030] Inputting the feature map into a posture estimation network to obtain key point features of the human body, wherein the posture estimation network is composed of a sampling module, a skip connection module, a residual module and a supervision module;

[0031] A convolution operation is performed on the key point features of the human body to obtain a preliminary heat map of the key points of the human body.

[0032] In one embodiment, the step of identifying the equipment, staff, and environment areas in the video frame based on the directional gradient features of each area in the video frame includes:

[0033] Converting the video frame into a grayscale image, and performing gamma correction on the grayscale image to obtain a gamma-corrected image;

[0034] Calculating the gradient magnitude and direction of each pixel in the gamma-corrected image to obtain a gradient map;

[0035] Dividing the gradient map into multiple connected regions, and summing the gradient amplitudes of the pixels within the connected regions to form a gradient direction histogram;

[0036] Normalizing the gradient direction histogram to obtain a feature vector;

[0037] The feature vector is compared with a feature database to determine the equipment, staff, and environmental areas in the video frame.

[0038] In one embodiment, the step of identifying the operation behavior through behavior pattern coding to obtain an identification result includes:

[0039] Performing behavior pattern coding on the operation behavior to obtain a behavior pattern coding result;

[0040] Compare the behavior pattern encoding result with each pattern in the preset abnormal behavior pattern library and calculate the similarity score;

[0041] The recognition result is obtained according to the similarity score and a preset abnormal behavior score threshold.

[0042] In one embodiment, after the step of sending a warning message to a preset device and marking the operation behavior when the operation behavior is abnormal, the method further includes:

[0043] Creating an index table based on the annotated content in the video frame and the position of the video frame corresponding to the annotated content in the surveillance video;

[0044] When a search instruction is received, obtaining a search target according to the search instruction;

[0045] The index table is searched according to the search target to obtain a surveillance video segment corresponding to the search target.

[0046] In one embodiment, after the step of sending a warning message to a preset device and marking the operation behavior when the operation behavior is abnormal, the method further includes:

[0047] Obtain historical surveillance videos;

[0048] Analyze the historical surveillance video according to the annotations in the historical surveillance video, and generate and save a video summary;

[0049] When there is no abnormal behavior annotation in the annotation and no abnormal event in the video summary, the historical surveillance video is deleted.

[0050] In addition, to achieve the above objectives, the present application also proposes a station-wide intelligent video control device, which includes:

[0051] An image recognition module is used to identify equipment, staff, and environmental areas in the video frame based on directional gradient features of each area in the video frame;

[0052] a labeling module, configured to label the device, the staff, and the environmental area in the video frame;

[0053] An operation behavior recognition module is used to determine the operation behavior of the staff member based on the posture of the staff member in the continuous video frames when the labeling is completed;

[0054] An abnormal behavior recognition module is used to recognize the operation behavior through behavior pattern coding and obtain a recognition result;

[0055] The exception handling module is used to send a warning message to a preset device and mark the operation behavior when the recognition result is an abnormal behavior.

[0056] In addition, to achieve the above-mentioned purpose, the present application also proposes a full-site intelligent video control device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the full-site intelligent video control method as described above.

[0057] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the full-site intelligent video control method as described above are implemented.

[0058] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the full-site intelligent video control method as described above.

[0059] One or more technical solutions proposed in this application have at least the following technical effects:

[0060] First, the system identifies equipment, personnel, and environmental areas based on the directional gradient features of each region in the video frame. This gradient-based recognition method accurately captures edge and texture information in the image, enabling highly robust object differentiation even under complex backgrounds and lighting conditions. The system then labels the identified objects within the video frames, clearly identifying each object's location and type, significantly improving the readability and practicality of the monitoring system. Furthermore, the system analyzes the changes in personnel posture across consecutive video frames to determine their operational behavior. By capturing dynamic changes in real time, it accurately identifies the personnel's specific movements, providing a more detailed monitoring method for safety management. The system then identifies operational behavior through behavioral pattern coding and rapidly outputs the recognition results. Finally, if the recognition result indicates abnormal behavior, the system sends a warning message to a pre-set device and further labels the abnormal operation behavior, promptly notifying relevant personnel to take measures and prevent potential safety incidents. This application achieves real-time and accurate recognition of dynamic operational behavior and abnormality warnings in industrial scenarios, significantly improving monitoring efficiency and safety management capabilities, reducing the need for manual intervention, and mitigating safety risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0062] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0063] Figure 1 A flowchart illustrating the first embodiment of the method for controlling the entire website's intelligent video is provided in this application;

[0064] Figure 2 A schematic diagram of the annotation scene provided in Example 1 of the method for controlling the entire site's intelligent video;

[0065] Figure 3 This is a flowchart of the second embodiment of the method for controlling the entire site's intelligent video.

[0066] Figure 4 This is a schematic diagram of the module structure of the whole-station smart video control device according to an embodiment of the present application;

[0067] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the full-site intelligent video control method in the embodiment of this application.

[0068] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0069] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0070] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0071] In modern industrial production, coal conveyor corridors, a critical material transportation link, face complex operating environments and diverse safety challenges, including densely packed equipment, high dust concentrations, and frequent human activity. These environments are prone to safety hazards such as belt tearing, misalignment, and improper operation. Currently, the industry primarily relies on monitoring systems utilizing high-definition cameras and sensors, but these systems lack a high level of intelligence and still require manual intervention, failing to achieve full automation and real-time warnings.

[0072] The main solution of this embodiment is that the system identifies equipment, personnel, and environmental areas based on directional gradient features in video frames, ensuring high-precision target differentiation even under complex conditions. The system also annotates the location and type of each target within the frame to improve monitoring readability. It then analyzes changes in personnel posture across consecutive frames to determine specific operational behaviors. It then rapidly identifies and outputs behavioral pattern codes. If abnormal behavior is detected, it immediately issues a warning and annotates it, notifying relevant personnel to promptly address the issue and prevent safety incidents.

[0073] It should be noted that the execution subject of the embodiments of the present application can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a video control system, etc. The following uses the video control system as an example to illustrate this embodiment and the following embodiments.

[0074] Based on this, the embodiment of the present application provides a method for controlling intelligent video of the entire station, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the site-wide intelligent video control method of this application.

[0075] In this embodiment, the site-wide intelligent video control method includes steps S10 to S50:

[0076] Step S10 , identifying equipment, staff, and environment areas in the video frame according to directional gradient features of each area in the video frame.

[0077] It should be noted that the directional gradient feature is a feature description used for image recognition and target detection. It describes the shape and texture information of an image by calculating and statistically analyzing the gradient direction histogram of a local area of the image. Specifically, the main idea of the directional gradient feature is that the appearance and shape of a local target can be well described by the directional density distribution of the gradient or edge, because the gradient information is mainly concentrated near the edge and contour. The gradient calculation formula is:

[0078] G x =I*S x , G y =I*S y

[0079] Where I is the pixel matrix of the input image (grayscale image or single-channel image), S x Is the horizontal Sobel operator (3×3 convolution kernel), used to detect vertical edges, S y is the vertical Sobel operator, used to detect horizontal edges, G x and G y are the gradient components of the image in the horizontal and vertical directions.

[0080] Gradient amplitude and direction calculation formula:

[0081]

[0082] Among them, G is the gradient magnitude, which indicates the edge strength, and θ is the gradient direction (ranging from 0 to 2π), which indicates the edge direction.

[0083] Small blocks are selected in areas with complex textures to preserve details, and large blocks are selected in smooth areas to reduce noise. At this time, the directional gradient feature is used for adaptive blocking, and the formula is as follows:

[0084]

[0085] Among them, B i,j is the optimal block size selected at image position (i, j), σ(G i,j,k ) is the standard deviation of the gradient amplitude G in the region when the block size is k×k, which measures the local texture complexity. k is the candidate block size, which is used to adapt to targets of different sizes.

[0086] In low light conditions, we rely on thermal imaging, and in complex backgrounds, we rely on depth information. We use a multi-spectral gradient fusion method, and the formula is as follows:

[0087] G fused =αG RGB +βG thermal +γG depth

[0088] Among them, G RGB is the gradient of the visible light image (detecting color and texture), G thermal is the thermal imaging gradient (highlighting heating equipment or human body), G depth is the depth camera gradient (distinguishing foreground from background), α, β, γ are learnable weights (α+β+γ=1), which are dynamically adjusted through the cross-modal attention mechanism.

[0089] Directional gradient feature block normalization formula:

[0090]

[0091] Where v is the original vector of the gradient direction histogram within the block (e.g., a 9-bin histogram), ||v|| 2 It is the square of the L2 norm of the vector (that is, the sum of the squares of each element), and ε is a minimum value (such as 1e-5) to prevent the denominator from being zero, which leads to numerical instability.

[0092] In this example, equipment generally refers to the various mechanical devices and facilities used for transportation, control, and monitoring within the coal conveyor corridor, such as coal conveyor belts, belt drive units, coal drop pipes, coal crushers, iron removers, and cleaning devices. Workers are those who perform operations, inspections, and maintenance within the coal conveyor corridor. The environmental zone refers to all spaces and background areas within the coal conveyor corridor, excluding equipment and personnel. These zones primarily reflect the overall operating environment and physical state of the coal conveyor corridor.

[0093] It's understandable that, first, the video control system extracts single frames from the real-time surveillance video stream and divides each frame into multiple small local regions for detailed image analysis. Secondly, the system calculates the directional gradient features for each local region and compares these extracted directional gradient features with a pre-prepared feature database to identify equipment, personnel, and environmental areas in the video frame, enabling accurate classification and location of different targets.

[0094] As an example, the step of identifying the equipment, staff and environmental areas in the video frame based on the directional gradient features of each area in the video frame includes: converting the video frame into a grayscale image, and performing gamma correction on the grayscale image to obtain a gamma-corrected image; calculating the gradient amplitude and direction of each pixel point in the gamma-corrected image to obtain a gradient map; dividing the gradient map into multiple connected areas, and counting the sum of the gradient amplitudes of the pixels within the connected areas to form a gradient direction histogram; normalizing the gradient direction histogram to obtain a feature vector; comparing the feature vector with a feature database to determine the equipment, staff and environmental areas in the video frame.

[0095] A grayscale image is an image that converts the color information of each pixel in a color image into grayscale values. Grayscale values are usually represented by an 8-bit integer ranging from 0 (black) to 255 (white), with intermediate values representing different shades of gray.

[0096] Gamma correction is a nonlinear adjustment method for image brightness, used to compensate for the nonlinear characteristics of images on display devices. During image acquisition and display, image brightness may be distorted due to the nonlinear response of the device. Gamma correction adjusts the image's brightness values so that it is correctly displayed on the display device.

[0097] The gradient magnitude represents the rate of change in brightness of a pixel in an image, reflecting the strength of an edge. The gradient direction represents the direction of the brightness change, reflecting the direction of the edge. Gradient magnitude and direction are typically calculated by calculating the horizontal and vertical gradients of a pixel.

[0098] A gradient map is an image of the same size as the original image, where the value of each pixel represents the gradient magnitude at that point. Gradient maps are primarily used to highlight edges and textures in an image, as gradient magnitudes are typically higher in edge regions and lower in smooth regions. Gradient maps allow for more intuitive visualization of image structure, providing a foundation for subsequent feature extraction.

[0099] A connected region is a region of pixels in an image that share similar characteristics (such as gradient direction or grayscale value). In directional gradient feature extraction, the image is typically divided into multiple small connected regions (or cells) to statistically analyze the gradient information in these local regions. The size and shape of these connected regions can be adjusted to suit the specific task; for example, they can be rectangular, circular, or other shapes.

[0100] The sum of the gradient magnitudes is the sum of the gradient magnitudes of all pixels within a connected region. It quantitatively describes the edge strength within the connected region. By calculating the sum of the gradient magnitudes, we can assess the significance of the edges within the region and provide data support for subsequent feature statistics.

[0101] The gradient direction histogram is the core part of the directional gradient feature, which is used to describe the distribution of gradient directions in connected areas. Specifically, for each connected area, the system counts the gradient directions of all pixels in the area and records their distribution in a histogram. The horizontal axis of the histogram represents the gradient direction (usually divided into multiple intervals, such as 0° to 180°), and the vertical axis represents the sum of the gradient amplitudes in each direction interval. Through the gradient direction histogram, the texture and shape information in the connected area can be captured, providing key features for subsequent target recognition.

[0102] The eigenvector is a numerical vector obtained by normalizing the data in the gradient directional histogram. The purpose of normalization is to eliminate the influence of factors such as lighting and contrast on the features, making the features more robust. The eigenvector is the final representation of the directional gradient feature, which can effectively describe the shape and texture information of the object in the image.

[0103] A feature database is a collection of feature vectors for various known targets (such as equipment, personnel, and environmental areas). These feature vectors are obtained by extracting and training features from a large number of sample images and serve as reference models. In practice, the video control system compares the extracted feature vectors with the models in the feature database and determines the target category in the video frame by calculating similarity or distance. The accuracy and completeness of the feature database directly impacts the accuracy of target recognition.

[0104] First, the video control system converts the video frames into grayscale images, removing color information and retaining only brightness information. This allows for more efficient processing of edge and texture features in the image. The system then performs gamma correction on the grayscale image, adjusting the image's brightness distribution and enhancing contrast, making it easier to analyze under varying lighting conditions. Next, the system calculates the gradient magnitude and direction of each pixel in the gamma-corrected image to generate a gradient map, highlighting edges and structural information. The system then divides the gradient map into connected regions and calculates the sum of the gradient magnitudes of the pixels within each region to form a gradient direction histogram, which captures the texture and shape characteristics of the local area. The system then normalizes the gradient direction histogram to generate a feature vector, eliminating the effects of lighting and contrast variations and making the features more robust. Finally, the system compares the feature vector with a pre-set feature database. By matching similarities, the system identifies the equipment, personnel, and environmental areas in the video frame, enabling accurate identification and classification of different targets.

[0105] Step S20: marking the equipment, the staff, and the environmental area in the video frame.

[0106] It should be noted that annotation refers to the clear marking and classification of equipment, staff, and environmental areas identified in the video frame for subsequent monitoring, management, and analysis. The purpose of annotation is to distinguish these targets from the complex background and give them specific identification or labels so that the monitoring system can quickly identify and understand the video content. The specific forms of annotation may include the following: (1) Bounding box: In the video frame, a rectangular box is drawn for each identified equipment, personnel, environmental area, etc. to frame the location and range of the target. The bounding box is usually represented by coordinates (such as the coordinates of the upper left corner and the lower right corner), which can clearly define the boundaries of the target. (2) Category label: A category label is assigned to each target to clarify its identity. For example, "conveyor belt", "staff", "wall", etc. These labels are usually displayed in text form on the video frame to help the system quickly identify the target type. (3) Color coding: Different colors are used to mark different types of areas or targets. For example, red represents equipment, green represents staff, and blue represents environmental areas. Color coding can intuitively distinguish targets and facilitate rapid visual recognition. (4) Unique identifier: For targets that need to be further tracked (such as specific equipment or personnel), a unique identifier (such as a number or ID) can be assigned. This helps to continuously track the same target in multiple frames of video and achieve more refined management. (5) Text description: Add a short text description to the video frame to explain the status of the target or related information, for example, "the belt conveyor is operating normally", "the person is not wearing a safety helmet", etc. This annotation method can provide more detailed information and assist the system in making decisions. Region annotation formula:

[0107] P outout =BilinearInterp(P input , Grid(x,y))

[0108] Among them, P input is the region of interest (RoI) in the input feature map, Grid(x,y) divides the RoI into a grid of fixed size, and BilinearInterp is a bilinear interpolation algorithm that calculates the pixel value of each sampling point in the grid to avoid the quantization error of traditional RoIPooling.

[0109] It is understandable that, first, after identifying the equipment, personnel, and environmental areas in the video frame, the video control system will select an appropriate annotation method based on the target type. For example, it will draw a bounding box for the environmental area and label its category name, assign color codes or unique identifiers to personnel, and label the function or status of the equipment through color or text descriptions. Secondly, the system will overlay this annotation information on the video frame in a visual manner to ensure that the annotation content is clear, visible, and easy to understand. The annotation information provides a clear reference basis for automated analysis and early warning, improves monitoring efficiency and accuracy, and ensures the safe operation and management efficiency of the coal transportation corridor.

[0110] Please refer to Figure 2 , Figure 2 A schematic diagram of the labeled scene provided for Example 1 of the full-station intelligent video control method of this application shows a number of key components and devices, which are clearly labeled with yellow markings. Various valves, pumps, sensors and other control devices can be seen in the figure, which are arranged along the pipeline to monitor and regulate the flow of fluids. Each labeled point may represent a specific device or monitoring point, and this information helps to quickly identify various parts of the system for maintenance or troubleshooting. In addition, this detailed labeling also helps the video control system to identify and track the status of each component through image recognition algorithms, and realize real-time monitoring and management of the operating status of equipment in industrial scenarios, thereby improving safety and efficiency. Through such labeling, training data can be provided for the behavior recognition model, enabling it to learn the normal operating behavior of different components, and then identify abnormal behavior, issue warnings in time, and prevent potential industrial accidents.

[0111] Step S30: When the labeling is completed, the operating behavior of the staff is determined according to the posture of the staff in the continuous video frames.

[0112] It should be noted that posture refers to the posture and movement state of the human body in an image or video. It is determined by detecting the key points of the human body (such as joints, head, hands, etc.) and their relative position relationships. The position and distribution of these key points can reflect the posture information of the human body, such as standing, bending, stretching, and other actions.

[0113] Operational behaviors refer to the series of actions or activities performed by workers in specific scenarios. These behaviors can be identified and understood through changes in posture. For example, in a coal conveyor corridor, workers' operational behaviors may include checking equipment, moving items, and operating control panels. These behaviors are identified by analyzing changes in human posture in consecutive video frames.

[0114] It is understandable that in this embodiment, after the annotation is completed, the system first extracts the key point information of the staff in the continuous video frames, estimates the posture of these key points, and forms a skeletal structure diagram of the human body. Then, the system constructs the spatiotemporal information of these key points (i.e., the position changes of the key points in the continuous frames) into a spatiotemporal graph model. The nodes in the model represent the key points, and the edges represent the spatial and temporal relationships between the key points. Next, the system performs a convolution operation on the spatiotemporal graph model to extract high-level features, and classifies the behavior through a classifier (such as Softmax), and determines the specific operational behavior of the staff based on the classification results.

[0115] Step S40: Identify the operation behavior through behavior pattern coding to obtain an identification result.

[0116] It should be noted that behavioral pattern encoding refers to the conversion of human posture and movement characteristics into standardized quantifiable data for describing and identifying specific behavioral patterns. In this way, the system can decompose complex movements into a series of recognizable patterns for subsequent analysis and classification.

[0117] The recognition result refers to the conclusion about the operation behavior drawn by the system after encoding and analyzing the behavior pattern. It indicates the specific operation behavior that the staff is performing, such as normal operation, violation behavior or other specific actions.

[0118] As you can understand, the system first converts the worker's posture information (such as key point locations and movement trajectories) extracted from continuous video frames into standardized behavioral pattern codes. The system then inputs these behavioral pattern codes into a pre-trained behavior recognition model, which compares the codes with the characteristics of known behavioral patterns to determine the worker's ongoing operational behavior.

[0119] As an example, the step of identifying the operation behavior through behavior pattern coding and obtaining the identification result includes: performing behavior pattern coding on the operation behavior to obtain a behavior pattern coding result; comparing the behavior pattern coding result with each pattern in a preset abnormal behavior pattern library to calculate a similarity score; and obtaining the identification result based on the similarity score and a preset abnormal behavior score threshold.

[0120] Behavior pattern encoding involves converting an action into a standardized sequence of numerical values or symbols through specific algorithms (such as key point detection and feature extraction) to describe the characteristics of that action. For example, by detecting the position and movement trajectory of key points on the human body, these values are converted into a series of numerical values that uniquely represent a specific behavior pattern.

[0121] The preset abnormal behavior pattern library is a collection of all known abnormal behavior patterns. These patterns are trained using a large amount of sample data. Each pattern corresponds to a specific abnormal behavior, such as "not wearing a helmet," "entering a hazardous area," or "improper operation." Each pattern in the pattern library has a corresponding feature vector or code, which is used to compare with the behavioral pattern code detected in real time.

[0122] The similarity score refers to the degree of match between the real-time detected behavior pattern code and the patterns in the preset abnormal behavior pattern library. The similarity score is typically a value between 0 and 1, with higher values indicating greater similarity. For example, if the real-time detected behavior pattern code is compared with the "not wearing a helmet" pattern in the pattern library, the resulting similarity score is 0.85, indicating that the two are very similar.

[0123] The preset abnormal behavior score threshold is a pre-set value used to determine whether the similarity score is high enough to identify a particular abnormal behavior. For example, if the threshold is set to 0.7, then when the similarity score is greater than or equal to 0.7, the system will identify the behavior as abnormal; if it is less than 0.7, the behavior is not considered abnormal. Suppose the system detects a staff member's behavior pattern coded as [0.1, 0.5, 0.8, 0.3]. The preset abnormal behavior pattern library contains a pattern coded as "entering a dangerous area" as [0.2, 0.4, 0.7, 0.3]. By calculating the similarity score between the two, the result is 0.8. If the preset abnormal behavior score threshold is 0.7, then because the similarity score of 0.8 is greater than the threshold, the system will identify the staff member's behavior as abnormal, indicating "entering a dangerous area."

[0124] First, the system encodes the operator's behavior pattern. By extracting the key locations and motion features of the operator from the video frame, it converts these into a standardized sequence of numerical values, generating the resulting behavior pattern encoding. This process aims to transform complex motion information into quantifiable data for subsequent processing and analysis. Secondly, the system compares the resulting behavior pattern encoding with each pattern in a pre-set abnormal behavior pattern library, calculating a similarity score between the two using cosine similarity or Euclidean distance algorithms. This process quantifies the degree of match between the behavior pattern encoding and known abnormal behavior patterns, providing a basis for identification. Finally, the system determines the calculated similarity score based on a pre-set abnormal behavior score threshold. If the similarity score is greater than or equal to the threshold, the operator is deemed abnormal and the corresponding recognition result is output. If the similarity score is less than the threshold, the operator is deemed normal. In this way, the system can accurately identify abnormal behavior, promptly detect potential safety hazards, and effectively enhance the intelligence and security of the monitoring system.

[0125] Step S50: When the recognition result is that the operation is abnormal, a warning message is sent to a preset device and the operation behavior is marked.

[0126] It should be noted that abnormal behavior refers to behavior that does not conform to normal operating specifications or preset behavioral patterns. For example, in a coal conveyor corridor, workers not wearing safety equipment, entering dangerous areas, or improper operation are all considered abnormal behavior. Warning messages are prompts generated by the system when abnormal behavior is identified. They are used to notify relevant personnel or systems of potential safety hazards or violations. Preset devices are terminals or devices pre-set by the system to receive warning messages, such as display screens in monitoring centers, administrators' mobile phones, or security alarm systems.

[0127] It is understandable that, first, when the system detects that the recognition result is abnormal behavior, it will trigger the built-in warning generation module to format the specific information of the abnormal behavior (such as behavior type, occurrence time, location coordinates, etc.) into a structured warning message. This is done to ensure the accuracy and readability of the information and facilitate the rapid communication of key content. Secondly, the system sends the warning information to pre-designated devices such as computers in the monitoring center, managers' mobile phones or other terminal devices through preset communication protocols (such as network transmission, SMS push or local notification). This process ensures that relevant personnel can receive alerts in a timely manner and take corresponding measures, thereby effectively shortening the response time. Finally, the system will mark the abnormal behavior in the corresponding video frame, such as by drawing a bounding box, adding text descriptions or using highlight colors to clearly identify the location and type of the abnormal behavior. Such marking not only allows the system to quickly locate the problem in the real-time picture, but also provides intuitive evidence for subsequent event investigation and analysis, further improving the practicality and management efficiency of the monitoring system.

[0128] As an example, when the operation behavior is an abnormal behavior, after the step of sending a warning message to a preset device and marking the operation behavior, it also includes: establishing an index table based on the marked content in the video frame and the position of the video frame corresponding to the marked content in the surveillance video; when a search instruction is received, obtaining a search target according to the search instruction; searching the index table according to the search target to obtain a surveillance video segment corresponding to the search target.

[0129] Annotation content refers to the specific information about the operation behavior annotated in the video frame, including the type of abnormal behavior, the location where the behavior occurred (such as coordinate information), timestamp, and possible additional descriptions (such as "personnel number" or "equipment number"). This information is displayed on the video frame in a visual manner (such as text, bounding boxes, color markings, etc.) to intuitively indicate abnormal behavior.

[0130] An index table is a data structure used to record the positional relationship between annotations and corresponding video frames in surveillance videos. It typically contains detailed information about the annotations (such as the behavior type, timestamp, frame number, and time offset), as well as the frame's offset or timestamp within the surveillance video file. The index table's purpose is to quickly locate video frames containing specific annotations, thereby improving video retrieval efficiency.

[0131] A search command refers to a query request entered by a user to find specific events or behaviors in the monitoring system. The search command can be a query command based on keywords (such as "not wearing a safety helmet"), time range (such as "January 1, 2024 08:00-10:00"), personnel number or other specific conditions. The user submits a retrieval request to the system through the search command to quickly find relevant monitoring video content.

[0132] A search target is the specific content a user wants to find, typically annotated information related to abnormal behavior. For example, a user might want to find all "entering a dangerous area" events within a certain time period, or all abnormal behavior records for a specific person. The search target is parsed based on the search instruction and guides the system's search within the index table.

[0133] Surveillance video segments are continuous video clips extracted from surveillance videos that contain specific annotations. After the system retrieves relevant annotations from the index table based on the search target, it extracts the video segments containing the annotations based on the corresponding video frame positions. Surveillance video segments can range from a few seconds to several minutes, depending on the duration of the annotations and user needs.

[0134] First, the system will traverse all annotated video frames, extract the annotated content and the location information of the frame in the surveillance video, and store this information in the form of key-value pairs in the index table. The purpose of this is to provide structured data support for subsequent rapid retrieval, greatly reducing the amount of calculation during retrieval. Secondly, when the system receives a search instruction input by the user, it will parse the keywords or conditions in the instruction and convert these conditions into specific search targets, clarifying the content that the system needs to find and ensuring the accuracy of the retrieval. Finally, the system searches the index table according to the search target, quickly locates the video frame containing the target, and then extracts the complete surveillance video segment based on the frame's location information for the user to view. This process not only improves retrieval efficiency, but also allows users to quickly obtain the required monitoring content, significantly improving the practicality of the monitoring system and user experience.

[0135] As an example, when the operation behavior is an abnormal behavior, after the step of sending a warning message to a preset device and marking the operation behavior, it also includes: obtaining historical surveillance video; analyzing the historical surveillance video according to the annotations in the historical surveillance video, generating and saving a video summary; when there is no abnormal behavior annotation in the annotation and there is no abnormal event in the video summary, deleting the historical surveillance video.

[0136] Historical surveillance video refers to recorded surveillance video data stored in the system. These videos record surveillance scenes over a period of time and are typically used for post-event analysis, incident investigation, or long-term storage. They are the product of the system's continuous monitoring of the monitored area and include both normal operations and possible abnormal events.

[0137] A video summary summarizes and summarizes the content of historical surveillance videos. By extracting key frames, important events, or behavioral snippets, a short video clip or text description is generated. The purpose of a video summary is to quickly convey the core information in a video, helping users quickly understand the key points of a surveillance video without having to watch the entire video. It typically includes a brief description of important time points, event types, or unusual behaviors.

[0138] Absence of abnormal behavior annotation means that no abnormal behavior is marked in the annotation information of the video frame.

[0139] Abnormal events are behaviors or events recorded in surveillance video that do not conform to normal operating procedures or pre-set safety standards. These events may include personnel violations (such as entering prohibited areas or not wearing safety equipment), abnormal equipment operation (such as equipment failure or belt tearing), or any other situation that may affect safety and normal operation.

[0140] First, the system retrieves recorded historical surveillance videos from storage devices. These videos record past surveillance scenes and provide a data foundation for subsequent analysis. Then, based on the annotation information in the historical surveillance videos, the system analyzes the video content frame by frame, extracting key frames containing abnormal behavior or important events. These key frames are spliced together in chronological order to generate a video summary, which is then saved to a database or file. The video summary quickly conveys the core information of the video, allowing users to quickly understand the surveillance content. Finally, the system checks the annotation information and video summary. If there are no abnormal behavior annotations in the annotations and no abnormal events are recorded in the video summary, it indicates that the historical surveillance video contains no important or abnormal information. The system will delete the historical surveillance video to save storage space and optimize data management.

[0141] This embodiment provides a method for intelligent video control throughout a site. First, the system identifies equipment, personnel, and environmental areas based on the directional gradient features of each region in a video frame. This gradient-based recognition method accurately captures edge and texture information in the image, enabling highly robust target differentiation even under complex backgrounds and lighting conditions. Next, the system labels the identified targets within the video frames, clearly identifying each target's location and type, significantly improving the readability and practicality of the monitoring system. Furthermore, the system analyzes changes in personnel posture within consecutive video frames to determine their operational behavior. By capturing dynamic changes in real time, it accurately identifies the specific actions of personnel, providing a more detailed monitoring method for safety management. Subsequently, the system identifies operational behavior through behavioral pattern coding and rapidly outputs the recognition results. Finally, if the recognition result indicates abnormal behavior, the system sends a warning message to a pre-set device and further labels the abnormal operation behavior, promptly notifying relevant personnel to take measures and prevent potential safety incidents. This embodiment achieves real-time and accurate recognition of dynamic operational behavior and abnormality warnings in industrial scenarios, significantly improving monitoring efficiency and safety management capabilities, reducing the need for manual intervention, and mitigating safety risks.

[0142] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 , Figure 3 This is a flow chart of the second embodiment of the site-wide smart video control method of this application. Step S30 of the site-wide smart video control method includes steps S31 to S35:

[0143] Step S31: After the labeling is completed, the posture of the image area corresponding to the worker is estimated to obtain a preliminary heat map of the key points of the human body.

[0144] It should be noted that the image area corresponding to the staff refers to the part marked in the video frame that contains the staff. In the surveillance video, the position and size of the staff may change over time. Therefore, the marked image area provides a clear range for subsequent analysis, ensuring that the system only processes the part containing the staff, avoiding unnecessary waste of computing resources.

[0145] Pose estimation is a computer vision technology used to identify and track human posture. It analyzes pixel information in images or video frames, detects the positions of key points on the human body, and infers the overall posture of the human body. The purpose of pose estimation is to extract the human body's posture from the image and provide basic data for subsequent behavioral analysis. The heat map generation for pose estimation is usually implemented using a convolutional neural network, and its formula can be expressed as:

[0146] H=σ(W*I+b)

[0147] Where I is the input image region, W is the weight of the convolution kernel, b is the bias term, σ is the activation function, and H is the generated heat map.

[0148] Human keypoints are the key parts used to describe human posture, such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. These keypoints are the core of posture estimation. By detecting the position and relative relationships of these points, the system can accurately infer the posture and movement of the human body. For example, by detecting keypoints on the shoulders and elbows, the degree of arm bending can be determined.

[0149] A preliminary heatmap is an intermediate result generated during the pose estimation process. It represents the possible locations of key points in the image. A heatmap is a two-dimensional array of the same size as the input image, where the value of each pixel represents the probability of that location being a key point. For example, a heatmap for a head key point would have higher values for pixels at the head location and lower values for pixels elsewhere. This preliminary heatmap is generated using a deep learning model (such as a convolutional neural network) and subsequently processed to precisely locate the key points.

[0150] It can be understood that, first, the system locates the image area corresponding to the staff member in the annotated video frame. This area is clearly identified during the annotation process to ensure the accuracy and efficiency of subsequent processing; secondly, the system applies a deep learning algorithm to the image area for posture estimation, analyzes the pixel information in the image area through a convolutional neural network, predicts the position of the key points of the human body, and generates a heat map for each key point. The pixel value in the heat map indicates the probability that the position is a key point; finally, the system integrates the heat maps of all key points into a preliminary heat map, providing a basis for subsequent precise positioning of key points and posture recognition.

[0151] As an example, when the labeling is completed, the posture of the image area corresponding to the staff member is estimated to obtain a preliminary heat map of the key points of the human body. The steps include: when the labeling is completed, the image area corresponding to the staff member is input into the backbone network to obtain a feature map, and the backbone network is composed of multiple convolutional networks and pooling networks; the feature map is input into the posture estimation network to obtain the key point features of the human body, and the posture estimation network is composed of a sampling module, a jump connection module, a residual module and a supervision module; a convolution operation is performed on the key point features of the human body to obtain a preliminary heat map of the key points of the human body.

[0152] The backbone network refers to the basic neural network architecture used to extract image features. It is usually composed of multiple convolutional layers and pooling layers. It is used to extract multi-level feature representations from the input image. Its function is to convert the input image data into feature maps with semantic information, providing basic features for subsequent tasks (such as posture estimation).

[0153] A feature map is a representation of image features after processing by a convolutional neural network. It is a multidimensional array that contains local and global feature information of the input image. Each channel of the feature map corresponds to a different feature of the input image, such as edges, texture, or shape. The feature map is usually smaller than the input image but contains higher-level semantic information.

[0154] A CNN (Convolutional Neural Network) is a deep learning architecture used to process data with a grid structure, such as images. It extracts local features of the input data through convolutional layers. Convolutional layers use convolution kernels to slide over the input data, performing convolution operations and generating feature maps. Convolutional networks can automatically learn features in images, such as edges, textures, and shapes. A pooling network is a network structure that includes a pooling layer. The pooling layer is used to reduce the spatial dimension of the feature map, reducing the amount of computation and the number of parameters, while retaining important features. Common pooling operations include max pooling and average pooling, which respectively take the maximum or average value of a local region, thereby achieving feature downsampling.

[0155] A pose estimation network (PAN) is a neural network architecture specifically designed for human pose estimation. This architecture consists of a series of modules that gradually reduce and restore resolution. It typically further processes the feature maps extracted by the backbone network to predict the locations of key points. Its training process typically involves the following steps: First, a large amount of image data with the locations of key points annotated is collected. This data serves as training samples, providing learning targets for the network. These annotated images are then fed into the network, which uses forward propagation to calculate the predicted key point locations and compares them with the annotated true locations to calculate a loss function. Next, a backpropagation algorithm is used to adjust the network's weight parameters based on the loss function to minimize the difference between the predicted and true values. Finally, after multiple iterations of training, the network gradually learns how to accurately identify the locations of key points in the image, forming a model capable of efficient pose estimation. In the PAN, a sampling module adjusts the resolution of the feature maps, a skip connection module transfers feature information, a residual module enhances the network's training capabilities, and a supervision module optimizes network performance. These modules work together to enable the network to accurately detect human pose in complex image environments.

[0156] Human key point features refer to the feature information related to human key points (such as head, shoulders, elbows, etc.) extracted from the feature map. These features are used in subsequent convolution operations to generate heat maps of key points.

[0157] A sampling module is a network module used to adjust the resolution of feature maps. It can resize feature maps through upsampling or downsampling operations. Upsampling increases the resolution of feature maps, while downsampling reduces it. In pose estimation networks, sampling modules are used to align feature maps at different levels for feature fusion. A skip connection module, such as a residual connection, is a connection method used in neural networks to directly transfer features. Skip connections allow certain layers in a network to connect directly to subsequent layers, bypassing intermediate layers, thereby avoiding vanishing or exploding gradients and improving network training efficiency and performance. In pose estimation networks, skip connections can directly transfer low-level features to higher-level layers, enhancing feature representation. A residual module is a network module based on residual learning and is commonly used to construct residual networks. By introducing skip connections, a residual module enables the network to learn the residual (i.e., the difference) between input and output, thereby simplifying network training and increasing network depth and complexity. In pose estimation, a residual module can enhance the network's ability to model complex poses. A supervised module is a module used for supervised learning in a network and typically includes a loss function and an optimizer. The supervision module optimizes network performance by calculating the difference (loss) between the network output and the target and updating the network parameters using backpropagation. In pose estimation networks, the supervision module ensures that the network can accurately predict the positions of key points on the human body. A convolution operation involves applying a convolution kernel to the input data and performing a sliding calculation. The convolution kernel slides pixel by pixel across the input data, calculating the dot product between the convolution kernel and the local region of the input data to generate a new feature map.

[0158] First, the system inputs the labeled worker image regions into the backbone network. The backbone network extracts the image's feature map through multiple layers of convolution and pooling operations. Convolution is used to extract local features, while pooling reduces the spatial dimension of the feature map, reducing computational effort while retaining key information. This results in a semantically meaningful feature map. Secondly, the resulting feature map is input into the pose estimation network. This network uses a sampling module to adjust the feature map's resolution, a skip connection module to transfer low-level detail features to higher levels, a residual module to enhance the network's training capabilities, and a supervision module to optimize network performance through a loss function. These modules work together to enable the network to accurately extract key point features. Finally, a convolution operation is performed on the extracted key point features to generate a preliminary heat map of the human body key points. Each pixel value in the heat map represents the probability that the location is a key point, providing an important basis for subsequent key point localization and pose recognition.

[0159] Step S32: performing non-maximum suppression on the preliminary heat map to obtain key points and connection information.

[0160] It should be noted that non-maximum suppression (NMS) is a post-processing technique used to remove redundant detection results and retain the most significant feature points. In pose estimation, non-maximum suppression compares the pixel values of each pixel in the heat map with those in its neighborhood, retaining only the local maximum points, thereby eliminating duplicate or close key point detections and ensuring that each key point is detected only once. This process can significantly improve the accuracy and robustness of key point detection. Non-maximum suppression retention box conditions:

[0161] IoU(B i , B j )<τ

[0162] Among them, τ is the threshold (usually set to 0.5). If the IoU of two boxes exceeds this value, the box with low confidence is suppressed. IoU is the intersection over union ratio, which calculates the overlap ratio of two boxes.

[0163] Key points refer to the important parts used to describe human posture in human posture estimation. These points are the core of posture estimation. By detecting the position and relative relationship of these key points, the system can accurately infer the posture and movement of the human body. After non-maximum suppression, the position of the key points is more accurate, reducing false detection and redundant detection.

[0164] Connectivity information refers to the relative positional relationships between key points and is used to describe the human skeletal structure. For example, the connection between the shoulder and elbow represents an arm, and the connection between the hip and knee represents a leg. By determining the connectivity information of key points, the system can reconstruct the human skeletal structure, thereby more intuitively representing the human posture.

[0165] It can be understood that, first, the system performs local maximum detection on each pixel in the preliminary heat map, compares the pixel value of each pixel with that in its neighborhood, and retains only the local maximum points, thereby removing redundant detection results and avoiding repeated detections caused by similar high-probability points in the heat map. Secondly, the system extracts the precise position information of the key points of the human body based on the retained local maximum points. Finally, the system constructs the connection information between the key points based on the preset relationship between the key points (such as the connection between the shoulder and the elbow represents the arm) to form the skeletal structure of the human body. This process not only improves the accuracy of key point detection, but also provides clear structured information for subsequent posture reconstruction and behavior analysis, allowing the system to understand human posture and movements more intuitively.

[0166] As an example, the step of performing non-maximum suppression on the preliminary heat map to obtain key points and connection information includes: performing local maximum search on the preliminary heat map to obtain local maximum points; filtering the points to be deleted from the local maximum points to obtain key points, and the confidence of the points to be deleted is less than a preset confidence threshold; and determining the connection information of the key points based on the part affinity field.

[0167] Local maximum search involves comparing the values of each pixel in the initial heatmap with those of its neighbors, looking for the local maximum. This process scans the heatmap pixel by pixel using a sliding window, ensuring that each pixel's value is greater than all the values of its surrounding neighbors. The goal of local maximum search is to extract possible keypoint candidates from the heatmap.

[0168] A local maximum point refers to a point in the heat map whose pixel value is greater than all pixel values in the surrounding neighborhood. These points are usually candidate locations of key points because they have the highest confidence in the local area, indicating that the location is likely to be the actual location of a key point.

[0169] The points to be deleted are those with low confidence among the local maximum points. These points may be mistakenly identified as key point candidates due to noise, false detection or other factors. By setting a confidence threshold, these low-confidence points can be filtered out, thereby reducing false detection and redundant detection.

[0170] Confidence refers to the probability value of each pixel in the heat map indicating that the location is a key point. It reflects the model's "confidence" in whether the location is a key point. The higher the confidence, the more likely the location is to be a true key point.

[0171] The preset confidence threshold is a pre-set value used to filter out local maximum points with low confidence. Only points with confidence higher than or equal to the threshold will be retained as key points, while points with confidence lower than the threshold will be marked as points to be deleted and eventually filtered out. The setting of this threshold can be adjusted according to the specific application scenario and accuracy requirements.

[0172] PAFs (Part Affinity Fields) are a feature field used to describe the connection relationship between key points. It calculates the vector field between key points to represent the connection direction and strength between key points. Part affinity fields are used to determine the connection relationship between key points and reconstruct the human skeletal structure. For example, part affinity fields can be used to determine the connection relationship between the shoulder and elbow, thereby indicating the existence of the arm. PAFs formula:

[0173]

[0174] Where N is the number of keypoint pairs, d i is the distance between the i-th pair of key points, and σ is the standard deviation. PAFs represents the connection direction and strength between key points by calculating the vector field between key points.

[0175] PAFs loss calculation formula:

[0176]

[0177] Where c is the limb category, p is the pixel position in the image, and F c (p) is the predicted PAF vector (indicating limb direction and association strength), F c * (p) is the true PAF vector, generated from the labeled data.

[0178] First, the system searches for the local maximum of each pixel in the preliminary heat map. By setting a neighborhood range (such as a 3×3 window), it compares the values of the central pixel with those of the pixels around it. If the central pixel value is higher than the values of all the surrounding pixels, the point is considered a local maximum point. This process can effectively screen out possible key point candidate locations and avoid repeated detections caused by the spread of high-confidence areas in the heat map. Secondly, the system filters out points with confidence levels lower than a preset confidence threshold from these local maximum points, as they may be caused by noise or false detections. By setting a reasonable confidence threshold (such as 0.5 or 0.7), only points with confidence levels higher than the threshold are retained as final key points, thereby improving the accuracy and reliability of key point detection. Finally, the system determines the connection information of the key points based on the part affinity field. The part affinity field represents the connection direction and strength by calculating the vector field between the key points. Specifically, for each pair of key points (such as the shoulder and elbow), the system searches for the vector field related to these two key points in the heat map. The direction and strength of the vector field reflect the connection relationship between the key points. For example, the direction of the vector field points from the shoulder to the elbow, and the intensity is high, indicating that there is a connection relationship between the two key points. By analyzing these vector fields, the system connects the key points according to the human skeletal structure, thereby reconstructing the complete human posture. This process can not only clearly represent the structure of the human body, but also provide important structured information for subsequent behavioral analysis.

[0179] Step S33: Connect the key points according to the connection information to obtain a posture estimation result.

[0180] It should be noted that the pose estimation result refers to the representation of the human body pose obtained by connecting key points. In pose estimation, by connecting key points according to the connection information, a skeleton structure representing the human body pose can be obtained. This skeleton structure is the pose estimation result. The pose estimation result is usually output in the form of images or data, which can intuitively display the human body's pose and movement, providing a basis for subsequent behavior analysis and movement understanding. Key point connection algorithm:

[0181]

[0182] Among them, C ij is the matching cost matrix, element C ij represents the distance between the i-th candidate key point and the j-th true key point, x ij is a binary variable, if x ij =1 means assigning candidate point i to real point j, with the constraint that each candidate point matches at most one real point, and vice versa.

[0183] It can be understood that first, the system identifies which key points have connection relationships based on predetermined connection information, such as recognizing that the connection between the shoulder and elbow represents the arm, and the connection between the hip and knee represents the leg, etc. Then, the system connects the key points in sequence according to these connection relationships to form a skeleton structure representing the human posture. This skeleton structure can clearly display the relative position and posture information of various parts of the human body. Finally, the system outputs this skeleton structure as the posture estimation result, which can be displayed visually on images or videos, helping the system to intuitively understand the posture and movements of the staff, while providing support for accurate behavior recognition and safety management.

[0184] As an example, the step of connecting the key points according to the connection information to obtain a posture estimation result includes: obtaining a key point-connection score table according to the connection information and the key points; selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table; adding the key point connection to the human body posture set; returning to the step of selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table until the key point-connection score table is empty; and obtaining a posture estimation result according to the key point connection construction in the human body posture set.

[0185] The keypoint-connectivity score table is a data structure that stores the connectivity between each pair of keypoints and their corresponding confidence scores. In pose estimation, each keypoint may have a connectivity relationship, and the connectivity score table assigns a score to each pair of potentially connected keypoints by calculating the confidence of these connectivity relationships (usually based on part affinity fields or other features). A higher score indicates a more reliable connectivity between the keypoints. For example, if the connectivity score between the shoulder and elbow is 0.9, while the connectivity score between the hip and knee is 0.85, then the shoulder and elbow connectivity is considered more reliable.

[0186] Keypoint connectivity refers to connecting two keypoints with a line segment in pose estimation to represent the body part between them. For example, the connection between the shoulder and elbow represents the arm, and the connection between the hip and knee represents the leg. Keypoint connectivity not only describes the spatial relationship between keypoints but also provides structured information for reconstructing human pose. By connecting all keypoints according to the connectivity relationship, a complete human skeleton structure can be formed.

[0187] The human pose set stores all keypoint connections, representing the pose of the entire human body. During the pose estimation process, the system gradually adds each pair of keypoint connections to this set, ultimately forming a complete human pose. Each keypoint connection in the human pose set represents a human part, such as an arm, leg, or torso. By combining all keypoint connections, the complete human pose can be reconstructed and output as a skeleton graph, providing a foundation for subsequent behavioral analysis and motion understanding.

[0188] First, based on the connection information and the positional relationships between keypoints, the system calculates a confidence score between each pair of keypoints. (Specifically, this is done by analyzing the part affinity field to determine the direction and strength of the connection between keypoints, and quantifying this direction and strength information into scores.) These scores are then recorded in a keypoint-connection score table. This quantifies the reliability of the connection between each pair of keypoints and provides a basis for subsequent connection selection. Second, the system selects the keypoint pair with the highest score from the score table as the keypoint connection, as the highest-scoring connection is most likely to be a true human part connection. This pair of keypoints is then removed from the score table to avoid repeated selection and ensure that each connection is processed only once. The system then adds the selected keypoint connections to the human pose set, gradually constructing a skeleton structure for the human pose. This process, through iterative selection and deletion operations, gradually refines the pose estimation results. Finally, when the keypoint-connection score table is empty, indicating that all possible connection relationships have been processed, the system constructs a complete human pose estimation result based on the keypoint connections in the human pose set. This intuitive representation of the human pose is presented in the form of a skeleton graph, providing accurate structured information for subsequent behavioral analysis, thereby achieving efficient and accurate pose estimation.

[0189] Step S34 , tracking the worker's motion trajectory according to the posture estimation results of the continuous video frames to obtain a motion trajectory map.

[0190] It's important to note that a motion trajectory graph is a graph that analyzes the worker's posture estimation results from consecutive video frames, plotting the movement of key body points over time. It visually illustrates the worker's movement changes and paths over time, typically presented as lines or curves within an image or video. The motion trajectory graph intuitively reflects the worker's movement patterns, speed, and direction, providing a basis for subsequent behavioral analysis.

[0191] As you can understand, the system first obtains pose estimation results for each frame from consecutive video frames. These results include the position information of key points on the human body. The system then tracks the position of the same key point in different frames, calculating its displacement and direction of motion in consecutive frames to plot its motion trajectory over time. The system then integrates the motion trajectories of all key points to form a complete motion trajectory diagram, which intuitively displays the worker's overall motion pattern and movement path.

[0192] Step S35: input the motion trajectory diagram into a behavior recognition model to obtain the operator's operating behavior. The behavior recognition model is obtained by training a three-dimensional convolutional network.

[0193] It should be noted that an action recognition model is a machine learning model used to identify and classify human motion behaviors. It analyzes input motion data (such as motion trajectory images) to determine the specific operational behaviors being performed by a worker. The action recognition model is obtained by training a 3D Convolutional Neural Network (CNN). First, a large amount of video data or motion trajectory images labeled with operational behaviors are collected. These data serve as training samples, providing a foundation for the model to learn. This data is then fed into the 3D convolutional network. The network uses convolutional and pooling layers to extract the spatiotemporal characteristics of the actions, capturing how they vary in time and space. Next, using the labeled action categories as supervisory signals, a loss function is used to calculate the difference between the predicted results and the true labels. The backpropagation algorithm is then used to update the network parameters, gradually optimizing the model's performance. After multiple iterations of training, the network learns the characteristic patterns of different operational behaviors, ultimately forming a behavior recognition model that can accurately identify them.

[0194] Operational behavior refers to the various actions or activities performed by workers in specific scenarios. In industrial scenarios, operational behavior may include normal workflows (such as operating equipment, moving items) or abnormal behaviors (such as illegal operations, entering dangerous areas). By analyzing the action trajectory graph through the behavior recognition model, these behaviors can be classified and identified for monitoring and management.

[0195] 3D CNN is a deep learning architecture specifically designed to process data with a temporal dimension (such as videos or motion trajectory graphs). Unlike two-dimensional convolutional networks (2D CNNs), 3D convolutional networks not only consider spatial features (such as pixel relationships in an image), but also introduce a temporal dimension, which can capture changes in motion over time.

[0196] It can be understood that first, the motion trajectory map generated from continuous video frames is input into the behavior recognition model. The motion trajectory map contains the movement information of the key points of the staff over time, which can reflect their motion patterns and behavioral characteristics. Next, the behavior recognition model uses its internal three-dimensional convolutional network structure to process the motion trajectory map. The three-dimensional convolutional network captures the temporal and spatial characteristics of the motion trajectory through the convolution layer, and extracts key feature information that can characterize different operational behaviors. Finally, based on the extracted feature information and the behavioral patterns learned during the training process, the model outputs the specific operational behavior category of the staff, such as normal operation, illegal operation, or other specific behaviors, thereby achieving accurate identification and classification of the staff's behavior, providing an important basis for monitoring and management.

[0197] In this embodiment, after labeling is complete, the system first performs pose estimation on the image region corresponding to the worker, generating a preliminary heat map of key points on the human body. This process quickly locates the possible locations of key points, providing basic data for subsequent analysis. Next, non-maximum suppression is performed on the preliminary heat map to remove redundant detections, retain the key points with the highest confidence, and determine the connectivity between key points, thereby improving the accuracy and reliability of key point detection. Key points are then connected based on this connectivity information to construct a complete human pose estimation result, forming a clear human skeleton structure and providing intuitive structured information for behavioral analysis. Subsequently, based on the pose estimation results from consecutive video frames, the worker's motion trajectory is tracked to generate a motion trajectory graph. This process captures dynamic behavioral changes and provides temporal information for behavior recognition. Finally, the motion trajectory graph is input into a trained behavior recognition model to identify the worker's specific operational behaviors. By learning the spatiotemporal characteristics of the movements, the model can accurately distinguish between normal and abnormal behaviors, thereby enabling real-time monitoring and early warning of worker behavior. The entire process achieves efficient conversion from image data to behavior recognition, significantly improving the intelligence level of the monitoring system and enhancing the proactive and effective safety management.

[0198] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the full-site intelligent video control method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0199] This application also provides a full-station smart video control device, please refer to Figure 4 The whole-station intelligent video control device includes:

[0200] An image recognition module 10 is configured to recognize equipment, staff, and environmental areas in a video frame based on directional gradient features of each area in the video frame;

[0201] a labeling module 20, configured to label the device, the staff, and the environmental area in the video frame;

[0202] An operation behavior recognition module 30 is used to determine the operation behavior of the staff member based on the posture of the staff member in the continuous video frames when the labeling is completed;

[0203] The abnormal behavior recognition module 40 is used to recognize the operation behavior through behavior pattern coding and obtain a recognition result;

[0204] The exception handling module 50 is used to send a warning message to a preset device and mark the operation behavior when the recognition result is that the operation is abnormal.

[0205] The full-station smart video control device provided in this application, which adopts the full-station smart video control method in the above-mentioned embodiment, can solve the technical problem of how to achieve real-time and accurate identification of dynamic operational behaviors and abnormal warning in industrial scenarios. Compared with the existing technology, the beneficial effects of the full-station smart video control device provided in this application are the same as the beneficial effects of the full-station smart video control method provided in the above-mentioned embodiment, and the other technical features of the full-station smart video control device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0206] The present application provides a site-wide smart video control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the site-wide smart video control method in the above-mentioned embodiment one.

[0207] Reference below Figure 5 , which shows a schematic diagram of the structure of a full-station smart video control device suitable for implementing the embodiments of the present application. The full-station smart video control device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The site-wide smart video control device shown is merely an example and should not limit the functions and scope of use of the embodiments of this application.

[0208] like Figure 5As shown, the whole-station smart video control device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to RAM (Random Access Memory) 1004. Various programs and data required for the operation of the whole-station smart video control device are also stored in RAM 1004. The processing device 1001, ROM 1002 and RAM 1004 are connected to each other via a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, an LCD (Liquid Crystal Display), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 can allow the full-station intelligent video control device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows a full-station intelligent video control device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.

[0209] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0210] The full-station smart video control device provided in this application, which employs the full-station smart video control method of the above-mentioned embodiment, can solve the technical problem of how to achieve real-time and accurate identification of dynamic operational behaviors and abnormal warnings in industrial scenarios. Compared with the prior art, the beneficial effects of the full-station smart video control device provided in this application are the same as those of the full-station smart video control method provided in the above-mentioned embodiment, and the other technical features of the full-station smart video control device are the same as those disclosed in the method of the previous embodiment, and are not further described here.

[0211] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0212] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0213] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the full-site intelligent video management method in the above-mentioned embodiment.

[0214] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash memory), optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0215] The above-mentioned computer-readable storage medium may be included in the whole-station smart video control device; or it may exist independently without being assembled into the whole-station smart video control device.

[0216] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the whole-station intelligent video control device, the whole-station intelligent video control device: identifies the equipment, staff and environmental areas in the video frame according to the directional gradient characteristics of each area in the video frame; marks the equipment, staff and environmental areas in the video frame; when the marking is completed, determines the operating behavior of the staff according to the posture of the staff in the continuous video frames; identifies the operating behavior through behavior pattern coding to obtain an identification result; when the identification result is an abnormal behavior, sends a warning message to the preset device and marks the operating behavior.

[0217] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0218] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0219] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0220] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned full-station intelligent video control method, and can solve the technical problem of how to achieve real-time and accurate identification of dynamic operating behaviors and abnormal warnings in industrial scenarios. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the full-station intelligent video control method provided in the above-mentioned embodiment, and will not be repeated here.

[0221] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned full-station intelligent video management method.

[0222] The computer program product provided in this application can solve the technical problem of achieving real-time, accurate identification of dynamic operational behaviors and abnormal warnings in industrial scenarios. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the full-station intelligent video control method provided in the above embodiment, and will not be elaborated here.

[0223] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A site-wide intelligent video control method, characterized in that: The method comprises: Identifying equipment, staff, and environmental areas in the video frame based on directional gradient features of each area in the video frame; marking the device, the staff, and the environmental area in the video frame; When the labeling is completed, determining the operator's operation behavior according to the operator's posture in the continuous video frames; Identify the operation behavior through behavior pattern coding to obtain an identification result; When the recognition result is that the operation is abnormal, a warning message is sent to a preset device and the operation behavior is marked.

2. The method according to claim 1, wherein When the labeling is completed, the step of determining the operator's operating behavior according to the operator's posture in the continuous video frames includes: When the annotation is completed, the posture of the image area corresponding to the worker is estimated to obtain a preliminary heat map of the key points of the human body; Performing non-maximum suppression on the preliminary heat map to obtain key points and connection information; Connecting the key points according to the connection information to obtain a posture estimation result; Tracking the worker's motion trajectory according to the posture estimation results of continuous video frames to obtain a motion trajectory map; The motion trajectory diagram is input into a behavior recognition model to obtain the operating behavior of the staff member, and the behavior recognition model is obtained by training a three-dimensional convolutional network.

3. The method according to claim 2, wherein The step of performing non-maximum suppression on the preliminary heat map to obtain key points and connection information includes: Performing a local maximum search on the preliminary heat map to obtain a local maximum point; Filtering the points to be deleted from the local maximum points to obtain key points, wherein the confidence of the points to be deleted is less than a preset confidence threshold; The connection information of the key points is determined according to the part affinity field.

4. The method according to claim 2, wherein The step of connecting the key points according to the connection information to obtain a posture estimation result includes: Obtain a key point-connection score table according to the connection information and the key points; Selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table; Adding the key point connection to a human body posture set; Returning to the step of selecting a pair of key points with the highest score from the key point-connection score table as a key point connection, and deleting the pair of key points with the highest score from the key point-connection score table until the key point-connection score table is empty; The posture estimation result is obtained by connecting and constructing the key points in the human posture set.

5. The method according to claim 2, wherein When the labeling is completed, the steps of performing posture estimation on the image area corresponding to the worker to obtain a preliminary heat map of key points of the human body include: After the annotation is completed, the image area corresponding to the worker is input into the backbone network to obtain a feature map. The backbone network is composed of multiple convolutional networks and pooling networks. Inputting the feature map into a posture estimation network to obtain key point features of the human body, wherein the posture estimation network is composed of a sampling module, a skip connection module, a residual module and a supervision module; A convolution operation is performed on the key point features of the human body to obtain a preliminary heat map of the key points of the human body.

6. The method according to claim 1, wherein The step of identifying the equipment, staff, and environment areas in the video frame according to the directional gradient features of each area in the video frame includes: Converting the video frame into a grayscale image, and performing gamma correction on the grayscale image to obtain a gamma-corrected image; Calculating the gradient magnitude and direction of each pixel in the gamma-corrected image to obtain a gradient map; Dividing the gradient map into multiple connected regions, and summing the gradient amplitudes of the pixels within the connected regions to form a gradient direction histogram; Normalizing the gradient direction histogram to obtain a feature vector; The feature vector is compared with a feature database to determine the equipment, staff, and environmental areas in the video frame.

7. The method according to claim 1, wherein The step of identifying the operation behavior through behavior pattern coding to obtain an identification result includes: Performing behavior pattern coding on the operation behavior to obtain a behavior pattern coding result; Compare the behavior pattern encoding result with each pattern in the preset abnormal behavior pattern library and calculate the similarity score; The recognition result is obtained according to the similarity score and a preset abnormal behavior score threshold.

8. The method according to claim 1, wherein After the step of sending a warning message to a preset device and marking the operation behavior when the operation behavior is abnormal, the method further includes: Creating an index table based on the annotated content in the video frame and the position of the video frame corresponding to the annotated content in the surveillance video; When a search instruction is received, obtaining a search target according to the search instruction; The index table is searched according to the search target to obtain a surveillance video segment corresponding to the search target.

9. The method according to any one of claims 1 to 8, characterized in that After the step of sending a warning message to a preset device and marking the operation behavior when the operation behavior is abnormal, the method further includes: Obtain historical surveillance videos; Analyze the historical surveillance video according to the annotations in the historical surveillance video, and generate and save a video summary; When there is no abnormal behavior annotation in the annotation and no abnormal event in the video summary, the historical surveillance video is deleted.

10. A station-wide intelligent video control device, characterized in that: The device comprises: An image recognition module is used to identify equipment, staff, and environmental areas in the video frame based on directional gradient features of each area in the video frame; a labeling module, configured to label the device, the staff, and the environmental area in the video frame; An operation behavior recognition module is used to determine the operation behavior of the staff member based on the posture of the staff member in the continuous video frames when the labeling is completed; An abnormal behavior recognition module is used to recognize the operation behavior through behavior pattern coding and obtain a recognition result; The exception handling module is used to send a warning message to a preset device and mark the operation behavior when the recognition result is an abnormal behavior.

Citation Information

Cited By

  • AI event compression transmission system for hydraulic engineering video monitoring

    CN121887960A