Dynamic graph modal semantic segmentation method and system for vehicle-mounted mining video monitoring

By adopting dynamic graph modal semantic segmentation method in the vehicle-mounted video monitoring system, combined with convolutional neural network and context information fusion technology, the challenges of complex environment and big data processing in mining area video monitoring are solved, and efficient mining area video monitoring and analysis are achieved.

CN119625323BActive Publication Date: 2025-05-16成都科瑞特电气自动化有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510157519.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The existing vehicle-mounted video monitoring methods face challenges such as light changes, occlusions, variable terrain and huge data volume when dealing with complex mining environments, making it difficult to achieve efficient mining area video monitoring.

Method used

The dynamic graph modal semantic segmentation method is used to obtain video information through the on-board camera, and the convolutional neural network is used to extract features, fuse dynamic graph modalities of time and spatial dimensions, and combine global and local context information to achieve efficient semantic segmentation of mining area videos.

Benefits of technology

It improves the identification accuracy and processing speed of mining area video monitoring, and can achieve real-time monitoring while ensuring identification accuracy, enhancing the understanding and analysis ability of mining area environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625323B_ABST
    Figure CN119625323B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision, and provides a method and system for semantic segmentation of dynamic graph modality for vehicle-mounted mining video monitoring. The present invention captures mining area videos through a vehicle-mounted camera, uses a pre-trained first convolutional neural network to deeply extract video frame features and fuse multiple layers of features to enhance representation capabilities, and then constructs a dynamic graph in the time dimension and a dynamic graph in the space dimension. After fusion, a dynamic graph modality is formed, and an attention mechanism is introduced to assign weights based on the similarity of feature vectors, focusing on the correlation between key areas. The global average pooling layer is used to extract context information, and the global and dynamic graph information are combined with the weighted original feature map to fuse. In addition, local context information is also combined to enhance feature understanding. Based on the fully fused feature map, an accurate semantic segmentation result is output, and a label representing the object or background category to which each pixel belongs is assigned, achieving a deep understanding and efficient analysis of mining area video information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and more specifically, to a method and system for dynamic graph modal semantic segmentation of vehicle-mounted mining video monitoring. Background Art

[0002] The contents of this section merely provide background information related to the present application and may not constitute prior art.

[0003] In the process of mining mineral resources, real-time monitoring of mining areas is a key link to ensure safe production and improve mining efficiency. Traditionally, monitoring of mining areas relies on manual inspections, which is not only time-consuming and labor-intensive, but also difficult to achieve comprehensive and real-time coverage of mining areas. With the rapid development of computer vision and deep learning technologies, vehicle-mounted video monitoring systems have gradually become an important means of mining area monitoring. However, existing vehicle-mounted video monitoring methods still face many challenges when dealing with complex mining environments.

[0004] First, the mining environment is complex and changeable, including different lighting conditions, obstructions, and changing terrain, all of which place high demands on the clarity and accuracy of the video. Secondly, there are many kinds of objects in the mining area, such as ore piles, equipment, personnel, etc. These objects vary in shape, size, and color, making it difficult for traditional image processing methods to accurately identify and classify them. In addition, the monitoring videos of mining areas often have a huge amount of data. How to efficiently process and analyze this data and extract useful information is another problem that needs to be solved.

[0005] In order to solve the above problems, researchers have begun to explore the use of deep learning technology for video monitoring in mining areas in the existing technology. For example, convolutional neural networks (CNNs) perform well in image feature extraction and classification, can automatically learn complex features in images, and accurately identify different objects. However, there are still deficiencies in directly applying existing deep learning models to video monitoring in mining areas. On the one hand, the video information in mining areas has rich spatiotemporal characteristics, that is, there is temporal continuity between video frames, and different areas within the frames contain different objects and backgrounds. How to effectively integrate these spatiotemporal information and improve the recognition accuracy of the model is a key issue. On the other hand, the amount of video data in mining areas is huge. How to improve the processing speed of the model while ensuring recognition accuracy is also an important challenge to achieve real-time monitoring. Summary of the invention

[0006] In order to solve the above technical problems, the purpose of this application is to provide a dynamic graph modal semantic segmentation method and system for vehicle-mounted mining video monitoring. By fusing the dynamic graph modalities of the time dimension and the space dimension, combining the global and local context information, and utilizing convolutional neural networks to perform feature extraction and semantic segmentation on mining area videos, feature representation can be enhanced, and the objects or background categories in the video can be accurately identified, thereby achieving efficient mining area video monitoring.

[0007] The purpose of this application is achieved through the following technical solutions:

[0008] In a first aspect, the present invention provides a dynamic graph modality semantic segmentation method for vehicle-mounted mining video monitoring, comprising:

[0009] Obtain mining area video information through vehicle-mounted cameras;

[0010] The pre-trained first convolutional neural network model is used to extract features from video frames, and the features of different convolutional layers are fused to obtain a feature set with richer feature representation;

[0011] By using the temporal continuity between video frames, the pixel differences between adjacent frames are calculated to obtain the inter-frame differential image; based on the inter-frame differential image, the motion area in the video is detected; the static background is identified from the initial frame or multiple consecutive frames of the mining area video information sequence through the visual background extractor, and the motion area is combined with the static background based on image superposition to construct a dynamic map in the time dimension;

[0012] The video frame is divided into multiple regions, each of which corresponds to a specific object or background; the pre-trained second convolutional neural network model is used to extract features from each region to obtain the feature vector of the region; based on the feature vector of the region, a dynamic graph of the spatial dimension is constructed;

[0013] The dynamic graphs of time dimension and space dimension are integrated to form a dynamic graph modality;

[0014] After feature extraction, for the feature vector of each region, based on the dynamic graph modality, the similarity between the feature vector of each region and the feature vector of other regions is calculated, and attention weights are assigned based on the similarity, so that the network pays attention to other regions that are most relevant to the current region to enhance feature representation;

[0015] The global context information is extracted by using the global average pooling layer. The global context information is matched with the weighted original feature map containing the dynamic graph modal information through upsampling. The two feature maps with the same spatial resolution are added or multiplied element by element at the corresponding pixel positions to realize the combination of the global context information and the dynamic graph modal information.

[0016] According to each area in the video frame or the local area around a specific pixel, the feature representation of these local areas is obtained, and the local context information is combined with the global context information by directly superimposing or weighted fusion features;

[0017] Based on the feature map that integrates global context information, dynamic graph modal information and local context information, the semantic segmentation result is output. The semantic segmentation result is a label map containing multiple pixels, where the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

[0018] Furthermore, the steps of extracting features from video frames using the pre-trained first convolutional neural network model and fusing features from different convolutional layers specifically include:

[0019] A first convolutional neural network model that has been pre-trained on a large image dataset is selected, where the first convolutional neural network model has multiple convolutional layers, and each convolutional layer can extract image features at different levels;

[0020] Input each frame of the mining area video information into the first convolutional neural network model, and extract the output feature map of each convolutional layer respectively;

[0021] Using the feature pyramid network structure, the feature maps of different convolutional layers are upsampled or downsampled to make the feature maps have the same spatial resolution;

[0022] The feature maps with the same spatial resolution are spliced ​​in the channel dimension to form a fused feature set, which contains a variety of feature information from low to high levels.

[0023] Furthermore, the dynamic graphs of the time dimension and the space dimension are fused to form a dynamic graph modality, which specifically includes:

[0024] Feature stitching is used to combine the dynamic change map obtained in the time dimension based on inter-frame difference and background subtraction with the static feature map obtained in the space dimension based on region division and feature extraction according to preset weights; among them, the dynamic map in the time dimension captures the motion trajectory and state changes of objects in the mining scene, and the dynamic map in the space dimension reflects the inherent attributes and relative position relationships of different areas in the scene.

[0025] Furthermore, the steps of calculating the similarity between each region feature vector and other region feature vectors and allocating attention weights based on the similarity specifically include:

[0026] The cosine similarity formula is used for measurement. The cosine similarity formula is:

[0027]

[0028] in, and Represent the feature vectors of two different regions, represents the cosine similarity between these two feature vectors;

[0029] Based on the similarity obtained by the cosine similarity formula, the attention weight is assigned, and the weight calculation formula is:

[0030]

[0031] in, Indicates Region to The attention weight of the region, and Respectively and The feature vector of the region; is the number of eigenvectors.

[0032] Furthermore, after outputting the semantic segmentation result, it also includes:

[0033] The label image is expanded to obtain an expanded label image; the label image is then corroded to obtain an eroded label image; the difference area between the expanded label image and the eroded label image is calculated to obtain an edge-enhanced label image of the object.

[0034] Furthermore, after outputting the semantic segmentation result, it also includes:

[0035] Each pixel in the semantic segmentation result is regarded as a node in the graph, and the connection between nodes is defined based on spatial proximity and color similarity;

[0036] By constructing a minimum spanning tree, the average feature distance between each node and the surrounding nodes is calculated. If the difference between the average feature distance of any node and the average value of its category exceeds a threshold, the node is reclassified into the category to which its nearest category node belongs to ensure consistency within the segmentation area.

[0037] Furthermore, after outputting the semantic segmentation result, it also includes:

[0038] Based on the semantic segmentation results, statistics are collected on specific objects or background categories in the mine video information to generate a report containing information on the number of objects, location distribution, and motion status, in order to facilitate mine management and safety monitoring.

[0039] In a second aspect, the present invention provides a dynamic graph modal semantic segmentation system for vehicle-mounted mining video monitoring, comprising:

[0040] A video information acquisition module is used to acquire video information of the mining area through a vehicle-mounted camera;

[0041] A feature extraction module is used to extract features from video frames using a pre-trained first convolutional neural network model, and to fuse features from different convolutional layers to obtain a feature set with richer feature representation;

[0042] The temporal dynamic map acquisition module is used to calculate the pixel difference between adjacent frames using the temporal continuity between video frames to obtain the inter-frame differential image; detect the motion area in the video based on the inter-frame differential image; identify the static background from the initial frame or multiple consecutive frames of the mining area video information sequence through the visual background extractor, combine the motion area with the static background based on image superposition, and construct a dynamic map in the time dimension;

[0043] The spatial dynamic map acquisition module is used to divide the video frame into multiple regions, each region corresponds to a specific object or background; use the pre-trained second convolutional neural network model to extract features from each region to obtain the feature vector of the region; and construct a dynamic map of the spatial dimension based on the feature vector of the region;

[0044] The fusion module is used to fuse the dynamic graphs of the time dimension and the space dimension to form a dynamic graph modality;

[0045] The attention allocation module is used to calculate the similarity between the feature vector of each region and the feature vector of other regions based on the dynamic graph mode after feature extraction, and allocate attention weights based on the similarity, so that the network pays attention to other regions that are most relevant to the current region to enhance feature representation;

[0046] The global fusion module is used to extract global context information using the global average pooling layer, match the global context information with the weighted original feature map containing dynamic graph modal information through upsampling, and perform element-by-element addition or element-by-element multiplication of two feature maps with the same spatial resolution at corresponding pixel positions to achieve the combination of global context information and dynamic graph modal information;

[0047] A local fusion module is used to obtain feature representations of each region in the video frame or the local regions around a specific pixel, and combine the local context information with the global context information by directly superimposing or weighted fusion features;

[0048] The output module outputs the semantic segmentation result based on the feature map that integrates global context information, dynamic graph modal information and local context information. The semantic segmentation result is a label map containing multiple pixels, where the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

[0049] In a third aspect, the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps corresponding to the method in the first aspect are implemented.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps corresponding to the method in the first aspect.

[0051] In summary, the technical solution of the embodiment of the present application has at least the following advantages and beneficial effects:

[0052] The present invention captures mining area videos through a vehicle-mounted camera, uses a pre-trained first convolutional neural network to deeply extract video frame features and fuse multiple layers of features to enhance representation capabilities, and combines inter-frame difference and visual background extraction technology to construct a time-dimensional dynamic map containing motion information and static background. Subsequently, the video frame is partitioned and the second convolutional neural network is applied to extract features of each region to construct a spatial dimension dynamic map. The two are fused to form a dynamic map modality, and an attention mechanism is introduced to assign weights based on the similarity of feature vectors, so that the model focuses on the correlation between key areas. The global average pooling layer is used to extract context information, which is combined with the weighted original feature map, and the global and dynamic map information is fused through element-level operations. In addition, local context information is also combined to enhance feature understanding. Finally, based on the fully fused feature map, the method outputs accurate semantic segmentation results, assigns labels to each pixel representing the object or background category to which it belongs, and achieves a deep understanding and efficient analysis of mining area video information. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flow chart of a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring provided by the present invention;

[0054] Figure 2 A structural schematic diagram of a dynamic graph modal semantic segmentation system for vehicle-mounted mining video monitoring provided by the present invention;

[0055] Figure 3 The present invention provides a schematic structural diagram of an electronic device. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0057] like Figure 1 As shown, a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring proposed in an embodiment of the present application includes:

[0058] S101, obtaining video information of the mining area through a vehicle-mounted camera.

[0059] Specifically, vehicle-mounted cameras are installed on mining vehicles or other suitable locations to ensure comprehensive and real-time monitoring of various activities in the mining area. These cameras have high resolution and wide-angle field of view, and can capture key information such as vehicle driving, personnel movement, and equipment operation in the mining area. In the process of obtaining video information, the camera not only records the static scenes of the mining area, such as road layout, building location, etc., but more importantly, it captures dynamic changes, such as the movement trajectory of vehicles and the working status of personnel.

[0060] S102, using the pre-trained first convolutional neural network model to extract features from the video frame, and fusing features from different convolutional layers to obtain a feature set with richer feature representation.

[0061] Among them, the pre-trained first convolutional neural network model is used to perform in-depth feature extraction and analysis on each frame of the mining area video information, aiming to extract rich and multi-level feature information from the complex mining area scene. This step is a key link in the entire processing flow, which is directly related to the accuracy of subsequent analysis and understanding. Specifically, a first convolutional neural network model that has been pre-trained on a large image dataset was first selected. This model has multiple convolutional layers, each of which can extract features at different levels in the image, from low-level edges and texture information to high-level complex features such as shapes and patterns. These features together constitute a comprehensive description of the image content.

[0062] Next, each frame of the mining area video information is input into the pre-trained convolutional neural network model one by one. The model processes each frame layer by layer, starting from the first convolution layer, and extracts feature maps layer by layer. These feature maps contain abstract representations of the image at different levels and are the basis for subsequent analysis.

[0063] In order to make full use of the feature information from different convolutional layers, the system adopts a feature pyramid network structure. This structure upsamples or downsamples the feature maps of different convolutional layers so that they have the same spatial resolution. In this way, feature maps with the same spatial resolution can be spliced ​​in the channel dimension to form a feature set that integrates multiple feature information from low to high layers. This feature set not only contains the detailed information of the image, but also contains higher-level abstract features, providing a solid foundation for subsequent analysis.

[0064] S103, using the temporal continuity between video frames, calculates the pixel differences between adjacent frames to obtain an inter-frame differential image; detects the motion area in the video based on the inter-frame differential image; identifies the static background from the initial frame or multiple consecutive frames of the mining area video information sequence through a visual background extractor, combines the motion area with the static background based on image superposition, and constructs a dynamic map in the time dimension.

[0065] Specifically, the temporal continuity in the mining video information is used to capture and analyze dynamic changes in the video. This step is mainly divided into several key links: First, by calculating the pixel differences between adjacent video frames, the system is able to generate inter-frame difference images. The inter-frame difference image reveals the changes in the pixel level between adjacent frames, which usually correspond to the moving areas in the video. Due to the existence of activities such as vehicle driving and personnel movement in the mining environment, the inter-frame difference image can effectively highlight these dynamically changing parts.

[0066] Next, the system uses the inter-frame difference image to detect the moving area in the video. This process usually involves thresholding the difference image to distinguish the moving area from the static background. By setting the appropriate threshold, the system can accurately separate the moving object from the background, providing key information for subsequent analysis.

[0067] In order to build a more comprehensive dynamic map, a visual background extractor is also used to identify the static background in the video information sequence of the mining area. For example, the ViBe algorithm (Visual Background Extracto) is used to initialize the background model using the first frame of the video stream. This process can extract stable background information from the initial frame or multiple consecutive frames of the video, such as fixed elements such as roads and buildings. Based on this information, the system can combine the previously detected motion areas with the static background to build a composite map containing dynamic and static information in the time dimension. This dynamic map not only reflects the real-time activities in the mining area, but also provides an in-depth understanding of the environmental structure of the mining area. By combining the motion area with the static background, the system can more accurately describe and analyze various dynamic changes in the mining area, providing a strong basis for subsequent advanced analysis and decision support.

[0068] S104, dividing the video frame into a plurality of regions, each region corresponding to a specific object or background; extracting features from each region using a pre-trained second convolutional neural network model to obtain a feature vector of the region; and constructing a dynamic graph of a spatial dimension according to the feature vector of the region;

[0069] Specifically, the processing of mining video information is further refined, aiming to enhance the understanding of mining activities through the analysis of spatial dimensions. First, the video frame is divided into multiple regions. This step is based on the natural segmentation of video content. Each region corresponds to a specific object or background, such as a vehicle, a person, or a fixed mining facility. Such a division helps to decompose complex video scenes into more manageable and analyzable parts.

[0070] Next, in order to further explore the characteristics of each region, a pre-trained second convolutional neural network model was used for feature extraction. This model also uses a convolutional neural network (CNN) to capture a variety of features from edges and textures to more complex shapes and patterns through different levels of convolution kernels. These features include but are not limited to color histograms, edge directions, local binary patterns (LBP), histograms of directional gradients (HOG), etc., as well as high-level abstract features automatically learned by CNN. After this feature extraction process, each region will generate a high-dimensional feature vector, which compactly and comprehensively represents the unique visual attributes of the region.

[0071] Based on these regional feature vectors, a dynamic graph of spatial dimension is constructed to reflect the spatial layout of each region in the video. In addition, the potential connections and interactions between regions are revealed through the similarity measurement of feature vectors. For example, adjacent vehicle areas or personnel activity areas may show close connections in the graph due to similar motion patterns and characteristics. Such a dynamic graph of spatial dimension provides rich contextual information for subsequent analysis and helps to more accurately understand the activity patterns in the mining area.

[0072] S105, the dynamic graphs of the time dimension and the space dimension are fused to form a dynamic graph modality. The dynamic graphs of the time dimension and the space dimension are fused to form a more comprehensive and in-depth dynamic graph modality. This step is a key bridge connecting the temporal dynamic analysis and the understanding of spatial features, aiming to integrate the temporal changes and spatial layout in the mining area video information, and provide strong support for the subsequent complex scene analysis.

[0073] Specifically, the fusion process first relies on feature splicing, that is, combining the dynamic change map obtained based on inter-frame difference and background subtraction in the time dimension with the static feature map obtained based on region division and feature extraction in the space dimension according to the preset weights. The dynamic map in the time dimension mainly captures the motion trajectory and state changes of objects in the mining scene, such as the driving path of vehicles and the movement speed of personnel. The dynamic map in the spatial dimension focuses on reflecting the inherent attributes and relative positional relationships of different areas in the scene, such as the vehicle density and personnel activity intensity in a specific area.

[0074] When splicing features, the complementarity of time and space dimension information is considered, and weight distribution is used to ensure that the two can enhance each other rather than interfere with each other during the fusion process. The output of this step is a comprehensive dynamic graph modality that combines time dynamics and spatial features. It not only contains rich spatiotemporal information, but also has the potential to deeply analyze complex scenes in mining areas.

[0075] S106, after feature extraction, for the feature vector of each region, based on the dynamic graph modality, the similarity between the feature vector of each region and the feature vectors of other regions is calculated, and attention weights are assigned based on the similarity, so that the network pays attention to other regions that are most relevant to the current region to enhance feature representation.

[0076] Specifically, for each region's feature vectors, the similarity between them and other regions' feature vectors is calculated. Similarity can be calculated using a variety of methods, such as cosine similarity, Euclidean distance, etc. These methods can effectively measure the proximity of feature vectors in space, thereby reflecting the correlation between regions. Based on these similarity metrics, the system then assigns attention weights to each region, and the size of the weight is proportional to the similarity, which means that other regions with the most similar features to the current region will receive higher attention.

[0077] The introduction of the attention mechanism enables the network to automatically focus on other areas that are most relevant to the current area. This strategy of focusing on key information helps to enhance feature representation and improve the accuracy of subsequent analysis. Through the allocation of attention weights, the feature vector of each area not only contains its own information, but also incorporates information from other areas that are highly relevant to it, thereby improving the richness and representativeness of the feature vector as a whole.

[0078] In addition, this step also implies an interaction mechanism between regions, that is, different regions achieve information exchange and integration through similarity calculation and weight allocation. This interaction helps to reveal the potential connections between different objects or backgrounds in the mining area video, such as the activity association between vehicles and personnel, and the functional division of different operating areas.

[0079] In other words, by calculating the similarity between regions and assigning attention weights, not only the feature representation is enhanced, but also a more accurate and comprehensive feature basis is provided for subsequent advanced analysis tasks such as semantic segmentation. The specific calculation is as follows:

[0080] The cosine similarity formula is used for measurement. The cosine similarity formula is:

[0081]

[0082] in, and Represent the feature vectors of two different regions, represents the cosine similarity between these two feature vectors;

[0083] Based on the similarity obtained by the cosine similarity formula, the attention weight is assigned, and the weight calculation formula is:

[0084]

[0085] in, Indicates Region to The attention weight of the region, and Respectively and The feature vector of the region; is the number of eigenvectors.

[0086] S107, using the global average pooling layer to extract global context information, matching the global context information with the weighted original feature map containing dynamic graph modal information through upsampling, and adding or multiplying the two feature maps with the same spatial resolution element by element at corresponding pixel positions to realize the combination of global context information and dynamic graph modal information.

[0087] Specifically, step S107 aims to enhance the ability to understand the mining area video information by fusing global context information with dynamic graph modal information. This step first relies on the application of a global average pooling layer, which can extract global context information for the entire feature map. Global context information is crucial for understanding the overall scene layout, the relationships between objects, and potential activity patterns in an image or video. Through global average pooling, each feature channel is compressed into a single numerical value, which together constitutes a low-dimensional representation of the global scene, effectively capturing the global features in the video frame.

[0088] Subsequently, the low-dimensional representation of the global context information is converted to the same spatial resolution as the original weighted feature map containing the dynamic graph modal information using upsampling technology. The upsampling process ensures the consistency of the global context information and the dynamic graph modal information in the spatial dimension, providing a basis for subsequent feature fusion. This step reflects a meticulous analysis strategy from global to local, that is, combining global context information with local dynamic graph modal information to form a comprehensive understanding of the mining area video information.

[0089] In the feature fusion stage, the system uses element-by-element addition or element-by-element multiplication to combine global context information with dynamic graph modal information on feature maps with the same spatial resolution. The element-by-element addition method directly adds the values ​​of the two feature maps at corresponding pixel positions, thereby retaining their respective information and generating new feature combinations. The element-by-element multiplication method emphasizes the common features of the two feature maps and suppresses irrelevant features by calculating the product of the values ​​of the two feature maps at corresponding pixel positions. The choice of these two fusion methods depends on the specific task requirements and dataset characteristics, aiming to maximize the complementarity between global context information and dynamic graph modal information to improve the accuracy of subsequent analysis.

[0090] S108, according to each area in the video frame or the local area around a specific pixel, obtain feature representations of these local areas, and combine the local context information with the global context information by directly superimposing or weighted fusion features.

[0091] Among them, step S108 aims to enhance the in-depth understanding of the video content by combining local context information with global context information. This step begins to focus on the local area features around each area or specific pixel in the video frame after the global context information and dynamic graph modal information have been integrated.

[0092] Specifically, the system first defines a local region for each area or specific pixel in the video frame. This local region is usually a fixed-size window centered on the target area or pixel, which contains the contextual information around the target area or pixel. Then, the previously mentioned pre-trained convolutional neural network model or other feature extraction methods are used to obtain feature representations of these local regions. These feature representations can be low-level image features, such as color, texture, etc., or high-level abstract features, such as shape, pattern, etc., which together constitute a comprehensive description of the local area.

[0093] After obtaining the feature representation of the local area, the local context information is combined with the global context information. The combination can be direct superposition or weighted fusion. The direct superposition method is simple and intuitive. It splices the local context information with the global context information in the feature dimension to form a feature vector containing richer information. The weighted fusion method is more flexible. It assigns a weight to each feature according to the importance of the local context information and the global context information, and then performs a weighted sum. This method can more accurately reflect the relative importance of different information, thereby obtaining a more accurate feature representation. By combining local context information with global context information, the global scene layout and the relationship between objects in the video frame are captured, and the local details around each area or pixel are focused on.

[0094] S109, based on the feature map that integrates global context information, dynamic graph modal information and local context information, outputs the semantic segmentation result, which is a label map containing multiple pixels, where the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

[0095] Among them, step S109 integrates the information extracted and fused in all previous steps to produce the final semantic segmentation result. This step converts the feature map that integrates global context information, dynamic graph modal information and local context information into an intuitive and easy-to-understand semantic segmentation map. Specifically, this feature map now contains rich feature representations of each pixel in the video frame, which not only reflect the properties of the pixel itself, but also incorporate the global and local context information related to it.

[0096] To derive semantic segmentation results from feature maps, the system uses an advanced semantic segmentation algorithm. This algorithm traverses each pixel in the feature map and determines the category to which the pixel belongs, such as vehicle, person, building or other background element, based on its feature vector. This process essentially classifies the feature vector and maps it to a predefined category label. Since the feature vector incorporates global, dynamic and local information, the classification result can accurately reflect the actual situation in the video frame and maintain high accuracy even in complex and changing mining environments.

[0097] The final output of semantic segmentation is a label map containing multiple pixels, in which each pixel is assigned a category label. This label map intuitively shows the distribution of objects and background structure in the mining area video information, providing strong support for subsequent analysis and decision-making. For example, through the semantic segmentation results, the vehicle driving paths, personnel activity areas and potential safety hazards in the mining area can be identified, thus providing a scientific basis for the safety management and optimized scheduling of the mining area.

[0098] In addition, the accuracy of the semantic segmentation results also benefits from the fine processing and feature fusion in the previous steps. The introduction of global context information enables the system to grasp the overall scene layout and the relationship between objects in the video frame; the fusion of dynamic graph modal information captures the temporal changes and dynamic features in the video; the combination of local context information further enhances the system's ability to capture detailed information. These factors work together to make the final semantic segmentation results both comprehensive and accurate, laying a solid foundation for the intelligent management of mining areas.

[0099] Furthermore, after outputting the semantic segmentation result, it also includes:

[0100] The label image is expanded to obtain an expanded label image; the label image is then corroded to obtain an eroded label image; the difference area between the expanded label image and the eroded label image is calculated to obtain an edge-enhanced label image of the object.

[0101] Among them, edge enhancement aims to enhance the recognition accuracy of object edges in semantic segmentation results. After outputting the semantic segmentation results, a label map containing multiple pixels and their corresponding category labels is obtained. A series of fine morphological operations are performed on the label map to optimize edge details.

[0102] Specifically, the first step is to perform dilation. The dilation operation is a morphological filtering technique that expands the boundaries of objects in the image by sliding a structural element (a rectangular or elliptical window in this embodiment) on the image and calculating the local maximum value at each position. The purpose of this step is to expand the edges of the object outward, which helps to more clearly define the boundary between the object and the background in the subsequent steps. After the dilation process, we obtain an expanded label image in which the edges of the object are moderately enlarged, reducing the edge blurring problem caused by inaccurate segmentation.

[0103] Next, we perform erosion on the original label map. In contrast to dilation, erosion shrinks the boundaries of objects in the image by calculating the minimum value of the area covered by the structural element. This step helps remove small protrusions or noise on the edges of objects, making the outline of the object smoother and more compact. The result of the erosion process is a eroded label map, which reflects the boundary of the object after being appropriately shrunk.

[0104] Subsequently, the difference area between the expanded label map and the eroded label map is calculated by comparing the two label maps pixel by pixel, aiming to identify the boundary change areas caused by the expansion and erosion operations. These change areas actually reflect the enhancement effect of the edge of the object, because the expansion operation expands the outer boundary of the object, while the erosion operation shrinks the inner boundary of the object. The difference area between the two is the edge enhancement area of ​​the object. Through this calculation, an edge-enhanced label map is obtained, which significantly enhances the visibility and accuracy of the edge of the object while keeping the internal area of ​​the object unchanged.

[0105] Furthermore, after outputting the semantic segmentation result, it also includes:

[0106] Each pixel in the semantic segmentation result is regarded as a node in the graph, and the connection between nodes is defined according to spatial proximity and color similarity. By constructing a minimum spanning tree, the average feature distance between each node and the surrounding nodes is calculated. If the difference between the average feature distance of any node and the average value of its category exceeds a threshold, the node is reclassified into the category to which its nearest category node belongs to ensure consistency within the segmentation area.

[0107] Specifically, each pixel in the semantic segmentation result is considered as an independent node in the graph structure. The connections between these nodes are not established randomly, but based on two key factors: spatial proximity and color similarity. Spatial proximity ensures the natural connection between adjacent pixels, while color similarity further strengthens the connection between visually similar pixels. Such graph structure construction provides a solid foundation for subsequent node classification and reclassification.

[0108] Next, the minimum spanning tree algorithm in graph theory is used to traverse the entire graph structure. The minimum spanning tree is a special graph structure that connects all nodes in the graph while ensuring that the total weight of the connection (here, the average feature distance between nodes) is minimized. By calculating the average feature distance between each node and the surrounding nodes, the similarities and differences between nodes are quantified. This step is key to identifying potential classification errors because abnormal nodes (i.e., those that differ significantly from the average value of their category) tend to stand out in this step.

[0109] Once these abnormal nodes are identified, they are reclassified based on their similarity to the nearest class nodes. Specifically, if the difference between the average feature distance of a node and the average value of its class exceeds a preset threshold, the node will be reclassified into the class to which its nearest class node belongs. This step ensures consistency within the segmented region and reduces discontinuity or confusion caused by misclassification.

[0110] Furthermore, after outputting the semantic segmentation results, it also includes: according to the semantic segmentation results, statistics are performed on specific objects or background categories in the mine video information, and a report containing object quantity, location distribution and motion status information is generated to facilitate mine management and safety monitoring.

[0111] Specifically, after obtaining the semantic segmentation results, the system can automatically identify and classify each pixel in the video frame and assign it to the corresponding object or background category. Then, the system will count the number of these classified objects or backgrounds, such as the total number of vehicles, the number of personnel, and the presence or absence of specific facilities in the mine area. This information is crucial for understanding the real-time operating status of the mine area.

[0112] In addition to quantity statistics, the system will further analyze the location distribution of these specific objects or backgrounds. By locating the specific position of each object in the video frame, the system can draw a distribution map of them in the mining area. This distribution map intuitively shows the spatial layout of different objects or backgrounds in the mining area, which helps managers quickly identify key areas or potential risk points. Combined with dynamic graph modal information, the system can also analyze the motion state of specific objects. For example, for vehicles, the system can track key indicators such as their driving trajectory, speed changes, and dwell time; for personnel, the system can monitor their movement paths, activity areas, and behavior patterns. These motion state information is of great significance for evaluating the operating efficiency of mining areas, predicting potential safety hazards, and optimizing resource allocation.

[0113] Based on the same inventive concept, Figure 2 As shown, the present invention provides a dynamic graph modal semantic segmentation system for vehicle-mounted mining video monitoring, comprising:

[0114] The video information acquisition module 201 is used to acquire the video information of the mining area through the vehicle-mounted camera;

[0115] A feature extraction module 202 is used to extract features from video frames using a pre-trained first convolutional neural network model, and to fuse features from different convolutional layers to obtain a feature set with richer feature representation;

[0116] The time dynamic map acquisition module 203 is used to calculate the pixel difference between adjacent frames by using the time continuity between video frames to obtain the inter-frame differential image; detect the motion area in the video according to the inter-frame differential image; identify the static background from the initial frame or multiple consecutive frames of the mining area video information sequence through the visual background extractor, combine the motion area with the static background based on image superposition, and construct a dynamic map in the time dimension;

[0117] The spatial dynamic map acquisition module 204 is used to divide the video frame into multiple regions, each region corresponds to a specific object or background; use the pre-trained second convolutional neural network model to extract features from each region to obtain a feature vector of the region; and construct a dynamic map of the spatial dimension based on the feature vector of the region;

[0118] A fusion module 205 is used to fuse the dynamic graphs of the time dimension and the space dimension to form a dynamic graph modality;

[0119] Attention allocation module 206, for calculating the similarity between the feature vector of each region and the feature vectors of other regions based on the dynamic graph modality after feature extraction, and allocating attention weights based on the similarity, so that the network pays attention to other regions that are most relevant to the current region, so as to enhance feature representation;

[0120] A global fusion module 207 is used to extract global context information using a global average pooling layer, match the global context information with the weighted original feature map containing dynamic graph modal information through upsampling, and perform element-by-element addition or element-by-element multiplication of two feature maps with the same spatial resolution at corresponding pixel positions to achieve the combination of global context information and dynamic graph modal information;

[0121] A local fusion module 208 is used to obtain feature representations of each region in the video frame or the local regions around a specific pixel, and combine the local context information with the global context information by directly superimposing or weighted fusion features;

[0122] The output module 209 outputs a semantic segmentation result based on a feature map that integrates global context information, dynamic graph modal information, and local context information. The semantic segmentation result is a label map containing multiple pixels, where the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

[0123] Based on the same inventive concept, Figure 3 As shown, the present invention provides an electronic device, including: a memory 302, a processor 301, and a computer program stored in the memory 302 and executable on the processor 301, wherein the processor 301 implements a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring when executing the computer program.

[0124] Based on the same inventive concept, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring.

[0125] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring, characterized in that: include: Obtain mining area video information through vehicle-mounted cameras; The pre-trained first convolutional neural network model is used to extract features from the video frames, and the features of different convolutional layers are fused to obtain a feature set; Using the temporal continuity between video frames, the pixel differences between adjacent frames are calculated to obtain the inter-frame differential image; According to the inter-frame difference image, the moving area in the video is detected; the static background is identified from the initial frame or multiple consecutive frames of the mining area video information sequence through the visual background extractor, and the moving area is combined with the static background based on image superposition to construct a dynamic map in the time dimension; The video frame is divided into multiple regions, each of which corresponds to a specific object or background; the pre-trained second convolutional neural network model is used to extract features from each region to obtain the feature vector of the region; based on the feature vector of the region, a dynamic graph of the spatial dimension is constructed; The dynamic graphs of time dimension and space dimension are integrated to form a dynamic graph modality; After the second convolutional neural network model extracts features, for the feature vector of each region, based on the dynamic graph modality, the similarity between the feature vector of each region and the feature vectors of other regions is calculated, and attention weights are assigned based on the similarity, so that the second convolutional neural network model pays attention to other regions that are most relevant to the current region; The global context information is extracted using a global average pooling layer, and the global context information is converted to the same spatial resolution as the initial feature map containing dynamic spectral modal information after region-weighted processing through upsampling; Adding or multiplying the global context information and the dynamic graph modal information pixel by pixel at corresponding pixel positions at the same resolution to achieve the combination of the global context information and the dynamic graph modal information; According to each area in the video frame or the local area around a specific pixel, the feature representation of these local areas is obtained, and the local context information is combined with the global context information by directly superimposing or weighted fusion features; Based on the combination of the global context information and the dynamic graph modal information and the combination of the local context information and the global context information, a semantic segmentation result is output; the semantic segmentation result is a label map containing multiple pixels, wherein the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

2. According to claim 1, a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring is characterized in that: The step of extracting features from video frames using the pre-trained first convolutional neural network model and fusing features from different convolutional layers specifically includes: Selecting a first convolutional neural network model that has been pre-trained on a large image dataset, wherein the first convolutional neural network model has multiple convolutional layers, each of which can extract image features at different levels; Input each frame of the mining area video information into the first convolutional neural network model, and extract the output feature map of each convolutional layer respectively; Using a feature pyramid network structure, feature maps of different convolutional layers are upsampled or downsampled so that the feature maps have the same spatial resolution; Feature maps with the same spatial resolution are spliced ​​in the channel dimension to form a fused feature set, which contains a variety of feature information from low layers to high layers.

3. According to claim 1, a dynamic graph modal semantic segmentation method for vehicle-mounted mining video monitoring is characterized in that: The dynamic graphs of the time dimension and the space dimension are integrated to form a dynamic graph modality, which specifically includes: Feature stitching is used to combine the dynamic change map obtained in the time dimension based on inter-frame difference and background subtraction with the static feature map obtained in the space dimension based on region division and feature extraction according to preset weights; among them, the dynamic map in the time dimension captures the motion trajectory and state changes of objects in the mining scene, and the dynamic map in the space dimension reflects the inherent attributes and relative position relationships of different areas in the scene.

4. The method for dynamic graph modal semantic segmentation of vehicle-mounted mining video monitoring according to claim 1 is characterized in that: The step of calculating the similarity between each regional feature vector and other regional feature vectors and allocating attention weights based on the similarity specifically includes: The cosine similarity formula is used for measurement, and the cosine similarity formula is: in, and Represent the feature vectors of two different regions, represents the cosine similarity between these two feature vectors; Based on the similarity obtained by the cosine similarity formula, the attention weight is assigned, and the weight calculation formula is: in, Indicates Region to The attention weight of the region, and Respectively and The feature vector of the region; is the number of eigenvectors; For the The feature vector of a region.

5. The method for dynamic graph modal semantic segmentation of vehicle-mounted mining video monitoring according to claim 1 is characterized in that: After outputting the semantic segmentation result, the method further includes: The label image is expanded to obtain an expanded label image; the label image is then corroded to obtain an eroded label image; and a difference area between the expanded label image and the eroded label image is calculated to obtain an edge-enhanced label image of the object.

6. The method for dynamic graph modal semantic segmentation of vehicle-mounted mining video monitoring according to claim 5 is characterized in that: After outputting the semantic segmentation result, the method further includes: Each pixel in the semantic segmentation result is regarded as a node in the graph, and the connection between nodes is defined based on spatial proximity and color similarity; By constructing a minimum spanning tree, the average feature distance between each node and the surrounding nodes is calculated. If the difference between the average feature distance of any node and the average value of its category exceeds a threshold, the node is reclassified into the category to which its nearest category node belongs to ensure consistency within the segmented area.

7. The method for dynamic graph modal semantic segmentation of vehicle-mounted mining video monitoring according to claim 1 is characterized in that: After outputting the semantic segmentation result, the method further includes: According to the semantic segmentation results, statistics are taken on specific objects or background categories in the mining area video information, and a report containing information on the number of objects, location distribution, and motion status is generated to facilitate mining area management and safety monitoring.

8. A dynamic graph modal semantic segmentation system for vehicle-mounted mining video monitoring, characterized in that: The system comprises: A video information acquisition module is used to acquire video information of the mining area through a vehicle-mounted camera; A feature extraction module, used to extract features from video frames using a pre-trained first convolutional neural network model, and to fuse features from different convolutional layers to obtain a feature set; The time dynamic map acquisition module is used to calculate the pixel difference between adjacent frames by using the time continuity between video frames to obtain the inter-frame differential image; detect the motion area in the video based on the inter-frame differential image; identify the static background based on the initial frame or multiple consecutive frames of the mining area video information sequence through the visual background extractor, combine the motion area with the static background based on image superposition, and construct a dynamic map in the time dimension; The spatial dynamic map acquisition module is used to divide the video frame into multiple regions, each region corresponds to a specific object or background; use the pre-trained second convolutional neural network model to extract features from each region to obtain the feature vector of the region; and construct a dynamic map of the spatial dimension based on the feature vector of the region; The fusion module is used to fuse the dynamic graphs of the time dimension and the space dimension to form a dynamic graph modality; An attention allocation module, configured to calculate the similarity between the feature vector of each region and the feature vectors of other regions based on the dynamic graph modality after the feature extraction of the second convolutional neural network model, and to allocate attention weights based on the similarity, so that the second convolutional neural network model pays attention to other regions that are most relevant to the current region, so as to enhance feature representation; A global fusion module is used to extract global context information using a global average pooling layer, and convert the global context information into the same spatial resolution as the initial feature map containing dynamic graph modal information after regional weighting processing through upsampling; the global context information and the dynamic graph modal information are added or multiplied pixel by pixel at corresponding pixel positions at the same resolution to achieve the combination of the global context information and the dynamic graph modal information; A local fusion module is used to obtain feature representations of each region in a video frame or a local region around a specific pixel, and combine local context information with global context information by directly superimposing or weighted fusion features; The output module outputs a semantic segmentation result based on the combination of the global context information and the dynamic graph modal information and the combination of the local context information and the global context information; the semantic segmentation result is a label map containing multiple pixels, wherein the label of each pixel corresponds to the object or background category represented by the pixel in the mining area video information.

9. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps corresponding to the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps corresponding to the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Real-time semantic segmentation method based on multi-feature reuse

    CN116740359A

  • VR panoramic space information analysis method and system based on artificial intelligence

    CN118570688A