Mine scene segmentation method and system based on graph neural network and comparative learning

By applying graph-based neural network and comparison learning methods in mine scene segmentation, and using the encoder and decoder of the segmentation model for feature extraction and expansion, the problem of low image segmentation efficiency in mine scenes is solved, and efficient and accurate mining boundary recognition and segmentation effect is achieved.

CN119942101AActive Publication Date: 2025-05-06CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411760984.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

The image segmentation efficiency of mine scenes is low, and existing deep learning models are difficult to effectively classify and segment complex mine scenes, and rely on manual interpretation to lead to inefficiency.

Method used

The mining scene segmentation method based on graph neural network and contrast learning is adopted to extract feature in the mining area image through the encoder of the segmentation model, and the decoder is used to expand the feature representation into the original dimension to generate the segmentation result image. The segmentation model includes a pre-trained SAM backbone network and a LORA fine-tuning module, combining to optimize model parameters with comparative learning loss.

Benefits of technology

The efficiency and accuracy of mine scene image segmentation are improved, and the mine boundaries are quickly identified through automated image analysis and interpretation processes, adapting to complex environments and changing landform features, and the segmentation accuracy and robustness of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942101A_ABST
    Figure CN119942101A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a mine scene segmentation method and system based on a graph neural network and comparative learning, and the method can call a segmentation model after obtaining a mine region image. And performing feature extraction in the mine area image through an encoder of the segmentation model to obtain feature representation of the mine area image. And expanding the feature representation into an original dimension of the mine region image through a decoder of the segmentation model to generate and output a segmentation result image. Wherein the encoder of the segmentation model further comprises an LORA fine tuning module, and model parameters of the SAM backbone network can be adjusted by introducing a low-rank matrix in the training process. According to the method, deep learning can be applied to the field of mine identification, so that the efficiency and accuracy of mine monitoring are improved. The automatic image analysis and interpretation process is utilized to quickly complete the recognition of the mine boundary, and the segmentation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image segmentation, and in particular to a method and system for segmenting a mine scene based on graph neural network and contrastive learning. Background Art

[0002] Mine scene segmentation refers to the image segmentation process of segmenting mining land, pools, stockpiling areas, ancillary buildings and other mining scene areas in remote sensing images of mining areas through image feature recognition. Mine scene segmentation is a coarse-grained research object with complex and diverse features. Features of different areas are easily confused with each other, making it difficult to achieve large-scale accurate identification. Therefore, for mining scenes and related downstream tasks, it is necessary to rely on manual field surveys or manual interpretation of remote sensing images to identify and segment mining scenes. This image segmentation method is inefficient.

[0003] In order to improve the efficiency of mining scene segmentation, mining scene segmentation can also be performed based on image processing algorithms such as deep learning models. Image processing algorithms can read the color, texture, edge, shape and other features in the mining scene image, perform operations such as region segmentation, region merging, region splitting and boundary adjustment, and segment the corresponding feature areas from the mining scene image.

[0004] However, due to the insufficient amount of data available in the field of mine segmentation and the inadequacy of model design, even though the deep learning model has powerful image processing capabilities, it is still difficult to effectively classify and segment complex scene objects such as mines due to the huge intra-class gap, and it is impossible to efficiently and quickly complete the identification of mine boundaries. In addition, it relies too much on expert knowledge and requires a lot of manual visual interpretation, resulting in low efficiency in mine scene image segmentation. Summary of the invention

[0005] In view of this, an embodiment of the present application provides a mine scene segmentation method and system based on graph neural network and contrastive learning to solve the problem of low efficiency of mine scene image segmentation.

[0006] According to one aspect of the present application, a mine scene segmentation method based on graph neural network and contrastive learning is provided, the method comprising:

[0007] Acquire images of the mine area;

[0008] Calling a segmentation model, the segmentation model is a deep learning model obtained by contrastive learning training with an additional output head; the segmentation model includes an encoder and a decoder, the encoder includes a pre-trained SAM backbone network and a low-rank adaptive LORA fine-tuning module; the LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training; the decoder includes a feature propagation tool based on a graph neural network;

[0009] Extracting features from the mine area image by the encoder to obtain a feature representation of the mine area image;

[0010] Expanding the feature representation to the original dimension of the mine area image by the decoder to generate a segmentation result image;

[0011] The segmentation result image is output.

[0012] Optionally, the method further includes:

[0013] Acquire a training data set, wherein the training data set includes a large-scale data set and a mine scene data set; the large-scale data set and the mine scene data set are marked with segmentation result labels;

[0014] Inputting the training data set into the segmentation model to obtain a training result image output by the segmentation model;

[0015] Calculating contrastive learning loss based on the training result image and the segmentation result label;

[0016] Model parameters of the segmentation model are optimized using the contrastive learning loss.

[0017] Optionally, inputting the training data set into the segmentation model to obtain a training result image output by the segmentation model includes:

[0018] Perform pre-training on the SAM backbone network using the large-scale dataset;

[0019] Inputting the mine scene data set into the pre-trained SAM backbone network to obtain a feature representation set;

[0020] Based on the low-rank matrix, performing graph feature aggregation on the feature representation set to obtain aggregated image data;

[0021] The decoder performs upsampling on the aggregated image data to generate the training result image.

[0022] Optionally, calculating the contrastive learning loss based on the training result image and the segmentation result label includes:

[0023] enhancing the feature representation extracted by the encoder through the additional output head;

[0024] Get the predicted mask output by each round of training process and the corresponding segmentation result label;

[0025] Divide the training data set into a positive sample set and a negative sample set based on the segmentation result label, wherein the positive sample set includes pixel points of the same segmentation category as the segmentation result label; and the negative sample set includes pixel points of a different segmentation category from the classification label;

[0026] A contrastive learning loss is calculated based on the positive sample set and the negative sample set.

[0027] Optionally, calculating contrastive learning loss according to the positive sample set and the negative sample set includes:

[0028] According to the prediction mask and the segmentation result label corresponding to the prediction mask, construct a sampling pixel set, wherein the sampling pixel set includes pixel points whose segmentation result label is positive and whose prediction mask is negative;

[0029] Performing multiple random samplings from the sampling pixel set;

[0030] The contrastive learning loss is calculated based on the positive sample set and negative sample set based on the sampling results of multiple random sampling as anchor points.

[0031] Optionally, extracting features from the mine area image by the encoder to obtain a feature representation of the mine area image includes:

[0032] Reading raw data from the mine area image, the raw data including pixel values ​​of pixel points in the mine area image;

[0033] Obtaining a downsampling weight matrix and a downsampling bias vector of the SAM backbone network;

[0034] According to the downsampling weight matrix and the downsampling bias vector, mapping the original data to a target feature space to obtain hidden layer output data;

[0035] Based on a linear rectifier unit activation function, the hidden layer output data is converted into activation output data, wherein the linear rectifier unit activation function is used to set negative values ​​in the hidden layer output data to zero and keep positive values ​​unchanged;

[0036] A feature representation of the mine region image is generated based on the activation output data.

[0037] Optionally, generating a feature representation of the mine area image according to the activation output data includes:

[0038] Get the upsampling weight matrix and upsampling bias vector;

[0039] An output vector is calculated based on the activation output data, the upsampling weight matrix, and the upsampling bias vector to obtain a feature representation of the mine area image.

[0040] Optionally, the method further includes:

[0041] Obtaining a feature representation and a two-dimensional coordinate of a first pixel point and a second pixel point, wherein the first pixel point and the second pixel point are two pixel points in the mine area image;

[0042] Calculating the cosine similarity between the first pixel and the second pixel according to the feature representation;

[0043] Calculate the coordinate distance between the first pixel point and the second pixel point according to the two-dimensional coordinates;

[0044] Acquire a learnable parameter, and construct an adjacency matrix according to the learnable parameter, the cosine similarity and the coordinate distance;

[0045] An edge set is constructed according to the adjacency matrix, and a graph vector is constructed according to the edge set, wherein the graph vector includes a node set and an edge set.

[0046] Optionally, the method further includes:

[0047] Constructing a diagonal matrix based on the adjacency matrix, wherein the diagonal elements in the diagonal matrix are the sums of the corresponding row elements in the adjacency matrix;

[0048] performing normalization processing on the adjacency matrix according to the diagonal matrix to generate a normalized matrix, wherein the normalized matrix is ​​obtained by multiplying the adjacency matrix by an inverse square root of the diagonal matrix;

[0049] Multiplying the normalized matrix with the original feature matrix to perform node feature aggregation;

[0050] The node feature aggregation results are linearly transformed through the weight matrix to obtain a new feature representation.

[0051] According to another aspect of the present application, a mine scene segmentation system based on graph neural network and contrastive learning is provided, the system comprising:

[0052] An image acquisition module, used to obtain images of the mining area;

[0053] A model calling module is used to call a segmentation model, wherein the segmentation model is a deep learning model obtained by contrastive learning training with an additional output head; the segmentation model includes an encoder and a decoder, wherein the encoder includes a pre-trained SAM backbone network and a low-rank adaptive LORA fine-tuning module; the LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training; the decoder includes a feature propagation tool based on a graph neural network;

[0054] A feature extraction module, used for performing feature extraction in the mine area image through the encoder to obtain a feature representation of the mine area image;

[0055] A segmentation image generation module, used for expanding the feature representation to the original dimension of the mine area image through the decoder to generate a segmentation result image;

[0056] The result output module is used to output the segmentation result image.

[0057] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein when the processor executes the program, the above-mentioned mine scene segmentation method based on graph neural network and contrastive learning is implemented.

[0058] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned mine scene segmentation method based on graph neural network and contrastive learning is implemented.

[0059] By means of the above technical solution, the embodiment of the present application provides a method and system for segmenting a mine scene based on a graph neural network and contrastive learning, wherein the method can call a segmentation model after acquiring a mine area image. Feature extraction is performed in the mine area image through the encoder of the segmentation model to obtain a feature representation of the mine area image. And the feature representation is expanded to the original dimension of the mine area image through the decoder of the segmentation model to generate and output a segmentation result image. Among them, the encoder of the segmentation model also includes a LORA fine-tuning module, which can adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during the training process. The method can improve the efficiency and accuracy of mine monitoring by applying deep learning to the field of mine identification. By using the automated image analysis and interpretation process, the identification of the mine boundary is quickly completed, and the segmentation efficiency is improved. In addition, the graph convolutional network is used to capture the complex topological structure and pixel relationship in the image, and the distance relationship in the feature space is optimized in combination with the contrast loss, which can improve the segmentation accuracy and robustness of the model under different terrain conditions, and adapt to the complex environment and changing geomorphic features of the mine.

[0060] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0062] Figure 1 A schematic diagram of the mining scene segmentation process provided in an embodiment of the present application;

[0063] Figure 2 A schematic diagram of the client-server connection relationship provided in the embodiment of the present application;

[0064] Figure 3 A schematic flow chart of a method for segmenting a mine scene based on a graph neural network and contrastive learning provided in an embodiment of the present application;

[0065] Figure 4 A schematic diagram of the process of outputting the segmentation result image provided in an embodiment of the present application;

[0066] Figure 5 A schematic diagram of the segmentation model training process provided in an embodiment of the present application;

[0067] Figure 6 A schematic diagram of a graph feature aggregation process provided in an embodiment of the present application;

[0068] Figure 7 A schematic diagram of the loss calculation process provided in the embodiment of the present application;

[0069] Figure 8 A schematic diagram of an image conversion process provided in an embodiment of the present application;

[0070] Fig. 9 A schematic diagram of a feature transfer process provided in an embodiment of the present application;

[0071] Fig.10 A schematic diagram of the structure of a mine scene segmentation system based on graph neural network and contrastive learning provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.

[0073] In the embodiment of the present application, the mine scene segmentation refers to the image processing process of segmenting the mine scene areas such as mining land, water pool, stockpile area, and ancillary buildings in the mine area image by performing pixel feature recognition on the mine area image, such as Figure 1 shown.

[0074] The raw data for mining scene segmentation is the mining area image. The mining area image refers to the image data generated by remote sensing monitoring of the mining area. The mining area image can be obtained by remote sensing equipment such as high-resolution satellites, which can take images of the mining area. The mining area image can also be obtained by taking multiple local images by detection equipment such as visible radar and remote sensing aircraft, and then stitching them together according to the geographical location of the area corresponding to the local image.

[0075] The result data obtained by the segmentation of the mining scene is the segmentation result image. The segmentation result image is an image whose parameters such as size and resolution used to characterize the image specifications are the same as those of the mining area image. In order to characterize the segmentation result, the segmented area can be marked in the segmentation result image by color, boundary, marking symbol, etc. Among them, the pixels in each segmented area can have the same marking method. For example, if the segmentation result image is a binary image, then in the segmentation result image, the pixels in the mining area are set as the foreground color, and the pixels outside the mining area are set as the background color.

[0076] In an embodiment of the present application, the graph neural network (GNN) is a deep learning framework specifically used to process graph structure data. The graph neural network can convert the graph structure data into a standardized and standard representation by formulating certain strategies on the nodes (Node) and edges (Edge) in the graph, and input it into a variety of different neural networks for training. In the graph neural network (GNN), nodes and edges are the basic elements that make up the graph. Nodes are used to represent entities, which can be pixels, sets of pixels, entire images, etc. Nodes can be associated with a feature vector, and the associated feature vector can contain attribute information of the node. In GNN, nodes can serve as the center of information processing, that is, nodes receive information from neighboring nodes and update their own feature representations. When performing classification tasks, the node itself is the target of classification, and GNN needs to predict the category label of each node.

[0077] Edges are used to represent the relationship or interaction between nodes, such as the characteristic distance, actual distance, and correlation between pixels. In GNN, edges can be used to define the path for information to be transmitted between nodes. That is, nodes can obtain information about their neighboring nodes through edges. Edges can also be set with weights to represent the strength or probability of the relationship, and these weights can be used to adjust the strength of information transmission. Edges can also be directional, and the direction of the edge is used to determine the direction of information flow. That is, in a directed graph, the direction of the edge is used to indicate the direction of information flow; in an undirected graph, the edge represents a bidirectional information flow relationship between nodes.

[0078] In GNN, nodes and edges can jointly define the structure of the graph. By aggregating the information of neighboring nodes, nodes can capture the structural characteristics of the local neighborhood, while edges can provide the path of such aggregation. This structured information transmission and update mechanism enables GNN to effectively process graph data and achieve excellent performance in various graph-related tasks.

[0079] In an embodiment of the present application, the contrastive learning is an unsupervised learning method, and contrastive learning can learn the feature representation of data by comparing the similarities and differences between different samples. Contrastive learning can focus on constructing positive sample pairs and negative sample pairs. Among them, a positive sample pair refers to two samples with similar features, such as different angles of the same object or different lighting conditions in a picture. A negative sample pair refers to two samples with different features, such as images of different objects.

[0080] Contrastive learning can reduce the distance between positive sample pairs and expand the distance between negative sample pairs through learning models, thereby promoting the model's in-depth understanding of the data. Therefore, contrastive learning can learn meaningful feature representations without labels. By comparing sample pairs, high-quality feature representations can be extracted and applied to tasks such as classification, clustering, and retrieval.

[0081] Based on this, in the embodiment of the present application, the mine scene segmentation method based on graph neural network and contrastive learning is an image segmentation method that integrates graph neural network and contrastive learning, which is used to segment the mine area image to determine each scene area from the mine area image. After image segmentation, a segmentation result image can be obtained, and the segmentation result image can be used to perform subsequent related operations, such as mine area display, setting safety defense zones, industrial land planning, geological risk assessment, etc.

[0082] It should be noted that the mine scene segmentation method based on graph neural network and contrastive learning can be run on devices with data processing capabilities, such as computers, mobile terminals, servers, industrial hosts, smart wearable devices, etc. The device with data processing capabilities can be an independent device or a combination of multiple devices. That is, in a feasible implementation, the mine scene segmentation method can be run on a single device, that is, all the steps of the mine scene segmentation method are performed by a separate device. For example, the mine scene segmentation method can be integrated in an application installed on a personal computer, and with the user's interactive instructions, the personal computer can execute the corresponding steps of the method according to the application.

[0083] In another feasible implementation, the mine scene segmentation method can be run on multiple devices, that is, multiple devices need to cooperate with each other to execute all the steps corresponding to the mine scene segmentation method. Figure 2 As shown, a server and multiple clients connected to the server can be set, and the multiple clients can be operated by different users respectively. The client can send control instructions or application data to the server through the network. The server can further process, distribute and store the application data in response to the control instruction, and feedback the processing result to the client to realize the image segmentation function corresponding to the mine scene segmentation method.

[0084] For ease of description, in some embodiments of the present application, one or more devices that execute the mine scene segmentation method are collectively referred to as data processing devices. It should be understood that the data processing device refers to a combination of one or more devices, and other implementations associated with the data processing device described in the present application by those skilled in the art by replacing the device combination method also fall within the scope of protection of the present application.

[0085] like Figure 3 As shown, a mine scene segmentation method based on graph neural network and contrastive learning provided in an embodiment of the present application includes:

[0086] S101, obtaining a mine area image.

[0087] When performing the segmentation of the mine scene, the data processing device needs to first obtain the mine area image. The mine area image can be obtained by remote sensing monitoring of the mine area. In order to obtain the mine area image, the data processing device can establish a communication connection with the remote sensing monitoring device or a data storage device corresponding to the remote sensing monitoring device.

[0088] The mine area image acquired by the data processing device may be a real-time image. Taking the remote sensing monitoring device as an example, when acquiring the mine area image, the data processing device may send an image acquisition request to the remote sensing monitoring device. After receiving the image acquisition request, the remote sensing monitoring device may take an image of the mine area to obtain the mine area image. The taken mine area image is then sent to the data processing device.

[0089] The mine area image acquired by the data processing device may also be a historical image. Taking the data storage device as an example, the data storage device may establish a communication connection with both the data processing device and the remote sensing monitoring device. In the process of performing remote sensing monitoring, the mine area image acquired by shooting may be sent to the data storage device for storage. When the data processing device needs to acquire the mine area image, it may send an image acquisition request to the data storage device. After acquiring the image acquisition request, the data storage device feeds back the mine area image to the data processing device.

[0090] In some embodiments, after acquiring the mine area image, the data processing device may also pre-process the acquired mine area image. That is, the data processing device may pre-process the acquired mine area image by graying, removing noise, enhancing contrast, and histogram equalization. Pre-processing can reduce the amount of data in subsequent image segmentation processing, remove noise and interference in the image, make the image clearer, and improve the contrast and brightness of the image.

[0091] S102: Calling a segmentation model.

[0092] After acquiring the mine area image, the data processing device can call the segmentation model. The data processing device can call the segmentation model in different ways. In some embodiments, the data processing device includes a server and a client. The segmentation model can be stored in the server, and after acquiring the mine image, the client can send a model call request to the server. After obtaining the model call request, the server can send model parameters to the client so that the client builds a segmentation model according to the obtained model parameters. The server can also receive the model call request and the mine area image sent by the client, and run the segmentation model on the server, and the server uses the segmentation model to perform scene segmentation processing on the mine area image.

[0093] In some embodiments, the segmentation model may be stored in a local memory of the data processing device. After acquiring the mine image, the data processing device may directly call and run the application corresponding to the segmentation model from the local memory, and use the segmentation model to perform scene segmentation processing on the mine area image.

[0094] In order to perform mine scene segmentation, an additional output head can be set for the segmentation model so that the segmentation model is trained by contrastive learning through the additional output head to obtain a deep learning model suitable for scene segmentation. The segmentation model includes an encoder and a decoder, and the encoder includes a pre-trained segmentation model (Segment Anything Model, SAM) backbone network and a low-rank adaptation (Low-Rank Adaptation, LORA) fine-tuning module. The LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training. The decoder includes a feature propagation tool based on a graph neural network.

[0095] That is Figure 4 As shown, the segmentation model can use an encoder and decoder structure as the model structure for mine segmentation. Among them, the encoder uses the SAM pre-trained backbone network combined with LORA fine-tuning to extract features. The decoder uses a graph neural network as a feature propagation tool to enhance the decoding ability of the model. When training the segmentation model, you can also use contrastive learning additional output heads to enhance the feature perception ability of the model. The LORA fine-tuning module adjusts the weights of the model by introducing additional low-rank matrices, thereby achieving rapid adaptation to new tasks. LORA fine-tuning can make the change in weights during task adaptation low-rank, that is, it can be expressed as the product of two smaller matrices. For example, the original weight matrix W0 is kept unchanged, and the weight update ΔW is decomposed into the product of two low-rank matrices B and A, that is, W=W0+BA, where B∈R d×r and A∈R r×k , r is the rank of the two matrices, and r<<min(d,k).

[0096] During the training process, W0 is fixed and only B and A are training parameters. Compared with full parameter adjustment, LoRA fine-tuning can significantly reduce the number of parameters that need to be trained, thereby reducing computational and storage costs. In addition, during the forward propagation process, the input x is transformed by the original weight matrix W0 and by the low-rank matrices A and B at the same time, and the final output is the weighted sum of the two.

[0097] In the image segmentation model, the LoRA fine-tuning module can be used to fine-tune the SAM encoder to adapt to specific downstream tasks. By introducing the LoRA fine-tuning module in the SAM encoder, the model can be adjusted to adapt to new tasks without significantly increasing the model size, thereby improving the flexibility and adaptability of the model. LoRA fine-tuning is suitable for processing large-scale data sets and complex tasks, and can reduce the demand for training resources while maintaining model performance.

[0098] In order to perform mine scene segmentation based on graph neural networks and contrastive learning, the segmentation model needs to be pre-trained. Figure 5 As shown, in some embodiments, the data processing device may obtain a training data set. The training data set includes a sample image and a training label of the sample image. When performing image segmentation, the sample image is the original image, and the training label is the segmentation result label, that is, the segmentation mask corresponding to the sample image. The training data set may include a large-scale data set and a mine scene data set, and the corresponding large-scale data set and the mine scene data set are both marked with segmentation result labels.

[0099] A large-scale dataset is a dataset for image segmentation tasks. A large-scale dataset may include original images and segmentation masks from various fields, so that the model can be trained on a wide range of data and can be adapted to specific tasks. For example, the large-scale dataset may be the SA-18 large-scale dataset. The SA-18 large-scale dataset may include 100 million masks applied to 11 million images.

[0100] The mining scene dataset is also a dataset for image segmentation tasks. The mining scene dataset can include sample images of mining scenes and the segmentation masks corresponding to the sample images of mining scenes. The mining scene dataset can make the trained model more suitable for the technical field of mining scene segmentation.

[0101] After obtaining the training data set, the data processing device may input the training data set into the segmentation model to obtain the training result image output by the segmentation model. Figure 6 As shown, in order to obtain the training result image, after the data processing device inputs the training data set into the segmentation model, it can perform feature extraction and mask generation based on the encoder and decoder in the segmentation model, that is, use the large-scale data set to pre-train the SAM backbone network. For example, the SA-18 large-scale data set is used to pre-train the model. The SA-18 large-scale data set can contain images of various scenes and can be used to train the model to recognize and segment different objects. In the pre-training stage, the network parameters based on the SAM backbone network can be continuously iteratively modified so that the encoder has the ability to recognize and understand the basic features and objects in the image.

[0102] After pre-training, the data processing device can input the mine scene data set into the pre-trained SAM backbone network to obtain a feature representation set. Since the amount of mine scene data is small, it belongs to a task that lacks data. Therefore, for tasks with insufficient data, the LORA fine-tuning module can be used to fine-tune the SAM backbone network to adapt to the needs of these tasks. That is, the data processing device can perform graph feature aggregation on the feature representation set based on the low-rank matrix to obtain aggregated image data. After pre-training, the encoder based on the SAM backbone network can extract the features of the image and perform feature aggregation to form a high-level understanding of the image content.

[0103] The feature map processed by the SAM encoder is upsampled by the upsampling layer in the decoder to restore the resolution of the original image. Therefore, the data processing device can perform upsampling on the aggregated image data through the decoder to generate the training result image.

[0104] After obtaining the training result image, the data processing device can calculate the comparative learning loss based on the training result image and the segmentation result label. That is, after obtaining the training result image, the data processing device can compare the training result image with the segmentation result label to determine the difference between the training result image and the segmentation result label to evaluate the output accuracy of the current segmentation model. In order to evaluate the accuracy of the current segmentation model, the training loss between the training result image and the segmentation result label can be calculated by a loss function. Among them, the loss function can be based on one or more combinations of cross entropy loss, Dice loss, Focal loss, Tversky loss, structural similarity loss, mean square error loss, and Hausdorff distance loss algorithms to calculate the loss to obtain the training loss.

[0105] In some embodiments, the training loss may be a contrastive learning loss. In order to calculate the contrastive learning loss, the data processing device may enhance the feature representation extracted by the encoder through the additional output head. Then obtain the prediction mask output by each round of training process and the corresponding segmentation result label. Divide the positive sample set and the negative sample set in the training data set based on the segmentation result label. Among them, the positive sample set includes pixel points of the same segmentation category as the segmentation result label; the negative sample set includes pixel points of a different segmentation category from the classification label. Then calculate the contrastive learning loss based on the positive sample set and the negative sample set.

[0106] During model training, the segmentation model can use additional output heads to enhance the features extracted by the model encoder. LoRA is applied to fine-tune the SAM Encoder. When outputting the training result image, the upsampling module can be used to restore the high-dimensional feature embedding to the size of the original image and retain its feature dimension.

[0107] When calculating the contrastive learning loss, a sampling pixel set may be constructed according to the prediction mask and the segmentation result label corresponding to the prediction mask, wherein the sampling pixel set includes pixels whose segmentation result label is positive and whose prediction mask is negative.

[0108] like Figure 7 As shown, for each round of training, the SAM Encoder fine-tuned with LoRA will have two sets of outputs, one for the predicted mask P mask , and the other set is the F output by the feature decoder with the same size as the original image decoder According to P mask and the corresponding segmentation result label Label. Design the pixel-level correspondence for comparative learning. For pixel i, its positive sample is the pixel with the same label as the corresponding segmentation result Label, and the negative sample is the pixel with a different label. The following loss function is obtained:

[0109]

[0110] Among them, P i is the positive sample set corresponding to pixel i, N i is a set of negative samples, and τ is a hyperparameter used to control the degree of model convergence. mask and the corresponding Label, let Label be positive and P mask The set of negative pixels is U notreacll =Label-P mask .

[0111] This part of the pixel set should be the focus of optimization, so multiple random sampling is performed from the sampled pixel set during each training; the sampling results of multiple random sampling are used as the positive sample set and negative sample set of the anchor point to calculate the comparative learning loss. By performing multiple random sampling in the sampled pixel set and adding the sampling results as the positive and negative samples of the anchor point to the calculation, the model's feature perception ability for these pixels can be increased. The distribution of pixels in the feature space can also be clustered, and the cluster center is used as the anchor point, and the outliers are treated as difficult samples and added to the calculation of the loss function.

[0112] After calculating the contrastive learning loss in the manner provided in the above embodiment, the data processing device can use the contrastive learning loss to optimize the model parameters of the segmentation model, that is, use the contrastive learning loss for iterative optimization. During each iteration, the data processing device can judge the contrastive learning loss with the preset loss condition. If the contrastive learning loss meets the preset loss condition, such as the contrastive learning loss is less than or equal to the loss threshold, the segmentation model can be output. If the contrastive learning loss does not meet the preset loss condition, such as the contrastive learning loss is greater than the loss threshold, iterative optimization can continue to be performed until the contrastive learning loss meets the preset loss condition.

[0113] After the segmentation model is obtained through training using the model training method provided in the above embodiment, the data processing device can apply the segmentation model to perform scene segmentation. Therefore, after calling the segmentation model, the data processing device can input the mine area image into the segmentation model, so that the segmentation model can extract features from the mine area image and generate a segmentation result image.

[0114] S103: extracting features from the mine area image by using the encoder to obtain a feature representation of the mine area image.

[0115] After the mine area image is input into the segmentation model, the encoder in the segmentation model can first extract features from the mine area image. Since the SAM backbone network in the encoder can be fine-tuned by LoRA, the introduction of additional low-rank matrices through LoRA fine-tuning can achieve effective adjustment of the model at a relatively low parameter cost.

[0116] like Figure 8 As shown, in some embodiments, the data processing device can read raw data from the mine area image. The raw data includes pixel values ​​of pixels in the mine area image. Then, the downsampling weight matrix and downsampling bias vector of the SAM backbone network are obtained. According to the downsampling weight matrix and the downsampling bias vector, the raw data is mapped to the target feature space to obtain hidden layer output data.

[0117] After obtaining the mine area image, the pixel points in the mine area image can be traversed to obtain the original data x in , and then obtain the downsampling weight matrix W down and the downsampling bias vector b down Then calculate the hidden layer output data x according to the following formula hid ,Right now:

[0118] x hid =D down (x in )=W down × in +bdown

[0119] After calculating and obtaining the hidden layer output data, the data processing device also converts the hidden layer output data into activated output data based on a linear rectifier unit activation function, wherein the linear rectifier unit activation function is used to set negative values ​​in the hidden layer output data to zero and keep positive values ​​unchanged.

[0120] The Rectified Linear Unit (ReLU) activation function can map linearly inseparable data to a high-dimensional space, making it linearly separable. The ReLU activation function can transform the hidden layer output data x hid The negative values ​​in are set to zero, and the positive values ​​remain unchanged, so that the hidden layer output data x hid Transformed into activation output data x act ,Right now:

[0121] x act =ReLU(x hid )=max(0,x hid )

[0122] The hidden layer outputs data x hid Transformed into activation output data x act After that, the data processing device can generate a feature representation of the mine area image according to the activation output data. By obtaining an upsampling weight matrix and an upsampling bias vector, an output vector can be calculated according to the activation output data, the upsampling weight matrix and the upsampling bias vector to obtain the feature representation of the mine area image.

[0123] In getting the activation output data x act After that, the data processing device can obtain the upsampling weight matrix W up and the upsampling bias vector b up , and then calculate the output vector x according to the following formula out ,Right now:

[0124] x out =D up (x act )=W up × act +b up

[0125] S104. Expanding the feature representation to the original dimension of the mine area image through the decoder to generate a segmentation result image.

[0126] After obtaining the feature representation of the mine area image, the data processing device can upsample the feature representation through a decoder to expand the feature representation to the original dimension of the mine area image to generate a segmentation result image. The segmentation result image is consistent with the mine area image in parameters such as image size and resolution.

[0127] Since the decoder includes a feature propagation tool based on a graph neural network, the graph neural network model has a complex topological structure. Therefore, using graph neural networks to model open-pit mines can make the output more adaptable to the complex structure of the mine.

[0128] The mining area also contains multiple areas that can be subdivided. For example, the mining area can also include open-pit mining land, water pools, stockpiling areas, and ancillary buildings. In the image segmentation task, these areas will be divided into one category. However, the characteristics of these areas will be very different. Therefore, if we only rely on complex prompts (task descriptions) to guide the model to cross the de-segmentation of features, we will face a large intra-class feature gap. Therefore, the message passing and feature aggregation mechanism of the graph model can be used to highlight the characteristics of the mining part and enhance the model's segmentation reasoning ability for the mining area.

[0129] That is Fig. 9 As shown, in some embodiments, when performing a mine scene segmentation task, the data processing device can obtain a feature representation and two-dimensional coordinates of a first pixel point and a second pixel point. The first pixel point and the second pixel point are two pixel points in the mine area image. For example, after obtaining the feature representation output by the encoder, the decoder can read the data in the feature representation to obtain the feature representation and the two-dimensional coordinates. That is, F i is the feature of the first pixel (point i), F j is the feature of the second pixel (point j). i is the two-dimensional coordinate of the first pixel (point i) on the feature map, Pos j is the two-dimensional coordinate of the second pixel point (point j) on the feature map.

[0130] The cosine similarity between the first pixel and the second pixel is calculated based on the feature representation, and the coordinate distance between the first pixel and the second pixel is calculated based on the two-dimensional coordinates. For example, based on the cosine similarity algorithm, the feature F of the first pixel (point i) can be calculated. i , the feature F of the second pixel (point j) j , calculate the cosine similarity cos(F i , F j ). And according to the two-dimensional coordinates Pos of the first pixel (point i) on the feature map i, the two-dimensional coordinates of the second pixel (point j) on the feature map are Pos j , calculate the L1 distance between the coordinates of points i and j on the feature map, that is:

[0131] L1=||Pos i ,Pos j ||1

[0132] Then obtain learnable parameters, such as γ and λ, and construct an adjacency matrix A according to the learnable parameters, the cosine similarity and the coordinate distance, that is:

[0133]

[0134] Then, an edge set is constructed according to the adjacency matrix, and a graph vector is constructed according to the edge set, wherein the graph vector includes a node set and an edge set. For example, the characteristic distance between the feature graph nodes and the actual distance are fused as the edge weight. For the mine area image segmentation task, the feature point is the node, and the edge set is established based on the above formula. Then for a graph G = (V, E), there are N nodes v i ∈V, and the corresponding edge (v i , v j )∈E.

[0135] When establishing a feature graph, the weight of the edge between nodes should contain two meanings, namely the distance between pixel features and the actual topological distance of pixels on the graph. Therefore, the edge value between nodes can be established adaptively, and the part that is actually connected in the topological relationship can have a larger edge weight, thereby increasing the information transmission capacity of this part.

[0136] In order to perform feature transfer, in some embodiments, the data processing device can construct a diagonal matrix based on the adjacency matrix. The diagonal elements in the diagonal matrix are the sums of the corresponding row elements in the adjacency matrix. For example, for the adjacency matrix A, self-loops can be set by graph, that is, the connection between the node and itself, to convert the adjacency matrix A into an adjacency matrix containing self-loops. Then calculate the adjacency matrix The sum of each row of elements in constructs a diagonal matrix That is, the diagonal matrix Each diagonal element in is the adjacency matrix The sum of the i-th row is the degree of node i.

[0137] Then, the adjacency matrix is ​​normalized according to the diagonal matrix to generate a normalized matrix. The normalized matrix is ​​obtained by multiplying the adjacency matrix with the inverse square root of the diagonal matrix. After obtaining the diagonal matrix, the inverse square root of the diagonal matrix can be calculated. And multiply the adjacency matrix by the inverse square root of the diagonal matrix to obtain the normalized matrix

[0138] Then, the normalized matrix is ​​multiplied by the original feature matrix to perform node feature aggregation, and the node feature aggregation result is linearly transformed through the weight matrix to obtain a new feature representation. That is, the new feature representation X′ is calculated according to the following formula:

[0139]

[0140] in, is the adjacency matrix containing self-loops, X is the input feature matrix, W is the network parameter, is the diagonal matrix of the graph. Then the feature transfer can follow the basic GCN transfer method and use multi-layer convolution to learn the weights of information transfer.

[0141] S105: Output the segmentation result image.

[0142] After generating the segmentation result image, the data processing device can obtain the segmentation result image generated by the segmentation model and output the segmentation result image to downstream tasks or other applications. For example, the output segmentation result image can be displayed by a visualization application to show the distribution location of each scene in the mine area.

[0143] By applying the technical solution of this embodiment, the method can call the segmentation model after acquiring the mine area image. Feature extraction is performed in the mine area image through the encoder of the segmentation model to obtain the feature representation of the mine area image. And the feature representation is expanded to the original dimension of the mine area image through the decoder of the segmentation model to generate and output the segmentation result image. Among them, the encoder of the segmentation model also includes a LORA fine-tuning module, which can adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during the training process. The method can improve the efficiency and accuracy of mine monitoring by applying deep learning to the field of mine identification. Using automated image analysis and interpretation processes, the identification of mine boundaries is quickly completed, and the segmentation efficiency is improved. In addition, the graph convolution network is used to capture the complex topological structure and pixel relationship in the image, and the distance relationship in the feature space is optimized in combination with contrast loss, which can improve the segmentation accuracy and robustness of the model under different terrain conditions, and adapt to the complex environment and changing geomorphic features of the mine.

[0144] Further, as a specific implementation of the mine scene segmentation method based on graph neural network and contrastive learning described in the above embodiment, the embodiment of the present application provides a mine scene segmentation system based on graph neural network and contrastive learning, such as Fig.10 As shown, the system includes:

[0145] An image acquisition module, used to obtain images of the mining area;

[0146] A model calling module is used to call a segmentation model, wherein the segmentation model is a deep learning model obtained by contrastive learning training with an additional output head; the segmentation model includes an encoder and a decoder, wherein the encoder includes a pre-trained SAM backbone network and a low-rank adaptive LORA fine-tuning module; the LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training; the decoder includes a feature propagation tool based on a graph neural network;

[0147] A feature extraction module, used for performing feature extraction in the mine area image through the encoder to obtain a feature representation of the mine area image;

[0148] A segmentation image generation module, used for expanding the feature representation to the original dimension of the mine area image through the decoder to generate a segmentation result image;

[0149] The result output module is used to output the segmentation result image.

[0150] Based on the mine scene segmentation system based on graph neural network and contrastive learning provided in the above embodiment, the efficiency and accuracy of mine monitoring can be improved by applying deep learning to the field of mine identification. And by using the automated image analysis and interpretation process, the identification of mine boundaries can be completed quickly, which can reduce time costs and human resource consumption compared to manual interpretation methods, while reducing the misjudgment rate. And by using a graph convolutional network (GCN) to capture the complex topological structure and pixel relationship in the image, combined with contrast loss to optimize the distance relationship in the feature space, the segmentation accuracy and robustness of the model under different terrain conditions can be significantly improved, and it can better adapt to the complex environment and changing geomorphic features of the mine. In addition, the system is not only suitable for monitoring mine boundaries, but can also be extended to other types of complex scenes to improve the breadth and expansibility of technical applications.

[0151] It should be noted that for other corresponding descriptions of the functional units involved in the mine scene segmentation system based on graph neural network and contrastive learning provided in the embodiment of the present application, reference can be made to the corresponding descriptions in the mine scene segmentation method based on graph neural network and contrastive learning provided in the above embodiment, and will not be repeated here.

[0152] The embodiment of the present application also provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a bus, a processor, a memory and a communication interface, and can also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.

[0153] Those skilled in the art will appreciate that the structure of the above-mentioned computer device is only a partial structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components, or combine certain components, or have a different arrangement of components.

[0154] In one embodiment, a computer-readable storage medium is also provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0155] In one embodiment, a computer program product is also provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0157] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0158] Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0159] Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0160] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0161] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0162] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A mine scene segmentation method based on graph neural network and contrastive learning, characterized in that: The method comprises: Acquire images of the mine area; Calling a segmentation model, the segmentation model is a deep learning model obtained by contrastive learning training with an additional output head; the segmentation model includes an encoder and a decoder, the encoder includes a pre-trained SAM backbone network and a low-rank adaptive LORA fine-tuning module; the LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training; the decoder includes a feature propagation tool based on a graph neural network; Extracting features from the mine area image by the encoder to obtain a feature representation of the mine area image; Expanding the feature representation to the original dimension of the mine area image by the decoder to generate a segmentation result image; The segmentation result image is output.

2. The method according to claim 1, characterized in that The method further comprises: Acquire a training data set, wherein the training data set includes a large-scale data set and a mine scene data set; the large-scale data set and the mine scene data set are marked with segmentation result labels; Inputting the training data set into the segmentation model to obtain a training result image output by the segmentation model; Calculating contrastive learning loss based on the training result image and the segmentation result label; Model parameters of the segmentation model are optimized using the contrastive learning loss.

3. The method according to claim 2, characterized in that Inputting the training data set into the segmentation model to obtain a training result image output by the segmentation model includes: Perform pre-training on the SAM backbone network using the large-scale dataset; Inputting the mine scene data set into the pre-trained SAM backbone network to obtain a feature representation set; Based on the low-rank matrix, performing graph feature aggregation on the feature representation set to obtain aggregated image data; The decoder performs upsampling on the aggregated image data to generate the training result image.

4. The method according to claim 2, characterized in that: Calculating contrastive learning loss based on the training result image and the segmentation result label includes: enhancing the feature representation extracted by the encoder through the additional output head; Get the predicted mask output by each round of training process and the corresponding segmentation result label; Divide the training data set into a positive sample set and a negative sample set based on the segmentation result label, wherein the positive sample set includes pixel points of the same segmentation category as the segmentation result label; and the negative sample set includes pixel points of a different segmentation category from the classification label; A contrastive learning loss is calculated based on the positive sample set and the negative sample set.

5. The method according to claim 4, characterized in that Calculating contrastive learning loss according to the positive sample set and the negative sample set includes: According to the prediction mask and the segmentation result label corresponding to the prediction mask, construct a sampling pixel set, wherein the sampling pixel set includes pixel points whose segmentation result label is positive and whose prediction mask is negative; Performing multiple random samplings from the sampling pixel set; The contrastive learning loss is calculated based on the positive sample set and negative sample set based on the sampling results of multiple random sampling as anchor points.

6. The method according to claim 1, characterized in that Extracting features from the mine area image by the encoder to obtain a feature representation of the mine area image includes: Reading raw data from the mine area image, the raw data including pixel values ​​of pixel points in the mine area image; Obtaining a downsampling weight matrix and a downsampling bias vector of the SAM backbone network; According to the downsampling weight matrix and the downsampling bias vector, mapping the original data to a target feature space to obtain hidden layer output data; Based on a linear rectifier unit activation function, the hidden layer output data is converted into activation output data, wherein the linear rectifier unit activation function is used to set negative values ​​in the hidden layer output data to zero and keep positive values ​​unchanged; A feature representation of the mine region image is generated based on the activation output data.

7. The method according to claim 6, characterized in that Generating a feature representation of the mine area image according to the activation output data includes: Get the upsampling weight matrix and upsampling bias vector; An output vector is calculated based on the activation output data, the upsampling weight matrix, and the upsampling bias vector to obtain a feature representation of the mine area image.

8. The method according to claim 1, characterized in that The method further comprises: Obtaining a feature representation and a two-dimensional coordinate of a first pixel point and a second pixel point, wherein the first pixel point and the second pixel point are two pixel points in the mine area image; Calculating the cosine similarity between the first pixel and the second pixel according to the feature representation; Calculate the coordinate distance between the first pixel point and the second pixel point according to the two-dimensional coordinates; Acquire a learnable parameter, and construct an adjacency matrix according to the learnable parameter, the cosine similarity and the coordinate distance; An edge set is constructed according to the adjacency matrix, and a graph vector is constructed according to the edge set, wherein the graph vector includes a node set and an edge set.

9. The method according to claim 8, characterized in that The method further comprises: Constructing a diagonal matrix based on the adjacency matrix, wherein the diagonal elements in the diagonal matrix are the sums of the corresponding row elements in the adjacency matrix; performing normalization processing on the adjacency matrix according to the diagonal matrix to generate a normalized matrix, wherein the normalized matrix is ​​obtained by multiplying the adjacency matrix by an inverse square root of the diagonal matrix; Multiplying the normalized matrix with the original feature matrix to perform node feature aggregation; The node feature aggregation results are linearly transformed through the weight matrix to obtain a new feature representation.

10. A mine scene segmentation system based on graph neural network and contrastive learning, characterized in that: The system comprises: An image acquisition module, used to obtain images of the mining area; A model calling module is used to call a segmentation model, wherein the segmentation model is a deep learning model obtained by contrastive learning training with an additional output head; the segmentation model includes an encoder and a decoder, wherein the encoder includes a pre-trained SAM backbone network and a low-rank adaptive LORA fine-tuning module; the LORA fine-tuning module is configured to adjust the model parameters of the SAM backbone network by introducing a low-rank matrix during training; the decoder includes a feature propagation tool based on a graph neural network; A feature extraction module, used for performing feature extraction in the mine area image through the encoder to obtain a feature representation of the mine area image; A segmentation image generation module, used for expanding the feature representation to the original dimension of the mine area image through the decoder to generate a segmentation result image; The result output module is used to output the segmentation result image.

Citation Information

Patent Citations

  • Medical image segmentation method based on multi-level feature extraction and attention mechanism fusion

    CN117456183A

  • Dynamic decision image segmentation method based on SAM basic model

    CN118072378A

  • Medical image rapid segmentation method based on segmentation cutting model

    CN118229974A

  • Training method of image segmentation model, and image segmentation method and system

    CN118657944A