PC screen semantic segmentation method based on efficient attention mechanism

By introducing an efficient attention mechanism in semantic segmentation technology, combining the codec module and the Transformer adaptive module, the problem of poor single-objective semantic segmentation effect in the existing technology under small samples and specific applications is solved, and higher classification accuracy and generalization capabilities are achieved.

CN113869396BActive Publication Date: 2025-06-06HEFEI HIGH DIMENSIONAL DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111127462.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-06-06
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

The existing semantic segmentation technology performs poorly when dealing with single-objective semantic segmentation under small samples and specific applications, and lacks classification accuracy for samples with large differences in in-class features.

Method used

The PC screen semantic segmentation method based on an efficient attention mechanism is adopted, and the network model is built through the codec module and the Transformer adaptive module, and the data set and loss function are used for training, which improves the classification accuracy of the model for samples with large differences in the class characteristics.

Benefits of technology

The classification accuracy of samples with large differences in characteristics within the class has been improved, and the generalization ability and segmentation effect of the model have been enhanced, especially in small samples and specific application scenarios, which have significantly improved performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113869396B_ABST
    Figure CN113869396B_ABST
Patent Text Reader

Abstract

The present invention particularly relates to a method for semantic segmentation of a PC screen based on an efficient attention mechanism, comprising the following steps: S100, constructing a network model using a codec module and a Transformer adaptive module, the codec module being used to process an input image to obtain a feature map, and the Transformer adaptive module being used to correct the feature map; S200, training the network model using a data set and a loss function; S300, importing the image to be segmented into the trained network model for identification to obtain a segmented image. Here, by setting a codec module and using a conventional segmentation model for training, accurate classification of common samples can be achieved. On this basis, the previously trained codec module is shared, and a Transformer adaptive module is added for parameter optimization, so that the classifier can dynamically adapt to the test sample, thereby improving the classification accuracy of the model for samples with large intra-class feature differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image recognition, and in particular to a PC screen semantic segmentation method based on an efficient attention mechanism. Background Art

[0002] At present, computer vision technology is applied to multiple scenarios, including image classification, target detection, 3D reconstruction, and semantic segmentation. With the rapid development of Internet communications, the competitiveness of intelligent products requires technological breakthroughs in more advanced semantic scene understanding. Therefore, semantic segmentation, as a core issue of computer vision, can help more and more products automatically and efficiently understand the relevant knowledge or semantics in images or videos, thereby achieving intelligent goals, reducing manual interaction operations, and improving customer comfort. Currently, these products have been widely used in autonomous driving, human-computer interaction, computational photography, image search engines, augmented reality and other fields.

[0003] The semantic segmentation problem in computer vision is essentially a process from rough reasoning to refined reasoning. At the beginning, it can be traced back to the classification problem, that is, roughly predicting the object category in the input sample, and then the positioning and detection of the target object, which not only predicts the category of the object, but also provides additional information about the spatial position of each category, such as the center point or the border of the object area. On this basis, semantic segmentation can be understood as a fine-grained prediction in the detection field. The test image is input into the segmentation network so that the predicted heat map size is consistent with the input image. The number of channels is equal to the number of categories, which represents the probability of each spatial position belonging to each category, that is, it can be classified pixel by pixel.

[0004] Deep learning algorithms are the mainstream direction of semantic segmentation technology, and have made important breakthroughs and progress. The most prominent implementation is unmanned driving technology. Although the existing semantic segmentation technology has made more and more breakthroughs in several common applications and data scenarios, there are not many studies and works on semantic segmentation of small samples and semantic segmentation of single targets in specific applications. In commercial applications, the actual implementation of semantic segmentation technology in products is mainly affected by multiple factors such as the performance of deep models, hardware, and the cost of acquiring large-scale data sets.

[0005] The fully convolutional network (FCN) has become the cornerstone of deep learning technology applied to semantic segmentation problems. It can accept input images of any size, and upsample and decode the feature map of the last convolution of the encoding network through several deconvolution layers to restore it to the same size as the input image, so that a prediction can be generated for each pixel while retaining the spatial information in the original input image. Subsequently, based on the FCN network, a variety of semantic segmentation models were derived, such as the symmetric network U-net with jump connections between the encoder and decoder, the DeepLab series of networks that introduce dilated convolutions and use conditional random fields (CRF) for post-processing optimization, and ParseNet that combines contextual information for feature fusion. These algorithm models all have the following shortcomings: First, they are overly dependent on labeled data, and the cost of obtaining data is high; second, the segmentation effect is not good for samples with large internal differences, and the generalization ability is insufficient. Summary of the invention

[0006] The purpose of the present invention is to provide a PC screen semantic segmentation method based on an efficient attention mechanism to improve the accuracy of classifying samples with large intra-class feature differences.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a PC screen semantic segmentation method based on an efficient attention mechanism, comprising the following steps: S100, building a network model using a codec module and a Transformer adaptive module, the codec module is used to process the input image to obtain a feature map, and the Transformer adaptive module is used to correct the feature map; S200, training the network model using a data set and a loss function; S300, importing the image to be segmented into the trained network model for recognition to obtain a segmented image.

[0008] Compared with the prior art, the present invention has the following technical effects: by setting the codec module and adopting the conventional segmentation model for training, accurate classification of common samples can be achieved. On this basis, the previously trained codec module is shared, and the Transformer adaptive module is added for parameter optimization, so that the classifier can dynamically adapt to the test samples, thereby improving the classification accuracy of the model for samples with large intra-class feature differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a network model diagram of the present invention;

[0010] Figure 2 It is a schematic diagram of the structure for training the encoding and decoding module;

[0011] Figure 3 This is a schematic diagram of the structure of the Transformer adaptive module training;

[0012] Figure 4 It is a model diagram of the Transformer adaptive module in the present invention;

[0013] Figure 5 It is the original image and its corresponding heat map. DETAILED DESCRIPTION

[0014] Combine the following Figures 1 to 5 , the present invention is further described in detail.

[0015] See also Figure 1 , a method for semantic segmentation of PC screen based on efficient attention mechanism, including the following steps: S100, constructing a network model using a codec module and a Transformer adaptive module, the codec module is used to process the input image to obtain a feature map, and the Transformer adaptive module is used to correct the feature map; S200, training the network model using a data set and a loss function; S300, importing the image to be segmented into the trained network model for recognition to obtain the segmented image. Here, by setting the codec module and using a conventional segmentation model for training, accurate classification of ordinary samples can be achieved. On this basis, the previously trained codec module is shared, and the Transformer adaptive module is added for parameter optimization, so that the classifier can dynamically adapt to the test sample, thereby improving the classification accuracy of the model for samples with large intra-class feature differences.

[0016] There are many kinds of network model structures composed of codec modules and Transformer adaptive modules. The present invention adopts the following scheme: in the step S100, the network model is composed of codec modules and Transformer adaptive modules connected in series, the input end of the codec module is the input end of the network model, the output end of the codec module and the Transformer adaptive module is connected to the linear classifier, the linear classifier is used to classify the feature map to obtain the heat map, and the output end of the linear classifier is the output end of the network model. In this network model, when training, it is necessary to connect the output end of the codec module to the linear classifier to facilitate the training of the codec module; when the network model is trained and put into use, the codec module does not need to be connected to the linear classifier, and it is only connected to the Transformer adaptive module.

[0017] In order to facilitate the training of the network model, two groups of sample sets are selected in the present invention, wherein the first group of sample sets are samples containing a complete screen, and the second group of sample sets are samples containing a partial screen or a tilted screen. Using two different groups of sample sets, the codec module and the Transformer adaptive module can be trained respectively. Specifically, the data set includes the first group of sample sets and the second group of sample sets. Step S200 includes the following steps: S220, using the first group of sample sets to train the encoder and decoder, and updating the network parameters of the encoder and decoder; S230, fixing the network parameters of the encoder and decoder, using the second group of sample sets to train the Transformer adaptive module, and updating the network parameters of the Transformer adaptive module. For such a network model composed of multiple modules, if it is trained directly, it will be more complicated and difficult to adjust the parameters. Therefore, in the present invention, a multi-stage training method is adopted to train the codec module and the Transformer adaptive module one by one, so that the training of the network model can be conveniently completed, and the trained network model has a very good effect on screen segmentation.

[0018] Furthermore, the data set includes a public sample set, and the following steps are included before step S220: S210, pre-training the codec module using the public sample set to initialize the parameters of the codec module. Pre-training using the public sample set here can achieve the effect of giving the model prior information, which can accelerate the convergence speed of network model training.

[0019] The above public sample set can be the PASCAL dataset. The first and second sample sets can use cameras or mobile phones to collect PC screen data with different lighting conditions and backgrounds in daily office scenes, and then use the open source tool labelme to perform pixel-level category annotation on the samples to generate the corresponding label heat map. The labels are divided into two categories: background, category 0, and screen area (without border), category 1, such as Figure 5 As shown, Figure 5 In the figure, the left is the original image, the right is the heat map (the attached image shows a black and white image, but it is actually a color image), and the gray area corresponds to the screen (the gray area in the color image is displayed in red).

[0020] See also Figure 2Furthermore, the encoding and decoding module includes an encoder and a decoder. The encoder is composed of a feature extraction network of multiple convolutional layers, pooling layers and shuffleNet Unit modules stacked together. The decoder is composed of multiple transposed convolutions and ordinary convolutional layers. The encoding and decoding module is a relatively mature network module. Its structure can refer to the description in the paper SHUFFLESEG:REAL-TIME SEMANTIC SEGMENTATION NETWORK. The training of the encoding and decoding module includes the following steps: S211, inputting the original image into the encoding and decoding module; S212, the encoding and decoding module outputs a first feature map of the same size as the original image; S213, the linear classifier processes the first feature map to obtain a first predicted heat map; S214, calculating the first loss function according to the first predicted heat map and the labeled heat map corresponding to the original image, and optimizing the network of the encoding and decoding module according to the first loss function; in step S210, all the pictures in the public sample set are used to perform steps S211-S214; in step S220, all the pictures in the first group of sample sets are used to perform steps S211-S214. Through the above steps, the codec module can be easily trained. Among them, step S210 is to pre-train the codec module and initialize the parameters of the codec module; the first set of sample sets is used to train and fine-tune the codec module. In this stage, only the network parameters of the codec module are updated, and the Transformer adaptive module is not considered. The structural diagram of its training is shown in the figure. Figure 2 When training at this stage, the output of the encoding and decoding module is directly processed by the linear classifier to obtain the first prediction heat map, and is not output to the Transformer adaptive module, nor is the Transformer adaptive module adjusted.

[0021] See also Figure 3, further, after the codec module is trained using the public sample set and the first group of sample sets, the codec module must have the ability to segment the screen, but the codec module at this time only has a good segmentation ability for pictures containing a complete PC screen. For some special cases, the segmentation effect is general. In order to further improve its segmentation ability for samples with large intra-class feature differences, we also train the Transformer adaptive module. The specific training steps are as follows: S231, fix the network parameters of the codec module, and perform the following steps S232-S234 on all pictures in the second group of sample sets in sequence; S232, input the original image into the codec module, and the codec module outputs a first feature map with the same size as the original image; S233, input the first feature map into the Transformer adaptive module, and the Transformer adaptive module outputs a second feature map; S234, the linear classifier processes the second feature map to obtain a second predicted heat map, calculates the second loss function according to the second predicted heat map and the labeled heat map corresponding to the original image, and optimizes the network of the Transformer adaptive module according to the second loss function. During this phase of training, although the image will also be processed by the codec module, the network parameters of the codec module have been fixed after the previous training step. At this stage, the network parameters of the Transformer adaptive module can be easily trained and optimized. In addition, during this phase of training, the first feature map output by the codec module is used as the input of the Transformer adaptive module.

[0022] See also Figure 4Transformer is a model architecture proposed in a 2017 paper "Attention is All You Need". This paper only conducted experiments on machine translation, which completely defeated the SOTA at the time. And because the encoder side is parallel computing, the training time is greatly shortened. Its groundbreaking idea subverted the previous idea of ​​equating sequence modeling with RNN, and is now widely used in various fields of NLP. In the present invention, the Transformer adaptive module is used to further improve the effect of semantic segmentation of the PC screen. Specifically, the Transformer adaptive module includes a query matrix, a key matrix, a value matrix, a linear mapping layer and a multi-head attention module. The first feature map is processed according to the following steps to obtain the second feature map: A. The first feature map is divided into blocks to obtain a block sample sequence; B. The block sample sequence is multiplied by the query matrix, the key matrix and the value matrix respectively to obtain new matrices Q, K and V; C. The new matrix Q is transposed and multiplied with K, and then multiplied by a constant, a softmax operation is performed, and finally multiplied by the V matrix and output to the multi-head attention module. The linear mapping layer mainly performs operations such as dot multiplication of some matrices and softmax normalization, and does not contain learning parameters; D. The multi-head attention module is composed of multiple self-attention modules, and each module extracts important features of different regions in the input sample; E. The normalization layer performs a normalization operation on the extracted matrix and performs a residual connection with the output feature map of the key matrix to obtain the second feature map.

[0023] When training a network model, we need to construct a semantic segmentation loss function so that we can tune the network parameters based on the loss function.

[0024] Cross entropy loss is a common loss function, and its formula is as follows:

[0025] ;

[0026] Among them, p represents the probability that the predicted sample belongs to category 1, the value range of p is 0-1, and y represents the label category. We can use the following formula to describe the cross entropy.

[0027] ,in, ;

[0028] Generally, the loss weight coefficient can be introduced Indicates the contribution of the ratio of positive and negative samples to the total loss, and its form is as follows: , this formula can control the weights of positive and negative samples, but cannot control the weights of easy-to-classify and difficult-to-classify samples. Therefore, in the present invention, the first loss function and the second loss function are both focalloss, and the formulas are as follows:

[0029] ;

[0030] In the formula, It is a modulation parameter used to control the weight of reducing easy-to-classify samples, so that the model can focus more on difficult-to-classify samples during training. After introducing focal loss as the loss function, the trained network model performs better when performing PC screen speech segmentation.

Claims

1. A PC screen semantic segmentation method based on efficient attention mechanism, Features: The steps include: S100, constructing a network model using a codec module and a Transformer adaptive module, wherein the codec module is used to process an input image to obtain a feature map, and the Transformer adaptive module is used to correct the feature map, wherein: The encoding and decoding module includes an encoder and a decoder, wherein the encoder is composed of a feature extraction network stacked with multiple convolutional layers, pooling layers, and shuffleNet Unit modules, and the decoder is composed of multiple transposed convolutional layers and ordinary convolutional layers; The Transformer adaptive module includes a query matrix, a key matrix, a value matrix, a linear mapping layer and a multi-head attention module; The network model is formed by connecting the codec module and the Transformer adaptive module in series, the input end of the codec module is the input end of the network model, the output ends of the codec module and the Transformer adaptive module are connected to a linear classifier, the linear classifier is used to classify the feature map to obtain a heat map, and the output end of the linear classifier is the output end of the network model; S200, training the network model using the data set and loss function; S300, importing the image to be segmented into the trained network model for recognition to obtain the segmented image.

2. The PC screen semantic segmentation method based on the efficient attention mechanism as claimed in claim 1, Features: In step S200, the data set includes a first sample set and a second sample set, and training the network model includes the following steps: S220: training the encoder and the decoder using the first sample set to update network parameters of the encoder and the decoder; S230, fixing the network parameters of the encoder and the decoder, using the second sample set to train the Transformer adaptive module, and updating the network parameters of the Transformer adaptive module.

3. The PC screen semantic segmentation method based on the efficient attention mechanism as claimed in claim 2, Features: In step S200, the data set includes a public sample set, and before step S220, the following steps are also included: S210: Pre-train the codec module using the public sample set to initialize parameters of the codec module.

4. The PC screen semantic segmentation method based on the efficient attention mechanism as claimed in claim 3, Features: The training of the encoding and decoding module includes the following steps: S211, inputting the original image into the encoding and decoding module; S212, the encoding and decoding module outputs a first feature map having the same size as the original image; S213, the linear classifier processes the first feature map to obtain a first prediction heat map; S214, calculating a first loss function according to the first predicted heat map and the marked heat map corresponding to the original image, and optimizing the network of the encoding and decoding module according to the first loss function; In step S210, steps S211-S214 are performed using all the pictures in the public sample set; in step S220, steps S211-S214 are performed using all the pictures in the first group of sample sets.

5. The PC screen semantic segmentation method based on efficient attention mechanism as claimed in claim 4, Features: The training of the Transformer adaptive module includes the following steps: S231, fixing the network parameters of the encoding and decoding module, and sequentially performing the following steps S232-S234 on all pictures in the second sample set; S232, inputting the original image into the encoding and decoding module, and the encoding and decoding module outputs a first feature map having the same size as the original image; S233, inputting the first feature map into the Transformer adaptive module, and the Transformer adaptive module outputs a second feature map; S234. The linear classifier processes the second feature map to obtain a second predicted heat map, calculates a second loss function according to the second predicted heat map and the labeled heat map corresponding to the original image, and optimizes the network of the Transformer adaptive module according to the second loss function.

6. The PC screen semantic segmentation method based on efficient attention mechanism as claimed in claim 5, Features: In step S233, the first feature map is processed according to the following steps to obtain the second feature map: A. Divide the first feature map into blocks to obtain a block sample sequence; B. multiplying the block sample sequence with the query matrix, the key matrix, and the value matrix to obtain new matrices Q, K, and V; C. Transpose the new matrix Q and multiply it with the new matrix K, then multiply it by a constant, perform a softmax operation, and finally multiply it by the new matrix V and output it to the multi-head attention module; D. The multi-head attention module is composed of multiple self-attention modules, each of which extracts important features of different regions in the input sample; E. The normalization layer performs a normalization operation on the extracted matrix and then performs a residual connection with the output feature map of the key matrix to obtain the second feature map.

7. The PC screen semantic segmentation method based on efficient attention mechanism as claimed in claim 6, Features: The first loss function and the second loss function are both focal loss, and their formulas are as follows: FL(p t )=-(1-p t ) γ *log(p t ); Where: γ is the modulation parameter, which is used to control the weight of reducing easy-to-classify samples; p represents the probability that the predicted sample belongs to category 1, the value range of p is 0-1, and y represents the label category.

8. The PC screen semantic segmentation method based on efficient attention mechanism as claimed in claim 7, Features: The first sample set is samples containing a complete screen, and the second sample set is samples containing a partial screen or a tilted screen.

Citation Information

Patent Citations

  • Remote sensing image content description method based on variational self-attention reinforcement learning

    CN111126282A

  • One-dimensional convolution position coding method of visual depth adaptive neural network

    CN112801280A