A Sketch Semantic Segmentation Method Based on Point-Line Segment Hierarchical Interaction
Through the enhanced local feature aggregation module and dense connection codec structure, combined with the point-line segment hierarchical interaction module, the problems of insufficient local feature coding and loss of semantic details in sketch semantic segmentation are solved, and high-precision sketch semantic segmentation is achieved.
Patent Information
- Application Number
- CN202211070214.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-02
AI Technical Summary
The existing sketch semantic segmentation method fails to effectively encode rich semantic information in the local feature aggregation module, does not consider the different impacts of neighboring points on the central point, fails to fully encode point-level and line segment-level features, and is prone to lose semantic details during the encoding and decoding process.
Design an enhanced local feature aggregation module, build a densely connected codec structure, build a point-level and line segment hierarchical branch network, and perform feature interaction through the point-line segment hierarchical interaction module, enhance local feature encoding using the distance information of neighboring points, and increase jump connections to improve gradient information flow.
It improves the accuracy of sketch semantic segmentation, reduces the computational complexity, enhances the ability to understand semantic information, and reduces the loss of semantic details during feature selection and extraction.
Smart Images

Figure CN115457564B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to an image semantic segmentation method and a sketch semantic segmentation method based on deep learning. Background Art
[0002] Currently, sketch semantic segmentation methods are mainly based on deep learning theory. According to the format of the data to be processed, related methods can be divided into sequence-based segmentation methods, image-based segmentation methods, and point-based segmentation methods. Among them, point-based segmentation methods have gradually attracted the attention of scholars due to their low computational complexity. For the convenience of description, the module that continuously extracts and aggregates local features during the encoding stage in the present invention is called the local feature aggregation module. Currently, most point-based segmentation methods are solved by establishing a deep neural network based on an encoder-decoder structure. Yang et al. introduced a local-global framework to better capture sketch features and proposed the SketchGNN network based on graph convolutional neural networks. Wang et al. directly used the sampled points as input and proposed using filters of different scales to better capture the sketch structure. Wang et al. proposed the multi-modal data fusion network SPFusionNet, which captures sketch features from both the image and point set modalities simultaneously. Qi et al. proposed an end-to-end learning network SketchSegNet+ by transforming the sketch semantic segmentation problem into a sequence-to-sequence generation problem, which transforms the stroke sequence of points into a semantic label sequence. The above methods all take a point set as input. However, when encoding point features, the local feature aggregation module has an obvious weakness, that is, the contribution of neighboring points to the central point is the same. The present invention designs an enhanced local feature aggregation module to enhance the positive impact of useful points and reduce the negative impact of noise points. Based on the enhanced local feature aggregation module, the present invention establishes a densely connected encoder-decoder structure. Specifically, in the encoding stage, it is composed of four enhanced local feature aggregation modules, and in the decoding part, it is composed of three multi-layer perceptrons. The encoder-decoder structure involves as much high-level semantic information as possible through dense connections. Based on the encoder-decoder structure, the present invention establishes a point-level branch network and a line-segment level branch network for encoding point-level features and line-segment level features simultaneously. At the same time, considering that some important semantic details are easily lost during the complex encoding and decoding process, the present invention designs a point-line segment level interaction module by calculating the relationship between point-level features and line-segment level features. The point-line segment level interaction module is placed between the two branch networks to enhance and supplement the semantic details of the finally fused features. To sum up, the proposed method solves several key problems that traditional methods have not solved, including: (1) The local feature aggregation module does not encode rich semantic information and does not consider the different impacts of neighboring points on the central point. (2) It is not able to fully encode point-level features and line-segment level features simultaneously for sketch semantic segmentation. (3) It does not handle the loss of important semantic details during the encoding and decoding process and lacks the interaction process between the two levels. (4) The encoder-decoder structure lacks necessary skip connections, and each decoding layer designs less high-level semantics and lacks the ability to flow gradient information across layers. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a sketch semantic segmentation method based on point-line segment hierarchical interaction. The method takes a set of points and a set of line segments as input, is simple to train, and has high accuracy. It is a method based on deep learning. The present invention designs an enhanced local feature aggregation module for extracting and aggregating local features of the sketch, which makes the "contributions" of neighboring points in the local area to the central point different. Based on the enhanced local feature aggregation module, the present invention constructs a dual-branch structure for simultaneously encoding and processing sketch semantics from the point level and the line segment level, and realizes the interaction between point-level features and line segment-level features.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is: a sketch semantic segmentation method based on point-line segment hierarchical interaction, comprising the following steps:
[0005] Step S1, acquisition of point-level data and line-segment-level data: using the farthest point sampling method to sample the sketch sample, the sketch sample is converted into a point set with N points; using the average division method to divide several points on the same stroke into a line segment, and then the sketch sample is converted into a line segment set with S line segments;
[0006] Step S2: Building a sketch semantic segmentation network: Construct a densely connected encoder-decoder structure. The encoding stage of the densely connected encoder-decoder structure consists of several enhanced local feature aggregation modules, and the decoding stage consists of several multi-layer perceptrons. Based on the densely connected encoder-decoder structure, build a point-level branch network and a line-level branch network; build a point-line-segment level interaction module, and place several point-line-segment level interaction modules between the point-level branch network and the line-level branch network for interaction between point-level features and line-segment level features.
[0007] Step S3: Training the sketch semantic segmentation network: Fusing the output of the point-level branch network and the output of the line-level branch network. The specific fusion process is as follows:
[0008]
[0009] Among them, f point and f segment They represent the output of the point-level branch network and the output of the line-level branch network, respectively. fusion represents the fusion result, is a matrix transformation operation, It is a matrix copy operation; the fusion result is optimized using cross entropy loss. The specific process is defined as follows:
[0010]
[0011] Among them, y i,jis a one-hot vector, B represents the batch size, N represents the number of sampling points, and L represents the number of semantic labels;
[0012] Step S4, obtaining the semantic segmentation result: Input the test samples of the point set represented by N points and the line segment set represented by S line segments into the trained sketch semantic segmentation network to obtain the final segmentation result.
[0013] A further improvement of the technical solution of the present invention lies in that: the densely connected encoder-decoder structure in the step S2 is constructed based on the following features: (1) The densely connected encoder-decoder structure does not perform downsampling and upsampling operations on the sampling points, and the number of sampling points remains unchanged during the encoding and decoding process; (2) The densely connected encoder-decoder structure adds necessary skip connections, that is: the first decoding layer processes the features of the penultimate and antepenultimate encoding layers, the second decoding layer processes the features of the first decoding layer and the antepenultimate third encoding layer, and the third decoding layer processes the features of the second decoding layer and the first, second, and third encoding layers.
[0014] A further improvement of the technical solution of the present invention lies in that: the enhanced local feature aggregation module in the step S2 includes four processing steps inside, including: Given a center point or a center line segment, use the information collection process C to obtain the information that needs to be encoded for adjacent points or adjacent line segments; use a multi-layer perceptron for encoding; based on two distance information, use the weight parsing process W to obtain the weight information of adjacent points or adjacent line segments for the center point and the center line segment. According to the weight information of each adjacent point or adjacent line segment, use a multi-layer perceptron to obtain the final local features of the center point or the center line segment.
[0015] A further improvement of the technical solution of the present invention lies in that: the information collection process of the information collection process C includes two steps:
[0016] (1) Use the K-nearest neighbor algorithm based on the Euclidean distance to determine the local area of the center point or the center line segment. The calculation process for obtaining the adjacent area of the center point is as follows:
[0017]
[0018] where argmin refers to obtaining the k nearest points to the center point, (x 1,i , x 2,i ) is the coordinate of the center point, and (x 1,ij , x 2,ij ) is the coordinate of any adjacent point of the center point; the center line segment consists of s points. Therefore, the center line segment is represented as a point in a 2×s-dimensional space. Therefore, the surrounding area of the center line segment can be obtained using the following formula:
[0019]
[0020] where n is the dimension of the line segment, and n is equal to 2×s;
[0021] (2) Obtain the coding information of all points or all line segments within the local area. The coded information is represented as follows:
[0022]
[0023] where c i represents the position information of the central point or the central line segment, c ij -c i represents the relative position information of the neighboring point or the neighboring line segment, f i is the intermediate stage feature of the central point, f ij -f i is the intermediate stage feature of the standardized neighboring point, o ij is the stroke information.
[0024] A further improvement of the technical solution of the present invention lies in that: in the weight parsing process W, the weight information of each neighboring point or neighboring line segment is calculated according to the distances between the neighboring points or neighboring line segments and the central point or central line segment in the coordinate space and the feature space, and then the enhanced intermediate state feature is obtained according to the weight information. The specific calculation process is as follows:
[0025]
[0026]
[0027] h' i = h i +(w' i ×h i )
[0028] where v i and are the distance vectors of the central point or central line segment in the coordinate space and the feature space respectively, ε is an extremely small number to prevent division by zero operation; is a multi-layer perceptron, is a batch normalization operation, δ is a ReLU function, h i is the intermediate state feature, h i ' is the new intermediate state feature after using the weight information.
[0029] A further improvement of the technical solution of the present invention lies in that: the point-line level interaction module is placed in several stages of the decoding part. Given the point level feature f p , whose dimension is [C, N], and given the line level feature f s , whose dimension is [C, S], the relationship calculation formula between the point level feature and the line level feature is as follows:
[0030]
[0031]
[0032] in, is a multi-layer perceptron, represents a batch normalization operation, Represents the matrix deformation operation, f p ' and f s 'Multiply, the specific calculation method is as follows:
[0033]
[0034] in, represents matrix multiplication, δ represents the softmax function, With [N, S] dimensions, apply the softmax function to The last dimension is the relationship matrix m between the points and line segments. r , and then use the relationship matrix m r The enhanced point-level features can be obtained. The detailed calculation process is as follows:
[0035]
[0036]
[0037]
[0038]
[0039] in, is the pooling operation, σ represents the sigmoid function, w c refers to the weight of all channels, The enhanced point-level features obtained include the semantics of the line-level features and strengthen the semantics of the final fusion features.
[0040] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention are:
[0041] 1. The data processed by the present invention are point sets and line segment sets. Compared with most sketch semantic segmentation methods, the present invention has the advantages of fewer parameters, simple calculation, and high feasibility, which is conducive to the promotion and use of the present invention on portable devices;
[0042] 2. The enhanced local feature aggregation module of the present invention not only encodes rich geometric features, but also considers the different effects of neighboring points on the center point based on the distance information between the neighboring points and the center point, further improving the rationality of local feature encoding;
[0043] 3. The present invention proposes a densely connected encoding and decoding structure, which involves richer semantic information and enhances the ability of gradient information to flow across layers. Based on the densely connected encoding and decoding structure, the present invention fully encodes and processes the semantic information of sketches at both the point level and the line segment level, improving the method's ability to understand semantic information;
[0044] 4. The present invention designs a point-line segment level interaction module, which enhances the overall semantics of point level features by calculating and utilizing the relationship matrix between points and line segments, further reducing the problem of semantic detail loss caused by feature selection and extraction; BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of the method of the present invention;
[0046] Figure 2 is a framework diagram of the proposed model;
[0047] Figure 3 is an internal schematic diagram of the enhanced local feature aggregation module;
[0048] Figure 4 is an internal schematic diagram of the point-line segment level interaction module. DETAILED DESCRIPTION OF THE INVENTION
[0049] The following further describes the present invention in detail with reference to the embodiments:
[0050] Figure 1 is a flowchart of the method of the present invention, and the method includes the following:
[0051] Content 1. The present invention uses the SPG and SketchSeg-150K sketch semantic segmentation datasets for training and validation. The farthest point sampling method is performed on the samples in the SPG and SketchSeg-150K datasets, and each sample is transformed into a point set. The average division method is performed on the point set, that is, several points on the same stroke are divided into a line segment, so that each sample is transformed into a line segment set.
[0052] Content 2. Build the sketch semantic segmentation network of the present invention. The sketch semantic segmentation network proposed by the present invention is as Figure 2 shown, where ELFA represents enhanced local feature aggregation, PSLI represents point-line segment level interaction, MLP represents multi-layer perceptron, N represents the number of sampled points, and S represents the number of line segments. According to Figure 2, the proposed network is based on a dual-branch structure. The dual-branch consists of two sub-networks, including a point-level network branch and a line-segment-level network branch. The basic structures of the point-level network branch and the line-segment-level network branch are the same, that is, both are constructed by a densely connected encoder-decoder structure. The encoder-decoder structure consists of 4 encoding stages and 3 decoding stages. The encoding stage is mainly implemented through an enhanced local feature aggregation module, and the decoding part is implemented by a multi-layer perceptron. During the entire encoder-decoder process, no upsampling and downsampling operations are performed on the sampling points, that is, the number of sampling points remains unchanged. At the same time, between the two network branches, the present invention places several point-line segment level interaction modules to enhance the overall semantics of the point-level features and further supplement the semantic details of the final fused features. From the above discussion, it can be seen that the key components and structures designed in the present invention are: the enhanced local feature aggregation module, the point-line segment level interaction module, and the dual-branch structure. These three parts are discussed in detail below.
[0053] Content 2-1, the enhanced local feature aggregation module is the main feature encoding and extraction module of the sketch semantic segmentation method, and its internal processing process is as Figure 3 shown. The enhanced local feature aggregation module processes point-level data and line-segment-level data. Specifically, for any input point or line segment p i Based on different dilation rates r, use the information collection process C to obtain the relevant encoded information of the neighboring points or neighboring line segments of p i ; encode the encoded information using a multi-layer perceptron to obtain the intermediate state features; use the weight parsing process for the obtained intermediate state features to obtain the weight information of each neighboring point or neighboring line segment relative to the center; use a multi-layer perceptron and combine the weight information to obtain the final intermediate stage features. According to the discussion, the enhanced local feature aggregation module mainly includes: two parts, the information collection process C and the weight parsing process W.
[0054] Content 2-1-1, the information collection process C mainly adopts two key steps. (1) Determine the local area of the center point or center line segment. (2) Obtain the encoded information within the local area. Step (1) is implemented based on the K-nearest neighbor algorithm of Euclidean distance. Assume p i (x 1,i , x 2,i ) is the center point, and p ij (x 1,ij , x 2,ij ) is any neighboring point, then the neighboring area can be expressed as R ik ={p i1 , p i2 ,... p ik} and is calculated by the following formula.
[0055]
[0056] Among them, argmin refers to obtaining the distance from the center point p i The nearest k points. At the same time, the center segment consists of s points, so the center segment can be represented as a point in the 2×s space. Center segment s i It is represented as: (x 1,i ,x 2,i ,...x 2×s,i ), any adjacent line segment s ij It is represented as: (x 1,ij ,x 2,ij ,...x 2×s,ij ), then the adjacent area of the center segment can be expressed by the following formula:
[0057]
[0058] Where n is the dimension of the line segment, which is equal to 2×s. Based on the above two formulas, the present invention determines the local area of the center point or center line segment. Furthermore, the encoded information can be defined using the following formula.
[0059]
[0060] Among them, I ij is the neighboring point p ij or adjacent line segment s ij Information collected. c i is the location information of the center point or center line segment, c ij -c i is the relative position information of adjacent points or adjacent line segments, f i is the intermediate stage feature of the center point or center line segment, f ij -f i is the intermediate state feature after standardization, o ij Is the stroke information. c i and f i The information about the attributes of neighboring points or segments is determined, while the stroke information emphasizes the stroke characteristics of the sketch. The information collection process C uses a multi-scale technique to expand the local area, that is, using different dilation rates r to achieve this goal.
[0061] Content 2-1-2. The weight parsing process is the core of the enhanced local feature aggregation module and is used to solve the weights of neighboring points or neighboring line segments with respect to the central point or central line segment. The traditional local feature aggregation module uses a multi-layer perceptron to extract local features, which ignores the different impacts of neighboring points on the central point. In fact, the closer a neighboring point or neighboring line segment is to the central point or central line segment, the greater its impact. Therefore, each point or line segment should participate in feature operations with different weights based on distance information. Generally speaking, distance information refers to the Euclidean distance based on coordinate positions. However, the distance between two points in the feature space still has important significance. Therefore, the present invention designs a weight parsing process W based on the distance in the coordinate space and the distance in the feature space. For the convenience of discussion, the present invention takes point-level data as an example for illustration.
[0062] Suppose the information collected by the information collection process C is: I i , whose dimension is: [K + 1, 4 + 2C + 1], and the intermediate state feature is: h i , whose dimension is: [K + 1, C'], the central point and any one neighboring point are: p i (x 1,i , x 2,i ) and p ij (x 1,ij , x 2,ij ), the intermediate state features of the central point and any one neighboring point are: h i (h 1,i , h 2,i , …, h n,i ) and h ij (h 1,ij , h 2,ij , …, h n,ij ), then in the coordinate space, the distance between the neighboring point and the central point is defined as follows.
[0063]
[0064] At the same time, the Euclidean distance between the neighboring point and the central point in the feature space is defined as follows.
[0065]
[0066] where n is equal to C'. Through the above two formulas, the Euclidean distance vector v i in the coordinate space and the Euclidean distance vector in the feature space can be obtained. The present invention believes that the weight information is inversely proportional to the distance information, so the following weight information formula is designed.
[0067]
[0068] where represents the concatenation operation, and ε is a very small number to prevent division by zero. Furthermore, in order to obtain the weight information of all points, a multi-layer perceptron is used for calculation. The detailed calculation process is as follows.
[0069]
[0070] in, is a multi-layer perceptron, is batch normalization, and δ is the ReLU function. After obtaining the weight information, the intermediate state features are processed using the following formula.
[0071] h' i =h i +(w' i ×h i )
[0072] Among them, + and × represent element-by-element addition and element-by-element multiplication, respectively, h i ' is the new intermediate state feature after using the weight information. The weight parsing process strengthens the useful points or line segments in the local area and weakens the noise information, which makes the enhanced local feature aggregation module more capable of expressing semantics.
[0073] Content 2-2, the present invention designs a point-line segment interaction module for the interaction between point-level features and line-segment-level features. Usually, point-level features and line-segment-level features can be fused in the final stage, but this will tend to lose some important semantic details of the intermediate stage. At the same time, point-level features represent sketches from a low level and have a stronger perception of points, while line-segment-level features have a certain continuity of strokes, so they have a better perception effect on strokes. Therefore, based on the complementary relationship between point-level features and line-segment-level features, the present invention improves the semantics of the final fused features by calculating and utilizing the relationship between points and line segments. The internal principle of the point-line segment level interaction module is as follows Figure 4 As shown. Assuming that the point level feature f p The dimension is: [C, N], the line segment level feature f s The dimension is [C, S], where C, N, and S represent the number of channels, points, and segments, respectively. To obtain the relationship between point-level features and segment-level features, we designed the following calculation process.
[0074]
[0075]
[0076] in, is a multi-layer perceptron, represents a batch normalization operation, represents the matrix deformation operation. Further, fp ' and f s ' are multiplied, and the specific calculation method is as follows.
[0077]
[0078] Among them, represents matrix multiplication, δ represents the softmax function, has the dimension of [N, S]. Applying the softmax function to the last dimension, the relationship matrix m of points and line segments can be obtained r , and its dimension is: [N, S]. The relationship matrix m r takes the line segment level feature f s ' as the base. Therefore, multiplying m r with f s ' can obtain the enhanced point level feature containing partial line segment level information. The related calculation process is defined as follows.
[0079]
[0080]
[0081] Among them, is a matrix transformation operation, is a copy operation, and f p ”' represents the new point level feature. The above process improves the semantics of the point level feature based on the association between the two level features, thereby supplementing the semantic details of the final fused feature. In addition, the point-line segment level interaction also improves the features of useful channels and suppresses the features of noise channels by capturing the dependencies between channels. Correspondingly, for this purpose, the present invention implements a channel attention mechanism. The related process is as follows.
[0082]
[0083]
[0084] Among them, represents a pooling operation, represents a multi-layer perceptron, δ is a ReLU function, σ is a sigmoid function, and w c refers to the weights of all channels, Refers to enhanced point-level features. The multilayer perceptron and ReLU function help the point-line segment hierarchical interaction module achieve nonlinear mapping across multiple channels. The sigmoid function is used as a gating mechanism to enhance the features of multiple useful channels. In the point-line segment hierarchical interaction module, this invention enhances the relationships between points and segments in the spatial dimension and the useful information in the channel dimension. By enhancing these two dimensions, the resulting fused features possess stronger semantic perception capabilities.
[0085] In Sections 2-3, the proposed method is based on a dual-branch structure and has three basic features. (1) The dual-branch structure is implemented through a point-level branch network and a line-level branch network, and the two branch networks use similar encoding and decoding structures. (2) In terms of the number of sampling points, the dual-branch structure does not perform downsampling and upsampling operations. (3) The dual-branch structure adds necessary skip connections to improve the performance of sketch semantic segmentation.
[0086] As described in Section 2-3-1, point-level branching networks and line-segment-level branching networks extract sketch features at two different levels. Point-level branching networks use point sets as the basic processing unit. As a result, feature extraction tends to be limited to a local region, and the relationships between points are relatively weak. Line-segment-level branching networks use line segments as the processing unit, which has strong continuity and a certain degree of integrity in the extracted features. Compared to single-branch structures, dual-branch structures offer greater feature complementarity.
[0087] In Section 2-3-2, the two-branch codec architecture does not perform downsampling or upsampling operations, meaning the number of sampling points remains constant from one stage to the next. Typically, downsampling aims to reduce spatial resolution to minimize computational complexity and enhance high-level semantics. Compared to images, sketches contain a large amount of blank space and a relatively small number of sampling points, so reducing the number of sampling points has a minimal impact on computational complexity. More importantly, maintaining a constant number of sampling points effectively prevents the loss of important points and a degradation in sketch semantic segmentation performance.
[0088] In content 2-3-3, in the encoding phase, the number of feature channels increases with depth, and more feature channels can describe richer semantic information. At the same time, dense skip connections facilitate the cross-layer flow of gradient information and optimize the network framework. Based on this observation, the present invention uses skip connections at each decoding layer to more likely involve intermediate stage features. This allows each decoding layer to contain more semantic information, enriching the content of the information processed at each decoding layer.
[0089] Content 3: Training the sketch semantic segmentation network of the present invention. The output of the point-level branch network and the output of the line-level branch network are fused. The specific fusion process is as follows.
[0090]
[0091] Among them, f point and f segment They represent the output of the point-level branch network and the output of the line-level branch network, respectively. fusion represents the fusion result, is a matrix transformation operation, is a matrix copy operation. Furthermore, the fusion result is optimized using cross entropy loss. The specific process is defined as follows.
[0092]
[0093] Among them, y i,j is a one-hot vector, B represents the batch size, N represents the number of sample points, and L represents the number of semantic labels. The training samples from the SPG and SketchSeg-150K datasets are fed into the proposed sketch semantic segmentation network. The model is then optimized using the Adam algorithm for 100 rounds to obtain a trained sketch semantic segmentation network.
[0094] In content 4, the test samples from the SPG and SketchSeg-150K datasets are input into the trained sketch semantic segmentation network. The dual-branch structure of the proposed method automatically fuses the point-level semantic segmentation map and the line-level sketch semantic segmentation map to obtain the final sketch semantic segmentation result.
[0095] The above-described embodiment is only one embodiment of the present invention and is used to understand the method and its core concept, but does not limit the scope of the present invention. Without departing from the design concept of the present invention, any modification and improvement made by researchers in this field to the technical solution of the present invention should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A sketch semantic segmentation method based on point-line segment hierarchical interaction, characterized in that It includes the following steps: Step S1, obtaining point-level data and line-segment-level data: Using the farthest point sampling method to sample the sketch sample, and converting the sketch sample into a point set with N points; Using the average division method to divide several points on the same stroke into a line segment, and then converting the sketch sample into a line segment set with S line segments; Step S2, building a sketch semantic segmentation network: Constructing a densely connected encoder-decoder structure, where the encoding stage of the densely connected encoder structure consists of several enhanced local feature aggregation modules, and the decoding stage consists of several multi-layer perceptrons; Based on the densely connected encoder-decoder structure, building a point-level branch network and a line-segment-level branch network; Building a point-line segment level interaction module, and placing several point-line segment level interaction modules between the point-level branch network and the line-segment-level branch network for the interaction of point-level features and line-segment-level features; The inside of the enhanced local feature aggregation module contains four processing steps, including: Given a central point or a central line segment, using the information collection process C to obtain the information that the neighboring points or neighboring line segments need to be encoded; Using a multi-layer perceptron for encoding; Based on two types of distance information, using the weight parsing process W to obtain the weight information of the neighboring points or neighboring line segments for the central point and the central line segment, and according to the weight information of each neighboring point or neighboring line segment, using a multi-layer perceptron to obtain the final local feature of the central point or the central line segment; Step S3, training the sketch semantic segmentation network: Fusing the output of the point-level branch network and the output of the line-segment-level branch network, and the specific fusion process is as follows: f fusion = f point + T(R(T(f segment ))) Among them, f point and f segment represent the output of the point-level branch network and the output of the line-level branch network respectively, and f fusion represents the fusion result. T is a matrix transformation operation, and R is a matrix copying operation; the fusion result is optimized using cross-entropy loss, and the specific process is defined as follows: where y i,j is a one-hot vector, B represents the batch size, N represents the number of sampling points, and L represents the number of semantic labels; Step S4, obtaining the semantic segmentation result: Inputting the test samples of the point set represented by N points and the line segment set represented by S line segments into the trained sketch semantic segmentation network to obtain the final segmentation result.
2. The method for sketch semantic segmentation based on point-line segment hierarchical interaction according to claim 1, wherein: The densely connected encoder-decoder structure in step S2 is constructed based on the following features: (1) The densely connected encoder-decoder structure does not perform downsampling and upsampling operations on the sampled points, and the number of sampled points remains unchanged during the encoding and decoding process; (2) The densely connected encoder-decoder structure adds necessary skip connections, that is: the first decoding layer processes the features of the penultimate and antepenultimate encoding layers, the second decoding layer processes the features of the first decoding layer and the third last encoding layer, and the third decoding layer processes the features of the second decoding layer and the first, second, and third encoding layers.
3. A sketch semantic segmentation method based on point-line segment hierarchical interaction according to claim 1, characterized in that: The information collection process of the information collection process C includes two steps: (1) Using the K-nearest neighbor algorithm based on the Euclidean distance to determine the local area of the central point or the central line segment, and the calculation process for obtaining the neighboring area of the central point is as follows: where argmin means obtaining the k points closest to the center point, (x 1,i , x 2,i ) are the coordinates of the center point, and (x 1,ij , x 2,ij ) are the coordinates of any adjacent point of the center point; the central line segment consists of s points. Therefore, the central line segment is represented as a point in a 2×s-dimensional space. Therefore, the surrounding area of the central line segment is obtained using the following formula: where n is the dimension of the line segment, and n is equal to 2×s; (2) Obtaining the encoded information of all points or all line segments in the local area, and the encoded information is represented as follows: Among them, c i represents the position information of the center point or the center line segment, and c ij -c i represents the relative position information of the adjacent point or the adjacent line segment. f i is the intermediate stage feature of the center point, and f ij -f i is the intermediate stage feature of the standardized adjacent point. o ij is the stroke information.
4. A sketch semantic segmentation method based on point-line segment hierarchical interaction according to claim 1, characterized in that: For the weight parsing process W, according to the distances between the neighboring points or neighboring line segments and the central point or the central line segment in the coordinate space and the feature space, calculating the weight information of each neighboring point or neighboring line segment, and then obtaining the enhanced intermediate state feature according to the weight information, and the specific calculation process is as follows: w i ' = δ(B(M(w i ))) h i ' = h i + (w i ' × h i ) where v i and are the distance vectors of the center point or the center line segment in the coordinate space and the feature space respectively, ε is an extremely small number to prevent division by zero; M is a multi-layer perceptron, B is a batch normalization operation, δ is the ReLU function, h i is the intermediate state feature, and h i ' is the new intermediate state feature after using the weight information.
5. A sketch semantic segmentation method based on point-line segment hierarchical interaction according to claim 1, characterized in that The point-line level interaction module is placed in several stages of the decoding part. Given the point level feature f p , whose dimension is [C, N], and given the line level feature f s , whose dimension is [C, S], the relationship calculation formula between the point level feature and the line level feature is as follows: f p ' = T(B(M(f p ))) f s ' = B(M(f s )) where M is a multi-layer perceptron, B represents a batch normalization operation, T represents a matrix deformation operation, multiplying f p ' and f s ', and the specific calculation method is as follows: Among them, represents matrix multiplication, and δ represents the softmax function. has dimensions [N, S]. Applying the softmax function to the last dimension, the relationship matrix m of points and line segments can be obtained. r Furthermore, using the relationship matrix m r the enhanced point-level features can be obtained. The detailed calculation process is as follows: f p ”' = f p ” + T(R(f s )) w c = σ(M(δ(M(P(f p ”'))))) where P is the pooling operation, σ represents the sigmoid function, and w c refers to the weights of all channels, is the obtained enhanced point-level feature, and the enhanced point-level feature contains the semantics of the line segment-level feature, strengthening the semantics of the final fused feature.
Citation Information
Patent Citations
Multi-data fusion sketch image segmentation method, system and device and storage medium
CN110853039A
Sketch recognition method based on double-layer structure
CN114373077A