Serialized Sketch Generation and Reconstruction Model Training Method, Reconstruction Method and Device

Through the training method of serialized sketch generation and reconstruction, the model is trained by using neural network and Gaussian mixed model to reconstruct the incomplete hand-drawn sketch, which solves the problem that the existing technology cannot effectively repair the incomplete sketch, and achieves a stable recovery effect.

CN115457183BActive Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210988057.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2025-05-30
Estimated Expiration
2042-08-17

AI Technical Summary

Technical Problem

The prior art cannot effectively carry out sketch reconstruction and sketch repair for serialized incomplete hand-drawn sketches.

Method used

The serialized sketch generation and reconstruction model training method is adopted. By obtaining the hand-drawn sketch dataset, selecting graph nodes, building an adjacency matrix, and reconstructing the incomplete hand-drawn sketch using neural network models (including encoder, hidden layer, and decoder) and Gaussian hybrid model.

Benefits of technology

Effective reconstruction and repair of incomplete hand-drawn sketches is achieved, and the stability of the recovery effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457183B_ABST
    Figure CN115457183B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for training a serialized sketch generation and reconstruction model, and a reconstruction method, including: obtaining a hand-drawn sketch data set, selecting graph nodes from incomplete hand-drawn sketches according to a preset method, and using a time-based nearest neighbor algorithm to connect each graph node with its nearby graph nodes according to the drawing stroke timing sequence to form an adjacency matrix; using the obtained graph nodes, adjacency matrix, and vectorized timing data of the complete hand-drawn sketch as samples to construct a training sample set; obtaining an initial neural network model, which includes an encoder module, a hidden layer, and a decoder module connected in sequence; training the initial neural network with the training sample set, constructing a loss function between the reconstructed sketch and the complete hand-drawn sketch, and iterating the parameters of the initial neural network model according to the loss function to finally obtain a serialized sketch generation and reconstruction model. The present invention can realize the generation, reconstruction, and repair of serialized incomplete hand-drawn sketches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method for training a serialized sketch generation and reconstruction model, a reconstruction method, and an apparatus. Background Art

[0002] A sketch is a highly refined and abstract form of expression. In daily communication, people can use sketches to highly summarize the abstract concepts to be expressed. For humans, even children can easily use sketches to express real objects at the semantic level. However, for machines, sketches, which are visual images with simple lines and structures, are difficult to process, and the semantics behind the sketches often change due to minor changes in the sketches. At the same time, the sketches constructed in different people's minds are very different in details, and even the sketches of the same person for the same thing are not the same. There will be deviations in the length, position, and even angle of the lines. People will consciously add some personal understandings during the process of drawing sketches and will also unconsciously add the characteristics that the object itself should have. What machines are difficult to learn is precisely this unconscious behavior of humans, and this highly abstract expression is also the current research focus and key point.

[0003] In recent years, the related research in the field of sketches is not comprehensive, and the inherent problems such as small datasets, complex data structures, and high abstraction levels have not been solved. Therefore, the existing technologies have not been greatly developed. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for training a serialized sketch generation and reconstruction model, a reconstruction method, and an apparatus to eliminate or improve one or more defects existing in the prior art and solve the problem that the prior art cannot perform sketch reconstruction and sketch repair for serialized incomplete hand-drawn sketches.

[0005] On the one hand, the present invention provides a method for training a serialized sketch generation and reconstruction model, and the method includes the following steps:

[0006] Obtain a hand-drawn sketch dataset, where the hand-drawn sketch dataset includes a plurality of data entries, and each data entry includes the vectorized time-series data of a single incomplete hand-drawn sketch and the vectorized time-series data of its corresponding complete hand-drawn sketch;

[0007] Select a first set number of graph nodes from each stroke of the incomplete hand-drawn sketch according to a preset method, and connect each graph node with a second set number of adjacent graph nodes nearby according to the stroke time series of the incomplete hand-drawn sketch by using a time-based nearest neighbor algorithm to form an adjacency matrix;

[0008] Taking each graph node corresponding to each data bar in the hand-drawn sketch dataset, the adjacency matrix, and the vectorized time-series data of the complete hand-drawn sketch as samples, a training sample set is constructed;

[0009] An initial neural network model is obtained. The initial neural network model includes an encoder module, a hidden layer, and a decoder module connected in sequence. Among them, the encoder module includes a plurality of convolutional layers, a max-pooling layer, a batch normalization layer, a node embedding module, and a plurality of fully connected layers connected in sequence. The decoder module is a long short-term memory neural network, and a Gaussian mixture model is also connected after the decoder module. The initial neural network model inputs each graph node and the adjacency matrix in a single sample into the encoder module, extracts the visual feature vector of each graph node, performs feature propagation according to the visual feature vectors of each graph node and the corresponding adjacency matrix to obtain a graph node feature set, calculates a graph embedding expression vector according to the graph node feature set and the weights of each graph node, inputs the graph embedding expression vector into the fully connected layer to obtain distribution parameters, and calculates the mean and variance. The hidden layer reconstructs a standard normal distribution according to the mean and variance and randomly samples to obtain a hidden layer vector. The hidden layer vector is input into the decoder module to obtain the distribution parameters of each stroke. The Gaussian mixture model calculates the maximum expectation for the distribution parameters of each stroke to obtain the offsets of each stroke and the connection states between the front and back strokes that make up the reconstructed sketch, so as to obtain the reconstructed sketch;

[0010] The initial neural network is trained using the training sample set, a loss function between the reconstructed sketch and the complete hand-drawn sketch is constructed, and the parameters of the initial neural network model are iterated according to the loss function, and finally a serialized sketch generation and reconstruction model is obtained.

[0011] In some embodiments of the present invention, the vectorized time-series data of the incomplete hand-drawn sketch includes the sampling order, abscissa, and ordinate of each sampling point on each stroke that makes up the corresponding incomplete hand-drawn sketch, and the connection state between each sampling point and the front and back sampling points during the generation process of the incomplete hand-drawn sketch;

[0012] The vectorized time-series data of the complete hand-drawn sketch includes the sampling order, abscissa, and ordinate of each sampling point on each stroke that makes up the corresponding complete hand-drawn sketch, and the connection state between each sampling point and the front and back sampling points during the generation process of the complete hand-drawn sketch;

[0013] The connection state includes three states represented by One Hot encoding. The first state indicates that the current sampling point and the next sampling point are adjacent strokes in the same stroke. The second state indicates that the current sampling point is the last stroke of the current stroke. The third state indicates that the current sampling point is the last stroke of the current hand-drawn sketch.

[0014] In some embodiments of the present invention, selecting a graph node from an incomplete hand-drawn sketch according to a preset method further includes:

[0015] Selecting the starting points of the strokes constituting the incomplete hand-drawn sketch as graph nodes;

[0016] And / or, the Nth turning point of each stroke constituting the incomplete hand-drawn sketch is selected as a graph node.

[0017] In some embodiments of the present invention, the calculation formula for performing feature propagation is:

[0018]

[0019] Wherein, m represents the number of nodes in the incomplete hand-drawn sketch; a ij represents the link strength between the i-th graph node and the j-th graph node, wherein the link strength gradually decreases from near to far according to the drawing time sequence of each graph node, wherein the link strength of the self-connected graph node is the highest; Represents the visual feature vector of the i-th graph node in the incomplete hand-drawn sketch.

[0020] In some embodiments of the present invention, constructing a loss function between the reconstructed sketch and the complete sketch further includes:

[0021] A reconstruction loss and a perceptual loss between the reconstructed sketch and the complete sketch are constructed, and a joint loss is constructed according to the reconstruction loss and the perceptual loss.

[0022] In some embodiments of the present invention, the function calculation formula of the reconstruction loss is:

[0023]

[0024] Among them, L recon represents reconstruction losses; represents the posterior distribution probability of the graph node data of the reconstructed sketch; p θ (S|z) represents the distribution probability of the graph node data of the incomplete hand-drawn sketch.

[0025] In some embodiments of the present invention, the function calculation formula of the perceptual loss is:

[0026]

[0027] Among them, L percep represents the perceptual loss; l represents the level of the serialized sketch generation and reconstruction model; H l Indicates the height of each layer; W l Indicates the width of each layer; w l Indicates scaling of the activation channels of each layer; and respectively represent rasterization and feature extraction from the reconstructed sketch and the incomplete hand-drawn sketch.

[0028] In some embodiments of the present invention, the functional calculation formula of the combined loss is:

[0029] L total = (1 - p mask ) * L recon + p mask * L percep ;

[0030] wherein, L total represents the combined loss; p mask represents the probability that the nodes of the incomplete hand-drawn sketch are set to zero and hidden in the adjacency matrix; L recon represents the reconstruction loss; L percep represents the perceptual loss.

[0031] On the other hand, the present invention provides a method for serial sketch generation and reconstruction, the method comprising the following steps:

[0032] Obtain an incomplete hand-drawn sketch to be processed;

[0033] Input the incomplete hand-drawn sketch into the serial sketch generation and reconstruction model in the serial sketch generation and reconstruction model training method described in any one of the above, to obtain the reconstructed sketch corresponding to the incomplete hand-drawn sketch.

[0034] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the serial sketch generation and reconstruction model training method and the serial sketch generation and reconstruction method described in any one of the above.

[0035] The beneficial effects of the present invention are at least:

[0036] The present invention provides a method for training a serial sketch generation and reconstruction model, which uses vectorized time-series data to record the stroke positions and drawing processes of incomplete hand-drawn sketches and complete hand-drawn sketches, so that the model can mine the time-series features in the sketch drawing process to reconstruct the incomplete hand-drawn sketches. By selecting graph nodes with typical stroke features from the strokes of the hand-drawn sketches and constructing an adjacency matrix based on the edge relationships of the graph nodes, the hand-drawn sketches are transformed into a graph structure expression form, and a diversified hand-drawn sketch feature system is constructed to improve the restoration effect of the incomplete hand-drawn sketches. At the same time, a neural network model for incomplete hand-drawn sketch reconstruction and repair is constructed based on the variational autoencoder architecture, and a Gaussian mixture model is introduced for reconstruction, which improves the repair ability of the incomplete hand-drawn sketches and ensures stable reconstruction effects.

[0037] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art after study of the following, or may be learned by practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by means of the structures particularly pointed out in the specification and the drawings.

[0038] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:

[0040] Figure 1 It is a schematic diagram of the steps of the serialized sketch generation and reconstruction model training method in an embodiment of the present invention.

[0041] Figure 2 It is a schematic diagram of the structure of the serialized sketch generation and reconstruction model in an embodiment of the present invention.

[0042] Figure 3 It is a schematic diagram of the time series format of the hand-drawn sketch dataset in an embodiment of the present invention.

[0043] Figure 4 It is a schematic diagram of the preprocessing process of the incomplete hand-drawn sketches in the hand-drawn sketch dataset in an embodiment of the present invention.

[0044] Figure 5 It is a schematic diagram of the comparison of the reconstruction and repair of incomplete hand-drawn sketches with different degrees of incompleteness by the serialized sketch generation and reconstruction model in an embodiment of the present invention.

[0045] Figure 6 It is a schematic diagram of the clustering of the feature spaces of each incomplete hand-drawn sketch obtained by the t-SNE algorithm in an embodiment of the present invention.

[0046] Figure 7 It is a schematic diagram of the innovative sketch generated by the serialized sketch generation and reconstruction model based on feature space interpolation in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objects, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0048] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less relevant to the present invention are omitted.

[0049] It should be emphasized that the term "comprising / including" as used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0050] Here, it should also be noted that if not otherwise specified, the term "connection" in this text can not only refer to a direct connection, but also represent an indirect connection with an intermediate.

[0051] In the following, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0052] It should be emphasized here that the step labels mentioned hereinafter are not intended to limit the order of the steps. Instead, it should be understood that the steps can be executed in the order mentioned in the embodiments, or in an order different from that in the embodiments, or several steps can be executed simultaneously.

[0053] In order to solve the problem that the prior art cannot perform sketch reconstruction and sketch repair for serialized incomplete hand-drawn sketches, the present invention provides a method for training a serialized sketch generation and reconstruction model, as Figure 1 and Figure 2 shown, the method includes the following steps S101 to S105:

[0054] Step S101: Obtain a hand-drawn sketch dataset, where the hand-drawn sketch dataset contains a plurality of data entries, and each data entry includes the vectorized time-series data of a single incomplete hand-drawn sketch and the vectorized time-series data of its corresponding complete hand-drawn sketch.

[0055] Step S102: Select a first set number of graph nodes from each stroke of the incomplete hand-drawn sketch according to a preset method, and connect each graph node with a second set number of its nearby graph nodes in the order of the strokes of the incomplete hand-drawn sketch using a time-based nearest neighbor algorithm to form an adjacency matrix.

[0056] Step S103: Use the graph nodes, adjacency matrix, and vectorized time-series data of the complete hand-drawn sketch corresponding to each data entry in the hand-drawn sketch dataset as samples to construct a training sample set.

[0057] Step S104: Obtain an initial neural network model, which includes an encoder module, a hidden layer, and a decoder module connected in sequence. The encoder module includes multiple convolutional layers, a maximum pooling layer, a batch normalization layer, a node embedding module, and multiple fully connected layers connected in sequence, and the decoder module is a long short-term memory neural network. A Gaussian mixture model is also connected after the decoder module.

[0058] Step S105: Use the training sample set to train the initial neural network, construct a loss function between the reconstructed sketch and the complete hand-drawn sketch, and iterate the parameters of the initial neural network model according to the loss function, and finally obtain a serialized sketch generation and reconstruction model.

[0059] In step S101, the hand-drawn sketch dataset uses the Quick Draw dataset, which is the largest human sketch dataset to date, providing more than 50 million vector sketches across 345 object categories. Since the Quick Draw dataset is too large, the present invention only selects some categories with certain characteristics for research and model training needs.

[0060] In some embodiments, 17 hand-drawn sketches are selected from the QuickDraw dataset and divided into three categories according to set rules. The set rules can be: 1. Both complex and simple drawings are included, such as angels and belts; 2. Instances within a category show similar global appearances and differ only in very local subtle details, such as cats and pigs; 3. Common life object categories contain multiple sub-category variations, such as buses, umbrellas, and clocks. For each category, 70,000 hand-drawn sketches are selected to form a training set for model training.

[0061] In some embodiments, for each category, 2,500 hand-drawn sketches are selected to form a test set for model testing.

[0062] Most of the hand-drawn sketches in the Quick Draw dataset are incomplete and incomplete images. By using the vectorized time series data of the original incomplete hand-drawn images, we losslessly draw the complete hand-drawn images of the incomplete hand-drawn images in the form of visual raster images, and construct a hand-drawn sketch dataset with one-to-one correspondence between incomplete hand-drawn images and complete hand-drawn images.

[0063] The data in the hand-drawn sketch dataset is hand-drawn sketch data based on time series. Each hand-drawn sketch is composed of a group of lines in time series. At the same time, a line can be represented only by the positions of the sampling points at both ends and the relationship between the sampling points (whether the sampling points are connected). Therefore, the data structure of the human sketching process can be represented by the state of the strokes and the order and offset of the strokes. Considering that the time series stroke data of the hand-drawn sketch fully expresses the human drawing process, the sketch can be efficiently stored in the form of vectorized data.

[0064] In some embodiments, Figure 3 As shown, an image expressed in the form of vectorized time series data is called a vectorized image, and SVG image is a commonly used storage format for vectorized images.

[0065] The vectorized timing data of the incomplete hand-drawn sketch includes the sampling order, horizontal coordinate, vertical coordinate of each sampling point on each stroke constituting the corresponding incomplete hand-drawn sketch, and the connection status of each sampling point with the previous and next sampling points during the generation process of the incomplete hand-drawn sketch. The vectorized timing data of the complete hand-drawn sketch includes the sampling order, horizontal coordinate, vertical coordinate of each sampling point on each stroke constituting the corresponding complete hand-drawn sketch, and the connection status of each sampling point with the previous and next sampling points during the generation process of the complete hand-drawn sketch.

[0066] In some embodiments, the connection state includes three states represented by One Hot coding, the first one indicates that the current sampling point and the next sampling point are adjacent strokes in the same stroke, and the One Hot coding is (1,0,0); the second one indicates that the current sampling point is the last stroke of the current stroke, and the One Hot coding is (0,1,0); the third one indicates that the current sampling point is the last stroke of the current hand-drawn sketch, and the One Hot coding is (0,0,1). Among them, the One Hot coding is a single hot coding, also known as a single-bit effective coding, and its method is to use an N-bit state register to encode N states, each state has its own independent register bit, and at any time, only one of them is valid.

[0067] For example, the vectorized time series data of the incomplete hand-drawn sketch is represented as (x n ,y n ,p n1 ,p n2 ,p n3 ), which consists of two parts: the stroke offset and the stroke state. n ,y n ) represents the position of sampling point n in the stroke, (p n1 ,p n2 ,p n3 ) form a One Hot vector.

[0068] In some embodiments, when sampling each stroke of an incomplete hand-drawn sketch, due to different sampling rates of hardware devices and software, the curves in the incomplete hand-drawn sketch may cause an increase in the number of sampling points. A compression algorithm is used to approximate the curves into straight lines, and the average number of sampling points of the incomplete hand-drawn sketch is controlled within 100. Controlling the number of sampling points can also make the graph structure simpler. The processing and sampling of the hand-drawn sketch data set can be implemented offline or online.

[0069] In step S102, the incomplete hand-drawn sketches in the hand-drawn sketch dataset are preprocessed, such as Figure 4 As shown, the preprocessing process includes steps S1021 to S1023:

[0070] Step S1021: Draw a complete hand-drawn sketch of each incomplete hand-drawn sketch in the hand-drawn sketch dataset.

[0071] Step S1022: Select the graph nodes of each incomplete hand-drawn sketch.

[0072] Step S1023: Generate an adjacency matrix based on the selected graph nodes.

[0073] In step S1021, the vectorized time series data of each incomplete hand-drawn sketch in the hand-drawn sketch data set is used to non-destructively draw a complete hand-drawn sketch in the form of a visual raster image. This step can also be pre-performed in step S101 and only needs to be performed once.

[0074] In step S1022, in some embodiments, a graph node is selected from the incomplete hand-drawn sketch according to a preset method, wherein the preset method includes two methods:

[0075] One is to select the starting points of each stroke constituting the incomplete hand-drawn sketch as the graph node, and the other is to select the Nth turning point of each stroke constituting the incomplete hand-drawn sketch as the graph node. Among them, the starting point and the turning point are both subsets of the sampling points.

[0076] Taking the starting point of each stroke as the graph node is a simple and easy-to-understand method. During the process of sketching, the first stroke contains very rich information, including the direction of the stroke, the position of the stroke, and the relative position with the previous stroke. Taking the turning point as the graph node is relatively difficult to understand compared to the starting point. However, through experimental verification, it is found that taking the turning point as the graph node has a better effect. For a relatively short stroke, the starting point of the stroke can represent the state of the entire stroke. However, for a relatively long stroke, if only the starting point is considered, it is easy to cause the loss of subsequent information because the subsequent stroke may have a large position offset or even a change in the stroke direction. Therefore, the Nth turning point of each stroke is used as the graph node, retaining the continuity between the turning point and the front and back points, where N can take values of 3, 4, or 5. In some embodiments, the 4th turning point of each stroke constituting the incomplete hand-drawn sketch is selected as the graph node.

[0077] Select a set of graph nodes according to a preset method, as shown in formula (1):

[0078] V=(v 1 ,v 2 ,...,v m ); (1)

[0079] Where, V represents the set of graph nodes; v m represents the mth graph node; m represents the number of graph nodes and also represents the maximum number of nodes defined by the network.

[0080] m represents the maximum number of nodes defined by the network. During the deep learning calculation process, the number of nodes in the network is fixed. However, this does not affect the scalability of the network because a larger sketch can be scaled down through scaling changes to meet the range of the number of sampling points.

[0081] In step S1023, a time-based nearest neighbor algorithm is used to construct the edge connections between graph nodes. Specifically, for each graph node v i , it is connected to the nearby graph nodes in the drawing order of the original stroke sampling point sequence to form an adjacency matrix. The adjacency matrix represents the relationship between graph nodes and is represented by a ij . Specifically, a ij represents the link strength between the ith graph node and the jth graph node. According to experience, it is found that between two closest graph nodes, a ij =0.3, between the second closest graph nodes, a ij =0.2. When there is no connection between two graph nodes, a ij =0. For self-connected graph nodes, let a ii= 0.5 to achieve regularization. According to the above values, it can be seen that the link strength gradually decreases from near to far according to the painting time sequence of each graph node, and the link strength of the self-connected graph node is the highest.

[0082] In some embodiments, the graph node v i is connected to four nearby graph nodes according to the stroke time sequence of the incomplete hand-drawn sketch, and the time sequence relationship needs to satisfy: before the two graph nodes appear as the parent node of the graph node v i and after the other two graph nodes appear as the child node of the graph node v i

[0083] In step S103, the vectorized time sequence data of each graph node, adjacency matrix, and complete hand-drawn sketch corresponding to each incomplete hand-drawn sketch in the hand-drawn sketch dataset obtained in step S102 are used as samples to construct a training sample set.

[0084] In step S104, an initial neural network model is obtained. This model is designed based on the VAE (Variational Auto-Encoder) model. The VAE model plays a crucial role in the field of unsupervised learning and is often used for dataset preprocessing, data dimensionality reduction, or feature extraction. Its structure mainly consists of two parts, namely the encoder and the decoder.

[0085] In the present invention, the structure of the initial neural network model includes an encoder module, a hidden layer, and a decoder module connected in sequence. Among them, the encoder module includes a plurality of convolutional layers, a max-pooling layer, a batch normalization layer, a node embedding module, and a plurality of fully connected layers connected in sequence. The decoder module is a long short-term memory neural network, and a Gaussian mixture model is also connected after the decoder module.

[0086] In some embodiments, the encoder module includes 6 convolutional layers, a max-pooling layer, a batch normalization layer, a node embedding module, and 2 fully connected layers connected in sequence. Among them, the kernel size of the convolutional layers is 2×2.

[0087] In some embodiments, the encoder module can also adopt other backbone networks such as Convolutional Neural Network (CNN), Graph Neural Network (GNN), or Vision Transformer (ViT).

[0088] In some embodiments, the decoder module can also adopt a Transformer-related model.

[0089] ​The initial neural network model inputs each graph node and the adjacency matrix in a single sample into the encoder module, extracts the visual feature vector of each graph node, performs feature propagation based on the visual feature vectors of the graph nodes and the corresponding adjacency matrix to obtain a set of graph node features, calculates the graph embedding expression vector based on the set of graph node features and the weights of each graph node, inputs the graph embedding expression vector into the fully connected layer to obtain distribution parameters, and calculates the mean and variance; the hidden layer reconstructs the standard normal distribution based on the mean and variance and randomly samples to obtain the hidden layer vector; the hidden layer vector is input into the decoder module to obtain the distribution parameters of each stroke; the Gaussian mixture model calculates the maximum expectation for the distribution parameters of each stroke to obtain the offsets of each stroke and the connection states between the front and back strokes that make up the reconstructed sketch, so as to obtain the reconstructed sketch.

[0090] Specifically, in some embodiments, after extracting the visual feature vector of each graph node, feature propagation is performed based on the visual feature vector of each graph node and the corresponding adjacency matrix to update the graph node characteristics. Among them, the calculation formula for performing feature propagation is shown in formula (2):

[0091]

[0092] where m represents the number of graph nodes in the incomplete hand-drawn sketch; a ij represents the link strength between the i-th graph node and the j-th graph node. The link strength gradually decreases from near to far according to the painting time sequence of each graph node. Among them, the link strength of self-connected graph nodes is the highest; represents the visual feature vector of the i-th graph node in the incomplete hand-drawn sketch.

[0093] In some embodiments, in order to perform high-dimensional transformation on the extracted features so that the network can obtain higher-dimensional information of the graph structure of the incomplete hand-drawn sketch, feature transformation is performed on the graph embedding expression vector. The specific transformation calculation formula is shown in formula (3):

[0094] h = W ⊙ G; (3)

[0095] where the calculation formulas of W and G are shown in formula (4) and formula (5) respectively:

[0096] W = (w 1 , w 2 ,..., w m ); (4)

[0097]

[0098] where h represents the graph embedding expression vector; W represents the set of graph node weights; G represents the set of graph node features.

[0099] Then, use the fully connected layer to calculate the distribution parameters of the graph embedding expression vector and calculate the mean and variance. Referring to the principle of the VAE model, the model of the present invention assumes that all samples satisfy the Gaussian distribution. Therefore, the hidden layer vector finally input to the decoder needs to be sampled in a Gaussian distribution. However, the operation of sampling cannot be differentiated in the neural network, or rather, the derivative does not exist and backpropagation cannot be performed. Therefore, it is decided to sample a random sampling value from a standard normal distribution during the model calculation process and perform the operation of magnifying the mean and variance of this random sampling value, as shown in formulas (6) and (7):

[0100] z = μ + σ ⊙ N(0, 1); (6)

[0101]

[0102] where z represents the hidden layer vector after magnifying the mean and variance of the random sampling value; μ represents the mean; σ represents the variance; N(0, 1) represents sampling a random sampling value from a standard normal distribution; W μ and W σ represent the coefficients of the linear layer; h represents the graph embedding expression vector.

[0103] Input the hidden layer vector into the decoder module to obtain the distribution parameters of each stroke. The front end of the decoder is actually an LSTM (Long Short-Term Memory Neural Network). The hidden layer vector actually cannot directly generate the strokes of the sketch, but uses the feature information of this sample carried by the hidden layer vector as the hidden layer feature of the LSTM, so as to obtain a set of Gaussian mixture sampling parameters. During the model training process, the LSTM is actually a conditional generation process, and during the model testing process, the LSTM uses the stroke parameters output at the previous moment as the input and the hidden layer vector as the middle layer input of the LSTM to finally obtain the stroke parameters at the next moment.

[0104] In the decoder module, for each stroke, calculate the parameters of each stroke according to formula (8):

[0105] [h i ; c i = LSTM forward (x i , [h i-1 ; c i-1 ); (8)

[0106] where [h i ; c i represents the hidden state and cell state output by the LSTM at the current moment; x i represents the input of the LSTM at the current moment; [h i-1 ; ci-1 represents the hidden state and cell state of the LSTM at the previous moment.

[0107] The parameters of each stroke obtained through the decoder module, but the actual stroke offset cannot be directly obtained, and the parameters need to be sampled using the Gaussian Mixed Model (GMM).

[0108] The Gaussian Mixed Model is a clustering algorithm in the field of deep learning. GMM uses the Gaussian distribution as the parameter model and is trained using the Expectation-Maximization algorithm. The Gaussian distribution is a distribution form that is almost native to nature. Therefore, in the present invention, it is considered that the stroke parameters of human drawing sketches actually conform to the Gaussian distribution.

[0109] Using the Gaussian Mixed Model to calculate the maximum expectation for the distribution parameters of each stroke obtained by the decoder module, the parameter set of each stroke constituting the reconstructed sketch is obtained, as shown in formula (9):

[0110] y i =[(∏,μ x ,μ y ,δ x ,δ y ,ρ xy ) 1 ,...,(Π,μ x ,μ y ,δ x ,δ y ,ρ xy ) M ,(q 1 ,q 2 ,q 3 )]; (9)

[0111] Among them, μ x and μ y are the mean coefficients that uniquely determine the bivariate normal distribution; δ x and δ y are the deviation coefficients that uniquely determine the bivariate normal distribution; ρ xy is the covariance coefficient that uniquely determines the bivariate normal distribution; M represents a set of parameter sets for a specific stroke of a certain sample; (q 1 ,q 2 ,q 3 ) is the set of stroke offset states predicted by the GMM model.

[0112] The parameter set in formula (9) contains two main parts, namely the set of stroke offset states and the set of stroke touch states. Obtaining a complete and definite stroke state is equivalent to a regression task for the stroke offset. Although each stroke offset is sampled from the distribution parameters, these distribution parameters are actually obtained through regression; for the touch state, as described above, the state of a stroke is divided into three states. The first is that the current touch and the next touch are adjacent touches, indicating that drawing is in progress; the second is that the current touch is the last touch of the current stroke and the next touch is a brand-new stroke, indicating the end of a single stroke; the third is that the current touch is the last touch of the current incomplete hand-drawn sketch, indicating complete end. Therefore, these three states can be regarded as a classification task. The stroke touch state can be transformed from formula (9) into formula (10) and formula (11):

[0113]

[0114]

[0115] Among them, formula (10) and formula (11) are equivalent to the rewriting of formula (9), removing the part of (q 1 ,q 2 ,q 3 ). It can be understood that the state of a point now takes values from an independent normal distribution. That is, formula (9) is the representation of a point, while formula (10) and formula (11) are the methods for calculating this point. After obtaining the stroke offsets of each stroke that makes up the reconstructed sketch and the connection states between the front and back strokes, the reconstructed sketch can be obtained by drawing in the drawing order.

[0116] In step S105, the initial neural network is trained using the training sample set, a loss function between the reconstructed sketch and the complete hand-drawn sketch is constructed, so that the reconstructed sketch generated by the model gradually approaches the complete hand-drawn sketch, and the parameters of the initial neural network model are iterated to finally obtain a serialized sketch generation and reconstruction model.

[0117] In some embodiments, constructing the loss function between the reconstructed sketch and the complete sketch further includes:

[0118] Constructing the reconstruction loss and perceptual loss between the reconstructed sketch and the complete sketch, and constructing a joint loss according to the reconstruction loss and the perceptual loss.

[0119] In some embodiments, the functional calculation formula of the reconstruction loss is as shown in formula (12):

[0120]

[0121] Among them, L recon represents the reconstruction loss; Represents the posterior distribution probability of the graph node data of the reconstructed sketch; p θ (S|z) represents the distribution probability of the graph node data of the incomplete hand-drawn sketch.

[0122] In some embodiments, the functional calculation formula of the perceptual loss is shown in Formula (13):

[0123]

[0124] Among them, L percep represents the perceptual loss; l represents the layer of the serialized sketch generation and reconstruction model; H l represents the height of each layer; W l represents the width of each layer; w l represents scaling the activation channels of each layer; and respectively represent rasterizing and extracting features from the reconstructed sketch and the incomplete hand-drawn sketch.

[0125] Among them, SqueezeNet is selected as the feature extractor of the feature map in the network structure, also called the deep perceptual feature extractor. By calculating the features obtained by passing the rasterized input image and the output image through the feature extractor respectively, and the calculation process is shown in Formula (13), the pixel-level regularization loss between the feature maps, that is, the perceptual loss, can be obtained.

[0126] Formula (13) actually takes the square of the difference of each element of the feature maps corresponding to the two samples, representing the L2 regularization loss. In the process of network training, the perceptual loss is more used to measure the visual similarity between the generated image and the input image. The main problem to be solved by the present invention is the reconstruction and repair of incomplete sketches, indicating that in the case of incomplete sketches, even if the generated image is a task of filling in the image based on the original image, the perceptual loss is still not zero. However, when calculating the perceptual loss between the generated image and the sketches of other different samples, the obtained value is relatively large (with a large deviation). Therefore, the perceptual loss is to make the visual features of the sketch still be able to match the original image as much as possible under the condition of incomplete sketches.

[0127] In some embodiments, a joint loss is constructed according to the above reconstruction loss and perceptual loss, and the functional calculation formula of the joint loss is shown in Formula (14):

[0128] L total =(1 - p mask ) * L recon + p mask * L percep ; (14)

[0129] Among them, L totalDenotes the combined loss; p mask Denotes the probability that the incomplete hand-drawn sketch nodes are set to zero and hidden in the adjacency matrix; L recon Denotes the reconstruction loss; L percep Denotes the perceptual loss.

[0130] In the combined function, the reconstruction loss and the perceptual loss each have a weight coefficient whose sum is 1. This weight adjustment coefficient enables the input incomplete hand-drawn sketch samples to adjust the focus of the combined loss function under different degrees of incompleteness. For example, when p mask = 0, it means that all graph nodes are retained. At this time, the perceptual loss is actually useless, and only the reconstruction loss is needed to guide the network to generate a sketch more similar to the original sketch; conversely, when the hand-drawn sketch is incomplete, that is, p mask ≠ 0, the perceptual loss needs to be used, and as p mask increases, that is, the degree of incompleteness increases, the weight of the perceptual loss increases.

[0131] In some embodiments, as Figure 5 shown, six category samples of apple, umbrella, butterfly, cake, sheep, and angel are selected and input into the serialized sketch generation and reconstruction model. From top to bottom, p mask steps from 10% in 20% increments, demonstrating the model's ability to repair sketches under different degrees of incompleteness of the same samples.

[0132] According to Figure 5 shown, it can be intuitively found that when p mask ≤ 50%, the model can ensure good repair ability; when p mask = 50%, although there will be some details missing in the reconstructed sketch generated by the model. Exemplarily, such as the butterfly losing its antennae and the cake becoming a single layer, the basic outer contour is still very similar to the original sketch.

[0133] It should be noted that for the category of angel, the situations when p mask = 50% and p mask = 70% are relatively special. Figure 5 Among them, when p mask = 50%, the generated angel sketch is basically deformed, while when p mask=70% is better restored when the probability of incompleteness is higher. In fact, this is because the features covered by the former are mainly the connection between the head and the body and part of the wings. After this part is incomplete, the basic features of the angel are basically gone, while the incomplete part of the latter is insignificant compared to the important features. Therefore, it can be found that the repair ability of the serialized sketch generation and reconstruction model has a certain ability to repair incomplete sketches, especially when a small amount of important features are incomplete.

[0134] In some embodiments, in order to more intuitively demonstrate the feature extraction capability of the model, the hidden layer vector of the sample is reduced to 2 dimensions and visualized using the t-SNE clustering algorithm. The t-SNE algorithm is a nonlinear spatial clustering algorithm, which is mainly used to reduce a set of high-dimensional features to 2-dimensional or 3-dimensional space according to the principal component dimension, so as to facilitate visualization by researchers. The hidden layer features of the sample extracted by the model of the present invention are visualized, such as Figure 6 As shown in the figure, it can be clearly observed that the categories with greater differences have a larger Euclidean distance in space and are farther away, such as an umbrella and an angel. The categories with smaller differences have a smaller Euclidean distance in space and are closer, such as the Great Wall and a belt. Although there are great differences in semantic expression, in the process of sketching, the images of the two are basically similar, so in the process of dimensionality reduction visualization, the Euclidean space distance of these two categories with similar shapes is very close.

[0135] In some embodiments, the above-mentioned serialized sketch generation and reconstruction model has the ability to repair partially incomplete hand-drawn sketches. If two similar or different local sketches are input at the same time, ideally, the interpolation of the two hidden layer vectors should have the key features of both, generating novel or even imaginative sketch paintings. Specifically, this creative visual operation can be described as a vector arithmetic problem of the hidden layer, and the calculation formula is shown in formula (15):

[0136] h=k*h 1 +(1-k)*h 2 ; (15)

[0137] Among them, h 1 and h 2 are two different graph embedding expression vectors (the output of the encoder); k is the coefficient weight, which indicates the occupancy ratio of the feature in the new graph embedding expression vector h.

[0138] like Figure 7 As shown, some classic examples are given, and it can be found that the present invention has good creativity. The serialized sketch generation and reconstruction model can extract key visual semantics from two local sketches and combine them to recreate a novel and interesting sketch.

[0139] The present invention also provides a method for generating and reconstructing serialized sketches, which includes the following steps S201 to S202:

[0140] Step S201: Obtain a defective hand-drawn sketch to be processed.

[0141] Step S202: Input the defective hand-drawn sketch into the serialized sketch generation and reconstruction model in any of the serialized sketch generation and reconstruction model training methods as described above to obtain a reconstructed sketch corresponding to the defective hand-drawn sketch.

[0142] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of any of the serialized sketch generation and reconstruction model training methods and the serialized sketch generation and reconstruction methods as described above.

[0143] In summary, the present invention provides a method for training a serialized sketch generation and reconstruction model. It uses vectorized time-series data to record the stroke positions and drawing processes of defective hand-drawn sketches and complete hand-drawn sketches, so that the model can mine the time-series features in the sketch drawing process to reconstruct the defective hand-drawn sketches. By selecting graph nodes with typical stroke features from the strokes of hand-drawn sketches and constructing an adjacency matrix based on the edge relationships of graph nodes, the hand-drawn sketches are transformed into a graph structure representation form, and a diversified hand-drawn sketch feature system is constructed to improve the restoration effect of defective hand-drawn sketches. At the same time, a neural network model for reconstructing and repairing defective hand-drawn sketches is constructed based on the variational autoencoder architecture, and a Gaussian mixture model is introduced for reconstruction, which improves the repair ability of defective hand-drawn sketches and ensures stable reconstruction effects.

[0144] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0145] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0146] In the present invention, features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0147] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a serialized sketch generation and reconstruction model, characterized in that, the method comprises the following steps: Obtain a hand-drawn sketch dataset, the hand-drawn sketch dataset containing a plurality of data entries, each data entry including the vectorized time-series data of a single incomplete hand-drawn sketch and the vectorized time-series data of its corresponding complete hand-drawn sketch; Select a first set number of graph nodes from each stroke of the incomplete hand-drawn sketch according to a preset method, and use a time-based nearest neighbor algorithm to connect each graph node with a second set number of adjacent graph nodes according to the stroke drawing time series of the incomplete hand-drawn sketch to form an adjacency matrix; Use the graph nodes corresponding to each data entry in the hand-drawn sketch dataset, the adjacency matrix, and the vectorized time-series data of the complete hand-drawn sketch as samples to construct a training sample set; Obtain an initial neural network model, the initial neural network model including an encoder module, a hidden layer, and a decoder module connected in sequence; wherein, the encoder module includes a plurality of convolutional layers, a max pooling layer, a batch normalization layer, a node embedding module, and a plurality of fully connected layers connected in sequence; the decoder module is a long short-term memory neural network, and a Gaussian mixture model is further connected after the decoder module; the initial neural network model inputs the graph nodes and the adjacency matrix in a single sample into the encoder module, extracts the visual feature vector of each graph node, performs feature propagation according to the visual feature vectors of the graph nodes and the corresponding adjacency matrix to obtain a graph node feature set, calculates a graph embedding expression vector according to the graph node feature set and the weights of each graph node, inputs the graph embedding expression vector into the fully connected layer to obtain distribution parameters, and calculates the mean and variance; the hidden layer reconstructs a standard normal distribution according to the mean and variance and randomly samples to obtain a hidden layer vector; inputs the hidden layer vector into the decoder module to obtain the distribution parameters of each stroke; uses the Gaussian mixture model to calculate the maximum expectation for the distribution parameters of each stroke to obtain the offset of each stroke and the connection state between the front and back strokes constituting the reconstructed sketch, so as to obtain the reconstructed sketch; Use the training sample set to train the initial neural network, construct a loss function between the reconstructed sketch and the complete hand-drawn sketch, and iterate the parameters of the initial neural network model according to the loss function to finally obtain a serialized sketch generation and reconstruction model.

2. The method for training a serialized sketch generation and reconstruction model according to claim 1, characterized in that, the vectorized time-series data of the incomplete hand-drawn sketch includes the sampling order, abscissa, ordinate of each sampling point on each stroke constituting the corresponding incomplete hand-drawn sketch, and the connection state between each sampling point and the front and back sampling points during the generation process of the incomplete hand-drawn sketch; the vectorized time-series data of the complete hand-drawn sketch includes the sampling order, abscissa, ordinate of each sampling point on each stroke constituting the corresponding complete hand-drawn sketch, and the connection state between each sampling point and the front and back sampling points during the generation process of the complete hand-drawn sketch; The connection status includes three states represented by One Hot encoding. The first state indicates that the current sampling point and the next sampling point are adjacent strokes in the same stroke. The second state indicates that the current sampling point is the last stroke of the current stroke. The third state indicates that the current sampling point is the last stroke of the current hand-drawn sketch.

3. The method for training a serialized sketch generation and reconstruction model according to claim 2, wherein, selecting graph nodes from the incomplete hand-drawn sketch according to a preset method further includes: selecting the starting points of each stroke constituting the incomplete hand-drawn sketch as graph nodes; and / or, selecting the Nth turning point of each stroke constituting the incomplete hand-drawn sketch as a graph node.

4. The method for training a serialized sketch generation and reconstruction model according to claim 1, wherein, the calculation formula for performing feature propagation is: Among them, m represents the number of graph nodes in the incomplete hand-drawn sketch; a ij represents the link strength between the i-th graph node and the j-th graph node. The link strength gradually decreases from near to far according to the painting time sequence of each graph node. Among them, the link strength of the self-connected graph node is the highest; represents the visual feature vector of the i-th graph node in the incomplete hand-drawn sketch.

5. The method for training a serialized sketch generation and reconstruction model according to claim 1, wherein, constructing a loss function between the reconstructed sketch and the complete sketch further includes: constructing a reconstruction loss and a perceptual loss between the reconstructed sketch and the complete sketch, and constructing a joint loss according to the reconstruction loss and the perceptual loss.

6. The method for training a serialized sketch generation and reconstruction model according to claim 5, wherein, the function calculation formula for the reconstruction loss is: Among them, L recon represents the reconstruction loss; represents the posterior distribution probability of the graph node data of the reconstructed sketch; p θ (S|z) represents the distribution probability of the graph node data of the incomplete hand-drawn sketch.

7. The method for training a serialized sketch generation and reconstruction model according to claim 6, wherein, the function calculation formula for the perceptual loss is: Among them, L percep represents the perceptual loss; l represents the level of the serialized sketch generation and reconstruction model; H l represents the height of each layer; W l represents the width of each layer; w l represents scaling the activation channels of each layer; and respectively represent rasterizing and extracting features from the reconstructed sketch and the incomplete hand-drawn sketch.

8. The method for training a serialized sketch generation and reconstruction model according to claim 7, wherein, the function calculation formula for the joint loss is: L total = (1 - p mask ) * L recon + p mask * L percep ; Among them, L total represents the combined loss; p mask represents the probability that the incomplete hand-drawn sketch node is set to zero and hidden in the adjacency matrix; L recon represents the reconstruction loss; L percep represents the perceptual loss.

9. A method for generating and reconstructing a serialized sketch, wherein, the method includes the following steps: obtaining an incomplete hand-drawn sketch to be processed; inputting the incomplete hand-drawn sketch into the serialized sketch generation and reconstruction model in the method for training a serialized sketch generation and reconstruction model according to any one of claims 1 to 8 to obtain a reconstructed sketch corresponding to the incomplete hand-drawn sketch.

10. A computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the steps of the method for training a serialized sketch generation and reconstruction model and the method for generating and reconstructing a serialized sketch according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • System and method for exchanging isomerism CAD model data based on genetic algorithm

    CN103793535A

  • Method for generating indoor house type 3D data through radar ranging

    CN111308495A