Online handwriting sequence generation system and method based on diffusion model, and storage medium

Through an online handwriting generation system based on a diffusion model, high-quality handwriting sequences that conform to individual writing styles are generated using the diffusion model backbone neural network and information processing module. This solves the problems of unstable generation process and insufficient diversity in existing technologies, and achieves efficient handwriting data generation and recognition model training.

CN120599097APending Publication Date: 2025-09-05CHONGQING AOXIONG INFORMATION TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510684483.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have difficulty generating high-quality and diverse online handwriting sequence data, especially unable to effectively capture the characteristics of individual writing habits. Moreover, the generation process is unstable, making it difficult to meet the requirements of high-precision handwriting recognition models.

Method used

An online handwriting generation system based on a diffusion model is adopted. Through the diffusion model backbone neural network, content information processing module and style information processing module, the Markov chain is used to construct the diffusion process, gradually add and remove noise, optimize the neural network parameters, and generate a handwriting sequence that conforms to the individual's writing style.

Benefits of technology

Generate high-quality, smooth and realistic online handwriting sequences that conform to human writing habits, reduce data collection costs, and improve the training data quality of handwriting recognition models. It is suitable for fields such as electronic signatures, calligraphy practice, and artistic font design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599097A_ABST
    Figure CN120599097A_ABST
Patent Text Reader

Abstract

The invention discloses an online handwriting generation system based on a diffusion model, which comprises a backbone neural network, a content information processing module and a style information processing module, and is characterized in that online handwriting sequence data is collected to carry out optimization training on the diffusion model, and the backbone neural network carries out diffusion processing on the collected online handwriting data in the diffusion process; noise data approximate to Gaussian distribution are obtained, and noise in the diffusion process is estimated; in the denoising process, noise in the sampling random number is gradually removed based on estimated noise, basic handwriting data is obtained, the content information processing module and the style information processing module respectively obtain content features and style features of the online handwriting data, and the content features and the style features are synthesized into the online handwriting with the basic handwriting data. High-quality handwriting can be generated, and the method can be used for model training and testing application needing a large number of different styles of handwriting samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a high-quality online handwriting sequence data generation technology based on a diffusion model. Background Art

[0002] High-quality datasets are crucial for building effective models for handwritten electronic signature recognition. Handwritten databases typically contain numerous images of handwritten digits and their corresponding labels, providing essential training material for machine learning models. With this data, the models can identify different handwriting styles and achieve high-precision digit and signature recognition.

[0003] High-quality datasets can significantly improve a model's recognition accuracy and generalization capabilities. By training on a large and diverse set of handwritten digit samples, the model can better adapt to different writing styles and font variations, thereby improving its performance in electronic signature recognition and applications. Handwritten digit databases provide researchers with a platform for improving handwritten digit recognition algorithms. By continuously optimizing algorithms to adapt to these data, the advancement of artificial intelligence technology in the field of handwritten signature recognition can be promoted. Handwritten digit recognition has a wide range of real-world applications, such as automatic postal code sorting systems, bank check reading, and interactive interfaces for smart devices. High-quality datasets help achieve efficient and accurate recognition in these scenarios.

[0004] Handwritten handwriting training datasets can be obtained in a variety of ways. MNIST, for example, is the primary source of data for training handwritten digits, and the data volume is relatively large. However, the data required for training artificial intelligence requires not only large quantities but also diversity. Everyone's handwriting is unique. For example, different people will write the same numbers or words in different ways and styles. If a computer is trained solely on a single person's handwriting, it will only develop a digit recognition system that is unique to that person.

[0005] Traditional online handwriting data collection requires the use of devices such as handwriting devices, smart tablets, and smartphones, and the recruitment of large numbers of people to write on the device screens to obtain online handwriting data. This handwriting data contains rich information about the writer's biometrics and behavioral habits, but it is difficult to legally and effectively collect handwriting through signature collection. Therefore, it is crucial to acquire diverse handwriting training data that reflects the writer's handwriting style.

[0006] For example, in the publication number "CN117057310A", named "A font generation method and device based on diffusion model", the target font image is blurred with Gaussian noise to obtain an image with Gaussian distribution noise; the noise image with Gaussian distribution is processed by the reverse diffusion process to generate a font image with the target style by gradually eliminating the noise.

[0007] Publication number CN118365744A, titled "A Single-Sample Font Generation Method Based on Multi-Scale Style Fusion and Glyph Control," takes a text style image and content image as input, extracts features from the input style image, and collects the output of each layer of the encoder network to obtain multi-scale style features. The content image and noise image are then channel-concatenated and used as input to a noise prediction model. The trained model is then sampled to generate a font image with the target style and content. Controlling the glyph shape with the content image improves the accuracy of the character structure. Integrating multi-scale style into the diffusion model allows for more comprehensive preservation of style image information.

[0008] Publication No. CN118762103B, a single-sample handwritten text copying method based on a diffusion model. This method involves constructing a diffusion model generation network capable of copying any handwriting style, including a style feature enhancement module, a content encoder, an adaptive fusion module, and a conditional diffusion model. A handwriting sample image and a standard font image are used as style input and content input, respectively. The style features and content features are extracted through the content encoder and style encoder, respectively. Both the style and content features are then simultaneously input into the conditional diffusion model to generate handwritten text with the target style and content. The diffusion model generation network capable of copying any handwriting style is trained. The trained diffusion model generation network is then used to generate handwritten text that satisfies both the target style and content.

[0009] The aforementioned diffusion model-based technology still generates text images, but the generated images are static rather than sequential data of individual handwritten signatures with different characteristics, so their use cases are limited. The training model and generation target are font images, which are relatively simple. Fonts are often images with neat structure and clear strokes. Handwritten handwriting, on the other hand, is more chaotic and casual, with examples such as cursive and connected strokes, and the speed and strength with which each person imitates the same text varies. In addition, handwriting generated based on autoregressive models or variational autoencoders is disjointed and lacks the naturalness of writing, resulting in low quality and diversity. The training process of GAN-based models is more difficult, and is prone to model collapse or difficulty in convergence.

[0010] In summary, the existing technology for generating Chinese fonts aims at font images. Fonts often have neat structures and clear strokes. The goal is relatively simple and the generation method is easy to implement. The generated handwriting is a static image rather than a data sequence, and does not contain dynamic handwriting sequence feature data that conforms to the characteristics of personal writing habits. Its usage scenarios are limited, and it cannot obtain dynamic information of the handwriting sequence during writing. It cannot meet the requirements of high-precision handwriting recognition models for sample data sets, and cannot meet the requirements for various fonts and styles in Chinese character writing. For example, it brings difficulties in the restoration demonstration of various font handwritings in calligraphy teaching, artistic font design, and virtual writing experience. Summary of the Invention

[0011] In order to overcome the above-mentioned deficiencies in the prior art, the present invention provides an online handwriting generation system based on a diffusion model.

[0012] Based on the first aspect of the present invention, an online handwriting handwriting generation system based on a diffusion model is proposed, including: a diffusion model backbone neural network, a content information processing module, and a style information processing module. The diffusion model backbone neural network performs a series of diffusion processing on online handwriting handwriting sequence data to obtain pure noise with an approximate Gaussian distribution, and estimates the noise in the diffusion process; the neural network is trained using the online handwriting handwriting sequence data, and the parameters of the backbone neural network denoiser, content information processing module, and style information processing module are optimized. The denoiser removes the noise component in the random noise to obtain basic handwriting data, and the content information processing module and the style information processing module respectively obtain the content features and style features of the online handwriting sequence data. The trained neural network fuses the content features and style features of the basic handwriting data to generate handwriting.

[0013] Further preferably, the backbone neural network model predicts noise based on noise loss, the diffusion model performs noise processing on the online handwriting sequence data according to the predicted noise to obtain a series of noisy data, and the denoiser gradually removes random noise to obtain a series of denoised data; the diffusion model is optimized based on the noise estimation style loss and content loss, and the parameters of the denoiser, content information processing module and style information processing module in the diffusion model are updated.

[0014] Further preferably, a diffusion process of online handwriting sequence data is constructed using a Markov chain, and a series of noisy data is generated by adding estimated noise to the online handwriting sequence data at each time step, with the noisy data at each time step serving as the intermediate state of the Markov chain; the reverse denoising process generates basic handwriting data by gradually removing the estimated noise at each time step based on random noise.

[0015] Further preferably, the backbone neural network predicts the noise estimate for the next time step based on the noisy data, corresponding content information and style information at any time step in the forward diffusion process; the diffusion model is trained using the noise estimation loss, the difference between the estimated noise and the real noise is measured by the loss, the extracted content features and style features are constrained and optimized using the content feature loss and the style feature loss, and the point features are constrained and optimized using the feature distribution consistency loss; and the parameters of the noise estimation module, the information processing module and the style processing module are updated using stochastic gradient descent.

[0016] Further preferably, the diffusion model training uses a mean square error loss function to optimize noise estimation, uses content feature loss and style feature loss to train the content information processing module and the style information processing module, constrains and optimizes the extracted content features and style features, and establishes a final loss function for neural network optimization, specifically including, according to the formula:

[0017]

[0018] Determine the loss L of the neural network, where and Labels representing content information and style information respectively, Represents the output of the content information processing module, represents the output of the style information processing module, where P(x) and Q(x) are the distribution of a certain feature X of the collected online handwriting sample and the distribution of the feature X of the generated handwriting, D KL (P||Q) is the loss function for neural network optimization.

[0019] According to the second aspect of the present application, an online handwriting handwriting generation method based on a diffusion model is proposed, including: a diffusion process, an inverse denoising process, and an inference process. The diffusion process: gradually adding noise to the online handwriting sequence data until all sequence element values ​​of the handwriting sequence data obey the state of random Gaussian distribution; the inverse denoising process: gradually removing the noise in the random sequence through continuous iteration to generate basic handwriting data; the training process: using online handwriting data and noise estimation to train the diffusion model, content information processing module, and style information processing module to optimize network parameters; the fusion process: the backbone neural network fuses the basic handwriting data, content features, and style features to generate online handwriting sequence data.

[0020] Further preferably, during the diffusion process, the backbone neural network performs diffusion processing on the online handwriting sequence data to obtain noise with an approximate Gaussian distribution, estimates the noise added at each time step during the diffusion process, optimizes the diffusion model based on the noise estimation style loss and content loss, updates the parameters of the denoiser in the diffusion model, and the denoiser gradually removes the noise in any random sequence based on the estimated noise to obtain basic handwriting data.

[0021] Further preferably, a diffusion process of online handwriting sequence data is constructed using a Markov chain, and a series of noisy data is generated by adding estimated noise to the online handwriting sequence data at each time step, and the noisy data at each time step serves as the intermediate state of the Markov chain; the reverse denoising process gradually removes the estimated noise at each time step for the random sequence to generate basic handwriting data.

[0022] Further preferably, estimated noise is added to the online handwriting sequence data, and the noise data of the next time step is calculated based on the noise data of the previous time step, specifically including: t-1 , calling the formula:

[0023]

[0024] Calculate the noise data x at time step t t , where ε is the estimated noise, α t is a hyperparameter, and the hyperparameter α in (0, T) times of noise addition is constructed linearly according to the number of noise addition operations T. t .

[0025] Further preferably, the inverse denoising process is constructed using a Markov chain, the mean of the inverse denoising process distribution is determined according to the estimated noise predicted by the neural network, and the denoised data of the current time step is determined based on the denoised data of the previous time step, specifically including: based on the denoised data of the t-th time step, according to the formula:

[0026]

[0027] Calculate the denoised data after removing the noise in step t-1 to obtain the denoised data for each time step, where ε θ (x t ,t) is the predicted noise estimate, σ t is the standard deviation of the inverse denoising process distribution, is the mean of the distribution of the inverse denoising process.

[0028] Further preferably, if the inverse denoising process does not follow the Markov chain, based on the standard deviation of the inverse denoising process distribution, according to the denoised sequence data at time step t, the formula is called:

[0029]

[0030] Calculate the denoised data of any time step τ before [0, t), σ t is the standard deviation of the inverse denoising process distribution.

[0031] Based on the third aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is proposed, wherein the computer instructions are used to enable the computer to execute the method according to any one of the above items.

[0032] The present invention constructs an online handwriting generation method based on a diffusion model, which realizes the ability to gradually generate high-quality Chinese handwriting sequences from randomly initialized Gaussian noise. The method also supports users to input style information and content information to control and constrain the generated handwriting sequences.

[0033] This invention generates high-quality online Chinese handwriting (online handwriting is a sequence that includes handwriting coordinates, chronological order, and writing status, such as whether the pen is placed), rather than handwriting images. This handwriting generation technology allows users to customize their own writing style for use in electronic signatures, calligraphy practice, and document annotation, showcasing their individuality. The generated handwriting accurately simulates the user's actual writing style, providing a natural and smooth online handwriting effect. Furthermore, the generated diverse handwriting sequences can be used as training data for deep learning models, improving their performance and accuracy in handwritten electronic signature recognition or classification tasks.

[0034] The online Chinese handwriting generation method based on the diffusion model of the present invention ultimately generates a smooth and realistic online handwriting sequence with similar style and readable content, which conforms to the characteristics of human writing habit trajectory. At the same time, it can provide personalized input information (for example, input a user's real writing style, or specify the generated handwriting content), so that the generated handwriting can accurately meet the user's specific needs. It overcomes the problem that the static image of fonts or handwriting generated by the diffusion model method of the prior art cannot fully display the dynamic change process of the handwriting sequence. The online handwriting generation method based on the adversarial generative network has a relatively unstable training process and the network model is difficult to converge. The method based on the autoregressive model or the variational autoencoder is usually difficult to balance the quality and diversity of the generated handwriting, and performs insufficiently when processing the conditional input specified by the user.

[0035] At the same time, this invention can effectively reduce the cost of online handwriting data collection and improve its efficiency. Traditional online handwriting data collection requires the use of handwriting devices, smart tablets, smartphones, and other devices, as well as recruiting a large number of people to write on the devices to obtain online handwriting data. However, this invention uses an algorithmic model to automatically generate a large amount of rich and diverse handwriting data, reducing collection costs.

[0036] Secondly, this invention can provide training data for online handwriting recognition and comparison, reducing the development and application costs of online handwriting recognition and comparison models. Online handwriting recognition models require a large amount of high-quality training data. When actual data is insufficient or data diversity is limited, this invention can generate high-quality handwriting for training and testing applications. This invention is also applicable to fields such as handwritten electronic signatures or handwriting, artistic font design, and virtual writing experiences. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the process of online handwriting generation based on a diffusion model in an exemplary embodiment of the present invention;

[0038] Figure 2 A schematic diagram of an online handwriting generation model framework based on a diffusion model in an exemplary embodiment of the present invention;

[0039] Figure 3 The present invention is a schematic diagram of the structure of an electronic device for implementing the method of the present invention. DETAILED DESCRIPTION

[0040] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0041] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0042] This application aims to generate online handwritten character sequence data using a diffusion model. Forward noise addition and backward denoising training are used to optimize the parameters of the diffusion model's noise estimator, content information processing module, and style information processing module. The trained model fuses character type and style type with input random noise to generate new character sequences of specified character types and style types. This invention utilizes data preprocessing, a backbone network, and task supervision to process online handwriting sequences. This differs from existing techniques, which are limited to generating character graphics based solely on image data and lack data sequences that capture the characteristics of individual handwritten signatures.

[0043] The online handwriting generation model based on the diffusion model includes a backbone neural network, a content information processing module, and a style information processing module. The backbone neural network (diffusion model) is optimized and trained using online handwriting sequence data. The online handwriting data is then denoised to generate pure noise that approximates a Gaussian distribution, thereby estimating the noise during the diffusion process. The noise in the random sequence data is removed to obtain basic handwriting data. The content information processing module and the style information processing module respectively obtain content features and style features of the online handwriting data, which are then synthesized with the basic handwriting data to generate handwriting. The neural network includes a noise estimator (denoiser) for the diffusion model, a handwriting content processing model, and a handwriting style processing model. The backbone neural network serves as the noise estimator for the diffusion model.

[0044] Using the above online handwriting generation model, online handwriting generation is achieved based on the diffusion model. This includes:

[0045] Collect and obtain online handwriting data to obtain original handwriting sequence data. In order to facilitate the processing of the diffusion model and improve the efficiency and accuracy of model training, the handwriting data can be further preprocessed to obtain standardized original handwriting sequence data;

[0046] The backbone neural network is trained, and the training process includes a diffusion process and an inverse denoising process. Typical methods for the diffusion and inverse denoising processes include: The diffusion process continuously adds noise to the original handwriting sequence data until all sequence element values ​​in the handwriting sequence data follow a random Gaussian distribution. The inverse denoising process inputs the random noise sequence data and the content and style features obtained by the content information processing module and the style information processing module. The denoiser (i.e., the noise estimator) removes noise from the random sequence data through continuous iteration to generate basic handwriting data.

[0047] During the training process, the diffusion model is optimized based on an iterative process of a customized loss function, and the parameters of the noise estimator, the handwriting content processing model, and the handwriting style processing model in the diffusion model are updated to complete the training of the diffusion model; the customized loss functions include: noise estimation loss, handwriting style loss, content loss, online handwriting sequence distribution loss, etc.

[0048] In the inference process, based on the trained diffusion model, content features, style features, and random noise sequences are input. The denoiser removes noise from the random sequence data through continuous iteration to obtain basic handwriting data. The content information processing module and the style information processing module obtain content features and style features from the online handwriting data respectively. The backbone neural network fuses the content features and style features in the basic handwriting data to generate online handwriting.

[0049] like Figure 1The figure shows a schematic diagram of the online handwriting generation process based on the diffusion model in this exemplary embodiment. The online handwriting generation model based on the diffusion model includes: a backbone neural network, a content information processing module, and a style information processing module. The backbone neural network performs noise processing on the online handwriting data (forward noise adding process) to obtain noise with an approximate Gaussian distribution, and gradually estimates the noise in the diffusion process; the diffusion model is optimized and trained using the online handwriting sequence data, and the noise in the sequence data is removed to obtain the basic handwriting data. The content information processing module and the style information processing module respectively obtain the content features and style features of the online handwriting data, and synthesize them with the basic handwriting data to generate handwriting.

[0050] In order to train the model parameters, the present invention can use the Markov chain to construct the forward noise addition process of the online handwriting data, that is, the diffusion process; the online handwriting data is still Gaussian distributed after forward noise addition, and the reverse denoising process samples the random Gaussian distribution sequence data to obtain the basic handwriting data, and the basic handwriting data is generated through the above-mentioned reasoning process, which is also the handwriting sequence data.

[0051] This exemplary embodiment illustrates the training of an online handwriting generation model based on a diffusion model.

[0052] First, online handwriting data (character data) is collected using handwriting devices and smart tablets. This data includes the signature coordinates (x, y), pen pressure, point-to-point time difference, pen state (such as pen lift and pen down), and handwriting sequence time. To enable the diffusion model to more effectively process handwriting sequence data, the collected online handwriting data can be preprocessed. This includes normalizing the handwriting point coordinates (x, y) to map them to a uniform coordinate range and assigning specific markers to discrete pen states. Pen states can be represented by different numerical values, such as -1 and 1 (pen down is -1, pen lift is 1). After adding noise to the handwriting data, a threshold of 0 can be used to distinguish between pen down and pen up states. Pen pressure and point-to-point time difference are also normalized. If the original data lacks point-to-point time difference and pen pressure information, only the [x, y, pen state] information can be used. Finally, the handwriting data is uniformly sampled to a fixed length and preprocessed to obtain standardized handwritten electronic handwriting sequence data to meet the input requirements of the diffusion model. If the handwriting data is online handwritten signature data or other text line data, it can be processed through character segmentation before being used for training or generation.

[0053] The backbone network of the diffusion model, i.e., the noise estimator, can adopt a one-dimensional convolutional neural network, transformer, LSTM, TDNN or a combination thereof to support the sequence signal input network model.

[0054] like Figure 2The figure shows a schematic diagram of the online handwriting generation model framework based on the diffusion model in this exemplary embodiment. It includes a backbone neural network, a content information processing module, and a style information processing module. The backbone neural network is trained through noise addition and denoising, and noise is estimated to generate handwriting data.

[0055] Online sampling is used to obtain the initial handwriting sequence of the handwriting data sequence xtorget (the handwriting data sequence x0 is obtained at time step t0). The diffusion process performs noise processing on the collected handwriting data sequence in sequence to obtain a series of noisy data, such as adding noise to the data sequence of the previous time step at each time step to obtain the noisy data of each time step, until pure noise with an approximate Gaussian distribution is obtained; the noise estimator estimates the noise, and the denoising process performs a series of denoising processing on the random noise signal in sequence to obtain a series of denoised data, such as removing the estimated predetermined noise from the data sequence of the previous time step at each time step to obtain the denoised data of each time step; the noisy or / and denoised data sequence is input into the backbone neural network model, the noise is estimated in the prediction data based on the noise loss, the diffusion model is optimized based on the noise estimated style loss and content loss, the parameters of the denoiser in the diffusion model are updated, the training of the diffusion model is completed, the content information processing module is trained based on the content loss function, and the style information processing module is trained based on the style loss function.

[0056] The neural network inputs include handwriting sequence data, content information, and style information. The neural network comprises a content information processing module, a style information processing module, and a diffusion model backbone neural network. The diffusion model backbone neural network must be designed to ensure that the dimensions of the input sequence data match the dimensions of the output sequence data. The backbone neural network is used to estimate the noise in the diffusion process.

[0057] The neural network model includes a content information processing module, a style information processing module and a diffusion model backbone neural network. The neural network model is trained to optimize the parameters of the noise estimator, content information processing module and style information processing module.

[0058] The content information processing module encodes the online handwriting sequence into a 1*d1-dimensional content feature vector using an online sequence model. The style information processing module encodes the online handwriting sequence into a 1*d2-dimensional style feature vector using an online sequence model. The online sequence model can use a one-dimensional convolutional neural network, transformer, LSTM, TDNN, or a combination thereof that supports sequential signal input into the network model. The content information processing module can be trained before or alongside the diffusion model backbone neural network.

[0059] The diffusion model backbone neural network estimates the noise at each sequence point dimension. It inputs the handwritten electronic handwriting sequence data, content feature vectors, and style feature vectors. The handwritten electronic handwriting sequence data is fused with the content and style features at the point feature dimension, and the output is passed through the backbone neural network. Fusion at the point feature dimension can be performed using methods such as concat([x, y, s], content feature vector, style feature vector) or add(mlp([x, y, s], mlp(content feature vector), mlp(style feature vector)).

[0060] The forward noising process and reverse denoising process in the diffusion model are constructed. This embodiment uses a Markov chain to construct the noise addition and denoising processes of the diffusion model. The diffusion model generates a series of intermediate states by adding noise to the data. The noisy data at each step serves as the intermediate state of the Markov chain. Reverse denoising is then used to gradually remove the noise from the random data sequence and generate the target data.

[0061] The diffusion model backbone neural network is used to estimate noise. Markov chain theory requires that the estimated noise cannot be too large, so a multi-step iterative estimation of noise is required.

[0062] In implementation 1, the forward diffusion process involves gradually adding random Gaussian noise to the handwriting data sequence. After a sufficient number of operations, it eventually evolves into pure noise with an approximate Gaussian distribution. Pure noise with an approximate Gaussian distribution is obtained by performing T noise addition operations. For example, the optimal setting of the minimum number of noise addition operations T for adding random Gaussian noise is 1000. The degree of noise addition at each time step α is set. t is a hyperparameter. According to the noisy handwriting data obtained after the noise addition in the previous time step and the added random Gaussian noise, the noisy handwriting data of the current time step is calculated.

[0063] For example, according to the (t-1) time step noise sequence handwriting data x t-1 , calling the formula:

[0064]

[0065] Calculate the noisy handwriting data x after the t-th step of noise addition t , where ε is random Gaussian noise, x t-1 is the sequence data after adding noise in step t-1, α t is a hyperparameter (indicating the degree of noise addition at each time step), and the hyperparameter α within (0, T) times of noise addition is constructed linearly according to the number of noise addition operations. t , such as T = 1000, set α0 = 0.9999, α T =0.9998. The handwriting data at any time step includes the point sequence data at the current time step, such as x t =[[xt ,y t ,p t ,s t ]].

[0066] During the diffusion model training process, the reverse denoising process is exactly the opposite of the forward denoising process. The forward denoising process obtains pure noise with an approximate Gaussian distribution. The reverse denoising process, including the inference process, iteratively removes noise from the new random Gaussian noise to obtain the random Gaussian noise of the process, train the model, and optimize the network parameters.

[0067] After T times of denoising, the real online handwriting sequence is restored. The backbone neural network model of the diffusion model is used to estimate the noise size at each time step t, using ε θ (x t ,t) represents the noise estimate at time step t predicted by the neural network. The distribution of the inverse denoising process conforms to the Gaussian distribution. Therefore, according to the diffusion process, the mean and variance of the inverse denoising process are solved, and sampling from the Gaussian distribution is performed to achieve the purpose of step-by-step iterative denoising.

[0068] Through the training process, the input handwriting sequence is restored; through the inference process, a new handwriting sequence is generated.

[0069] The distribution of the inverse denoising process conforms to the Gaussian distribution. Therefore, according to the diffusion process, the mean and variance of the inverse denoising process are solved, and sampling from the distribution is performed to achieve the purpose of step-by-step iterative denoising.

[0070] Online handwriting data and noise estimation are used to train the diffusion model, content information processing module, and style information processing module to obtain the network optimization parameters of each module.

[0071] For example, the standard online handwriting data sequence obtained after preprocessing is recorded as x0. Through the forward diffusion process, a series of noise-added data of time steps are obtained. For any input data x at the tth time step, t , input into the neural network, at the same time, input content information and style information, after the standard online handwriting data sequence is fused with content features and style features, it is input into the backbone neural network of the diffusion model, and the output is the sequence noise data ε that combines content features and style features θ (x t ,t). Among them, the noise estimate ε at any time step t θ (x t ,t) dimension is the same as the dimension of the collected online handwriting data x0.

[0072] The mean square error loss function (MSE) is used in diffusion model training to optimize the noise estimator and calculate the weighted mean square error loss L of the feature sequence according to the formula noise :

[0073]

[0074] Among them, j represents the dimension index of the generated handwriting sequence point [such as: x, y, p, s], w j represents the weight coefficient of the j-th dimension, Represents the data value of the jth dimension at time step t, Represents the data value of the noise estimator input to the jth dimension at time step t prediction noise.

[0075] The above feature sequence weighting can enhance the learning of sequence stroke attribution.

[0076] Content feature loss and style feature loss are used in training to optimize the content information processing module and style information processing module. This type of loss is based on the classification loss function or the comparison loss function to constrain and optimize the extracted content features and style features. For example, according to the formula:

[0077]

[0078] Calculate content feature loss L content and style feature loss L style .in, and Labels representing content information and style information respectively, Represents the output of the content information processing module, represents the output of the style information processing module, and b represents the number of training set samples.

[0079] Taking into account the distribution of features within the strokes in the online handwriting sequence (such as position coordinate increments Δx or Δy), it is necessary to constrain the generated handwriting to have a similar feature distribution pattern. Assuming that P is the feature distribution of the training handwriting and Q is the feature distribution of the generated handwriting, the constraint is performed through the consistency loss of the feature distribution within the handwriting.

[0080] According to the formula:

[0081]

[0082] Establish a neural network to optimize the loss function D KL (P||Q), where P(x) and Q(x) are the characteristic X distribution of the collected online handwriting samples and the characteristic X distribution of the generated handwriting, according to the formula:

[0083] L=L noise +L content +L style +D KL (P||Q)

[0084] Determine the loss of the neural network.

[0085] The specific method of training the diffusion model may include using the pre-processed online handwriting data x0, obtaining a series of handwriting sequences with added noise through a forward diffusion process, inputting the sequences into the constructed neural network, and simultaneously inputting content information and style information, such as the input handwriting sequence data x0 at the tth time step. t , input neural network, output estimated noise ε θ (x t ,t). The diffusion model is trained using a noise estimation loss, which measures the difference between the estimated noise and the actual noise. Furthermore, the content information processing module and style information processing module in the neural network use content feature loss and style feature loss during training. These losses can constrain and optimize the extracted content and style features based on either classification or comparison-based loss functions. Furthermore, the generated sequence closely resembles the online sequence in terms of feature distribution.

[0086] The final loss function of the neural network can be the weighted sum of noise estimation loss, content feature loss, style feature loss, and feature distribution consistency loss. This loss is used as the optimization target of the neural network, and the mean square error loss L is used to calculate the loss. noise Calculate the difference between the estimated noise and the true noise and use (L content +L style ) calculates the difference between the content and style information and the true labels, and ultimately uses stochastic gradient descent to update the parameters of the noise estimation network, information processing module, and style processing module. After numerous iterations of optimizing the noise estimation network, the model accurately predicts the noise distribution in handwriting sequences and enhances the capabilities of the information processing and style processing modules. The resulting information processing module, style processing module, and backbone neural network are all sequence encoding networks.

[0087] After a large number of iterative optimizations of the noise estimation network, the model can accurately predict the noise distribution in handwriting sequences and enhance the capabilities of the information processing module and style processing module.

[0088] After the diffusion model training is completed, it can be used to generate online handwriting data.

[0089] Implementation 2: The diffusion model is trained. The noise estimation is optimized using the mean square error loss function (MSE) mentioned above. However, it is not limited to the mean square error loss. The Manhattan distance can be used to measure the difference between the estimated noise and the actual noise.

[0090] For example, according to the formula:

[0091] L noise =|ε θ(x t ,t)-ε| 2

[0092] Compute the difference between the estimated noise and the true noise.

[0093] At the same time, other losses can be used to jointly train the diffusion model, such as based on the sequence data x at any time step t t , noise estimation ε θ (x t ,t), hyperparameters, calling formula:

[0094]

[0095] Obtain the target handwriting data reconstructed by estimating the noise through the neural network Furthermore, we can use the SoftDWT (time series clustering) loss function or The coordinates (x, y) in the sequence are constrained by calculating the RMSE (root mean square error) loss to optimize the neural network parameters.

[0096] The distribution of the inverse denoising process conforms to the Gaussian distribution. According to the sampled random Gaussian noise, the standard deviation of the inverse denoising process distribution, and the predicted noise estimate, the noise is gradually removed to obtain the denoised handwriting data.

[0097] Implementation 3: For example, the inverse denoising process is used as a Markov chain. The denoised data of the current time step is determined based on the denoised data of the previous time step, and the mean of the inverse denoising process distribution is determined based on the noise estimate predicted by the neural network, and the noise is gradually removed to obtain the denoised handwriting data. For example, set the hyperparameter α t , construct (0,T) times hyperparameter α in a linear manner according to the number of denoising operations t , such as T = 1000, set α0 = 0.9999, α T =0.9998,

[0098] Based on the denoised data at time step t. According to the formula:

[0099]

[0100] Calculate the sequence data after removing the noise in step t-1, thereby obtaining the denoised data generated at each time step, until the denoising operation of the predetermined time step is completed, and obtain the handwriting data therein.

[0101] Among them, ε θ (x t ,t) is the noise estimate predicted by the neural network, σ t is the standard deviation of the inverse denoising process distribution, and the hyperparameter α in the diffusion model t The known parameters obtained According to the formula: Calculate the mean of the distribution of the inverse denoising process. After denoising T time steps, the result at t = 0 is the final generated online handwriting.

[0102] The generation process includes: after the diffusion model training is completed, the noise and data sequence distribution of the inverse denoising process is obtained, and the distribution conforms to the Gaussian distribution. The neural network performs a series of denoising processes on the random noise. By sampling the obtained inverse denoising process distribution, a denoised data sequence is obtained at each time step. At the same time, the denoised data sequence obtained at each time step is fused with content features and style features. Through a series of iterations, all noise is removed and the content features and style features are fused to generate handwriting sequence data. For example, the generation result from the iteration starting from time step T to t = 0 is the final generated online handwriting.

[0103] In the above embodiment, the handwriting data generation process still uses the inverse denoising process as a Markov chain. However, the inverse denoising process no longer follows the Markov chain.

[0104] In implementation mode 4, if the inverse denoising process does not follow the Markov chain, the distribution of the inverse denoising process conforms to the Gaussian distribution. Based on the standard deviation of the inverse denoising process distribution and the denoised sequence data at time step t, the formula is called:

[0105]

[0106] Calculate the denoised sequence data of any time step τ before [0, t). Where τ represents any time step before [0, t), σ t is the standard deviation of the inverse denoising process distribution.

[0107] Through this method, the inverse denoising process can skip time steps to remove noise and perform sequence data sampling. For example, sampling is performed at intervals of 20 time steps, which can greatly reduce the number of iteration steps.

[0108] The generation method based on Markov chain requires iterative generation step by step. For example, if the number of iterations is set to T = 1000, 1000 iterations are required to obtain the generated handwriting data. Assuming that the inverse denoising process no longer follows the Markov chain, the iterative calculation can jump time steps, such as iterating and sampling the denoising sequence every 20 time steps, which greatly reduces the number of iteration steps.

[0109] The present invention introduces the diffusion model into the field of online handwriting generation, integrates content information and style information into the diffusion model, and proposes a generation method that can generate rich and diverse online handwriting. Compared with traditional autoregressive models and GAN models, the handwriting generated by the diffusion model is more detailed and natural, and the training process avoids the problem of pattern collapse. During the training process based on the diffusion model, multiple loss functions are used for constraints and optimization. First, the difference between the estimated noise and the actual noise is calculated, and the difference between the content information and style information and the actual label is calculated. Finally, the parameters of the noise estimation network, information processing module, and style processing module are updated using stochastic gradient descent to optimize the online handwriting generation model.

[0110] The present invention can generate high-quality handwriting and can be used for model training and testing applications that require a large number of handwriting samples of different styles.

[0111] like Figure 3 As shown, electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. Various programs and data required for the operation of device 300 can also be stored in RAM 303. Computing unit 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0112] Multiple components within electronic device 300 are connected to I / O interface 305, including an input unit 306, an output unit 307, a storage unit 308, and a communication unit 309. Input unit 306 can be any type of device capable of inputting information into electronic device 300. Input unit 306 can receive input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 307 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 308 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 309 allows electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0113] The computing unit 301 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above. For example, in some embodiments, the reconstruction and decomposition of the muscle movement trajectory based on the original trajectory of the signature stroke, as well as the decomposition of its logarithmic velocity curve, can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via the ROM 302 and / or the communication unit 309. In some embodiments, the computing unit 301 can be configured to execute the signature handwriting dynamic acquisition implementation method by any other appropriate means (e.g., by means of firmware).

[0114] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0118] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0119] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0120] The applicant of the present invention has made a detailed illustration and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the detailed description of the above-mentioned implementation is not intended to limit the scope of the invention claimed for protection, but only represents the selected implementation methods of the present invention. Therefore, changes in the example numerical values ​​or replacement of functional modules should still fall within the scope of the present invention.

Claims

1. Online handwriting generation system based on diffusion model, characterized by: include: Diffusion model backbone neural network, content information processing module, style information processing module, the diffusion model backbone neural network performs a series of diffusion processing on the online handwriting sequence data to obtain pure noise with an approximate Gaussian distribution, and estimates the noise in the diffusion process; the neural network is trained using the online handwriting sequence data to optimize the parameters of the backbone neural network denoiser, content information processing module and style information processing module; the denoiser removes the noise component in the random noise to obtain basic handwriting data; the content information processing module and the style information processing module respectively obtain the content features and style features of the online handwriting sequence data; the trained neural network fuses the basic handwriting data with the content features and style features to generate handwriting.

2. The system according to claim 1, wherein: The backbone neural network model predicts noise based on noise loss. The diffusion model adds noise to the online handwriting sequence data according to the predicted noise to obtain a series of noisy data. The denoiser gradually removes random noise to obtain a series of denoised data. The diffusion model is optimized based on the noise estimation style loss and content loss, and the parameters of the denoiser, content information processing module and style information processing module in the diffusion model are updated.

3. The system according to claim 1, wherein: The diffusion process of online handwriting sequence data is constructed using Markov chain. A series of noisy data is generated by adding estimated noise to each time step of the online handwriting sequence data. The noisy data at each time step is used as the intermediate state of the Markov chain. The reverse denoising process generates basic handwriting data by gradually removing the estimated noise based on random noise at each time step.

4. The system according to claim 1, wherein: The backbone neural network estimates the noise in the point dimension of the online handwriting sequence and fuses the handwriting sequence data with the content features and style features in the point feature dimension. The fusion in the point feature dimension includes: concat([x,y,s], content feature vector, style feature vector, or add(mlp([x,y,s],mlp(content feature vector),mlp(style feature vector)).

5. The system according to any one of claims 1 to 4, characterized in that: The backbone neural network predicts the noise estimate for the next time step based on the noisy data, corresponding content information and style information at any time step in the forward diffusion process; the diffusion model is trained using the noise estimation loss, and the difference between the estimated noise and the true noise is measured by the loss; the extracted content features and style features are constrained and optimized using the content feature loss and style feature loss, and the point features are constrained and optimized using the feature distribution consistency loss; and the parameters of the noise estimation module, information processing module and style processing module are updated using stochastic gradient descent.

6. The system according to claim 5, characterized in that Diffusion model training uses the mean square error loss function to optimize noise estimation, and uses content feature loss and style feature loss to train the content information processing module and style information processing module. The extracted content features and style features are constrained and optimized, and the final loss function for neural network optimization is established. Specifically, according to the formula: L=L noise +L content +L style +D KL (P||Q) Determine the loss L of the neural network, where the noise estimate ε at any time step t θ (x t ,t) dimension is the same as the dimension of the collected online handwriting data, and Labels representing content information and style information respectively, Represents the output of the content information processing module, represents the output of the style information processing module, P is the feature distribution of online handwriting, and Q is the feature distribution of generated handwriting.

7. The online handwriting generation method based on the diffusion model is characterized by: include: Diffusion process, reverse denoising process, and training process. Diffusion process: gradually adding noise to the online handwriting sequence data until all sequence element values ​​of the handwriting sequence data obey the state of random Gaussian distribution; reverse denoising process: gradually removing noise in the random sequence through continuous iteration to generate basic handwriting data; Training process: Use online handwriting data and noise estimation to train and optimize network parameters of the diffusion model, content information processing module, and style information processing module; fusion process: The backbone neural network fuses basic handwriting data, content features, and style features to generate online handwriting sequence data.

8. The method according to claim 7, characterized in that During the diffusion process, the backbone neural network diffuses the online handwriting sequence data to obtain noise with an approximate Gaussian distribution, estimates the noise added at each time step during the diffusion process, optimizes the diffusion model based on the noise estimation style loss and content loss, and updates the parameters of the denoiser in the diffusion model. The denoiser gradually removes the noise in any random sequence based on the estimated noise to obtain basic handwriting data.

9. The method according to claim 7, characterized in that The diffusion process of online handwriting sequence data is constructed using Markov chain. A series of noisy data is generated by adding estimated noise to each time step of the online handwriting sequence data. The noisy data at each time step is used as the intermediate state of the Markov chain. The reverse denoising process gradually removes the estimated noise at each time step for the random sequence to generate the basic handwriting data.

10. The method according to claim 8 or 9, characterized in that Add estimated noise to the online handwriting sequence data, and calculate the noise data of the next time step based on the noise data of the previous time step, specifically including: t-1 , calling the formula: Calculate the noise data x at time step t t , where ε is the estimated noise, α t is a hyperparameter, and the hyperparameter α in (0, T) times of noise addition is constructed linearly according to the number of noise addition operations T. t .

11. The method according to claim 8 or 9, characterized in that For example, the inverse denoising process is constructed using a Markov chain. The mean of the inverse denoising process distribution is determined based on the estimated noise predicted by the neural network. The denoised data of the current time step is determined based on the denoised data of the previous time step. Specifically, based on the denoised data of the t-th time step, according to the formula: Calculate the denoised data after removing the noise in step t-1 to obtain the denoised data for each time step, where ε θ (x t ,t) is the predicted noise estimate, σ t is the standard deviation of the inverse denoising process distribution, is the mean of the distribution of the inverse denoising process.

12. The method according to claim 8 or 9, characterized in that If the inverse denoising process does not follow the Markov chain, based on the standard deviation of the inverse denoising process distribution, according to the denoised sequence data at time step t, the formula is called: Calculate the denoised data of any time step τ before [0, t), σ t is the standard deviation of the inverse denoising process distribution.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: in, The computer instructions are used to cause the computer to execute the method according to any one of claims 7 to 12.

Citation Information

Patent Citations

  • Font generation method and device based on diffusion model

    CN117057310A

  • Single sample font generation method based on multi-scale style fusion and font control

    CN118365744A

  • A single-sample handwritten text copying method based on diffusion model

    CN118762103B