Multi-style offline handwritten signature generation method based on diffusion model
Through the combination of the diffusion model and the style control module, high-fidelity generation of offline handwritten signatures for multiple styles is achieved, solving the problem of single style, low generation quality and insufficient details in the existing technology, and generating high-quality and diverse handwritten signatures.
Patent Information
- Application Number
- CN202510308854.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
When generating multi-style signatures, existing offline handwritten signature generation methods are difficult to take into account the coherence of handwriting, detail texture and dynamic characteristics of writing tools, resulting in the generated signatures being unable to meet the actual application needs in terms of style differences and realism.
The multi-style offline handwritten signature generation method based on the diffusion model is adopted. The original signature is forward noise-added through the forward processing module, the style control module generates a style modulation feature map, and the reverse processing module performs reverse noise denoising. Combining the style control module and the U-Net encoder decoder, precise control of the style is achieved.
High-quality and diverse handwritten signatures are generated, which improves the diversity and stylistic expression of the generated signatures, while ensuring the richness and fidelity of signature details, and overcoming the shortcomings in the prior art.
Smart Images

Figure CN120279567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a multi-style offline handwritten signature generation method based on a diffusion model. Background Art
[0002] Traditional offline handwritten signature generation methods mostly adopt generation methods based on template matching, rule deformation, or generative adversarial networks (GANs). There are certain limitations in the diversity, style control, and detail restoration of the generated signatures. Especially when generating multi-style signatures, it is often difficult to simultaneously take into account the handwriting coherence, detail texture, and dynamic characteristics of the writing tool, resulting in the generated signatures not meeting the actual application requirements in terms of style difference and authenticity.
[0003] In view of this, overcoming the defects of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a multi-style offline handwritten signature generation method based on a diffusion model to solve the problem that the existing technology cannot generate multi-style handwritten signatures.
[0005] The present invention adopts the following technical solutions: In a first aspect, the present invention provides a multi-style offline handwritten signature generation method based on a diffusion model, which uses a preset model to generate multi-style handwritten signatures. The preset model includes a forward processing module, a reverse processing module, and a style control module. The method includes: The forward processing module performs forward noise addition on the original signature to obtain a noise-added feature map; The style control module generates a style modulation feature map according to the style label; The reverse processing module uses the style modulation output feature to perform reverse denoising on the noise-added feature map to generate a handwritten signature of the corresponding style.
[0006] Preferably, the style control module includes a label processing unit, an attention gating unit, and a residual connection unit; The feature processing unit is used to perform a convolution operation on the style label to extract style features, and minimize the correlation of the style features to obtain a style embedding vector ; The attention gating unit is used to perform channel splicing on the style embedding vector and the intermediate feature map to perform a convolution operation on the spliced result, and then perform non-linear activation on the convolved result using the sigmoid function to generate a spatial attention weight map , where represents the sigmoid function, represents addition; The residual connection unit is used to output a style modulation feature map according to the spatial attention weight map and the intermediate feature map ; wherein, the intermediate feature map is generated by processing the noise-added feature map by the U-Net encoder in the reverse processing module, and
[0007] is a convolution operation.
[0007] Preferably, the minimization of the correlation of the style features is achieved by minimizing the correlation between different style features using an orthogonal constraint loss function, and the orthogonal constraint loss function is ; wherein, is a style embedding matrix, represents the Frobenius norm.
[0008] Preferably, the reverse processing module includes a U-Net encoder and a U-Net decoder with skip connections; the reverse processing module uses the style modulation output feature to perform reverse denoising on the noise-added feature map to generate a handwritten signature of the corresponding style, specifically including: The U-Net encoder extracts the intermediate feature map of the i-th level from the output feature map of the (i - 1)-th level, and further extracts the low-dimensional feature of the i-th level from the intermediate feature map of the i-th level; The U-Net decoder is used to upsample the low-dimensional feature of the i-th level to obtain the upsampling result of the i-th level, and superimpose the style modulation feature map with the upsampling result of the i-th level to obtain the output feature map of the i-th level; wherein, the noise-added feature map is used as the output feature map of the 0-th level; The output feature map of the last level is used as the handwritten signature of the corresponding style.
[0009] Preferably, the forward processing module performs forward noise addition on the original signature to obtain a noise-added feature map, specifically including: Performing time-step adaptive noise addition on the original signature, injecting hierarchical controllable noise based on a dynamic noise coefficient, and recording the noise distribution parameters at each time step until the signature features spread into a Gaussian distribution to obtain the noise-added feature map; wherein, at each time step , the noise coefficient is jointly determined by the style label and the current time step, specifically including: The noise coefficient ; wherein, and are the minimum and maximum noise intensities respectively, and are learning parameters, is the style label, is the total number of time steps; The noise type adopts a uniform noise distribution, and the noise at each time step is ; where, is the noise, is the noise amplitude, and U represents the uniform distribution.
[0010] Preferably, the preset model is pre-trained, and the loss function used in the training process is ; where, is the noise prediction network, t is the time step, is the pre-trained style discriminator, MMD is the maximum mean discrepancy metric, and are both preset weights, is the signature noise image corresponding to the time step t, is the style-generated signature image, is the expectation of the random variable.
[0011] Preferably, the training process of the preset model specifically includes: Construct an original dataset containing handwritten signature samples and multi-style labels, preprocess each handwritten signature sample in the original dataset to obtain preprocessed signature samples, and form a training dataset from the preprocessed signature samples and multi-style labels; Use the training dataset to train the initial model; where, the training includes the first-stage training and the second-stage training; In the first-stage training, keep the parameters of the style control module fixed and adjust the parameters of the forward processing module and the backward processing module; In the second-stage training, freeze the parameters of the forward processing module and the backward processing module and fine-tune the style control module.
[0012] Preferably, the preprocessing of each handwritten signature sample in the original dataset specifically includes: Perform a fast Fourier transform on the handwritten signature sample to filter out the high-frequency background noise component to obtain the filtered sample; Use an adaptive threshold to segment and extract the stroke connected domain from the filtered sample; Use a generative adversarial network to reconstruct the incomplete stroke edges of the stroke connected domain to restore the original writing dynamic features to obtain an intermediate sample; Perform data augmentation on the intermediate sample to obtain the preprocessed signature sample.
[0013] Preferably, the handwritten signature samples include real handwritten signature samples and forged signature samples generated by disturbing the styles and writing methods of real handwritten signature samples; the disturbances include one or more of disturbances in writing speed, disturbances in stroke pressure, and disturbances in rotation angle; The disturbance in writing speed is expressed as ; where , is the original speed, is the disturbance range; The disturbance in stroke pressure is expressed as ; where , is the original pressure, is the disturbance value, is the disturbance range; The disturbance in rotation angle is expressed as ; where represents the rotation matrix, is the rotation angle .
[0014] Preferably, the method further includes: Deploying a style discriminator, which is used to calculate a style similarity score according to the handwritten signature and the corresponding real signature , where is the sample format, is the i-th real signature, is the i-th generated forged signature; When the style similarity score is lower than a preset threshold, retrain the preset model.
[0015] In a second aspect, the present invention further provides a multi-style offline handwritten signature generation device based on a diffusion model, which is used to implement the multi-style offline handwritten signature generation method described in the first aspect. The device includes: At least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the multi-style offline handwritten signature generation method described in the first aspect.
[0016] In a third aspect, the present invention further provides a non-volatile computer storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the method described in the first aspect.
[0017] Fourthly, a chip is provided, including: a processor and an interface, which are used to call and run a computer program stored in a memory, and execute the method as described in the first aspect.
[0018] Fifthly, a computer program product containing instructions is provided. When the instructions run on a computer or a processor, the computer or the processor is enabled to execute the method as described in the first aspect.
[0019] By introducing a style control module into the diffusion model, the present invention realizes precise control of style, making it possible to generate multi-style handwritten signatures. Moreover, through the forward noise addition and reverse denoising processes of the diffusion model, the present invention effectively combines the style control mechanism to generate high-quality and diverse handwritten signatures. This embodiment not only improves the diversity of the generated signatures and the expressiveness of the style, but also ensures the richness and fidelity of the signature details, overcoming the defects in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is a schematic flowchart of a multi-style offline handwritten signature generation method based on a diffusion model provided by an embodiment of the present invention; Figure 2 is a schematic flowchart of a multi-style offline handwritten signature generation method based on a diffusion model provided by an embodiment of the present invention; Figure 3 is a schematic flowchart of a multi-style offline handwritten signature generation method based on a diffusion model provided by an embodiment of the present invention; Figure 4 is a schematic diagram of a preset model in a multi-style offline handwritten signature generation method based on a diffusion model provided by an embodiment of the present invention; Figure 5 is a schematic flowchart of a multi-style offline handwritten signature generation method based on a diffusion model provided by an embodiment of the present invention; Figure 6 is a schematic architecture diagram of a multi-style offline handwritten signature generation device based on a diffusion model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0023] Unless the context requires otherwise, throughout the specification and claims, the term "comprising" is to be construed in an open inclusive sense, i.e., "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples" or "some examples", etc., are intended to indicate that specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms are not necessarily referring to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any suitable manner, that is, although they may be carried in the embodiments or examples of the above terms due to reasons such as the order and position of appearance, they are not limited to being carried in a combined manner by one embodiment or example.
[0024] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality" is two or more. In addition, for example, in the description, for the same type of nouns, the method of adding "A" and "B" at the end is used to describe them as two independent individuals. In this case, the features defined with "A" and "B" are only used for the purpose of distinguishing similar individuals and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.
[0025] In the description of the present invention, there will be a description of the form "A and / or B" (where A and B are used to formally represent specific feature contents), and the corresponding description includes the following three combinations: only A, only B, and the combination of A and B.
[0026] As used in the present invention, "about", "substantially" or "approximate" includes the stated value and the average value within an acceptable deviation range of the specific value, where the acceptable deviation range is determined by those of ordinary skill in the art considering the measurements being discussed and the errors associated with the measurement of the specific quantity (i.e., the limitations of the measurement system).
[0027] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] Example 1: In recent years, as an emerging technology of generative models, diffusion models have achieved high-quality performance in generating image and sequence data by simulating the process of data gradually evolving from noise to the target distribution, providing a new technical idea for the multi-style generation of offline handwritten signatures. However, there is currently a lack of a complete technical solution for realizing the multi-style generation of offline handwritten signatures based on diffusion models. How to achieve precise control of styles and efficient parameter optimization while ensuring high-fidelity and diversity of the generated signatures remains a key problem to be solved urgently. To solve the problems of single style, low generation quality and insufficient details in the existing handwritten signature generation methods, Example 1 of the present invention provides a multi-style offline handwritten signature generation method based on a diffusion model, which uses a preset model to generate multi-style handwritten signatures. The preset model includes a forward processing module, a reverse processing module and a style control module. As Figure 1 shown, the method includes: In step 201, the forward processing module performs forward noise addition on the original signature to obtain a noise-added feature map.
[0029] In step 202, the style control module generates a style modulation feature map according to the style label.
[0030] In step 203, the reverse processing module uses the style modulation output feature to perform reverse denoising on the noise-added feature map to generate a handwritten signature of the corresponding style.
[0031] Among them, the forward processing module and the reverse processing module constitute the diffusion model. In this embodiment, by introducing a style control module into the diffusion model, precise control of styles is achieved, making it possible to generate multi-style handwritten signatures. Moreover, through the forward noise addition and reverse denoising processes of the diffusion model in this embodiment, the style control mechanism is effectively combined to generate high-quality and diverse handwritten signatures. This embodiment not only improves the diversity of the generated signatures and the expressiveness of styles, but also ensures the richness and fidelity of signature details, overcoming the defects in the prior art.
[0032] In a specific application scenario, as Figure 2 shown, the style control module includes a label processing unit, an attention gating unit and a residual connection unit; aiming to adjust the personalized style features of the signature during the generation process.
[0033] In step 301, the feature processing unit is used to perform a convolution operation on the style label to extract the style feature, and minimize the correlation of the style feature to obtain a style embedding vector ; Thus, it is ensured that the style embedding vectors of different styles are decoupled from each other in the latent space, avoiding interference between styles. This process can be understood as mapping style labels to a low-dimensional latent space, where each style label corresponds to an embedding vector. Assume there is a style embedding matrix , where is the dimension of the latent space, is the number of types of style labels, and each column corresponds to an embedding vector of a style label. By optimizing this matrix, the embedding vectors of different styles are made as independent as possible in the latent space.
[0034] Among them, the above is achieved by minimizing the correlation between different style features using an orthogonal constraint loss function, and the orthogonal constraint loss function is ; is the style embedding matrix, represents the Frobenius norm. This constraint ensures that the embedding vectors of style labels are orthogonal to each other in the latent space, thus effectively avoiding confusion and mutual interference between styles. During the optimization process, minimizing this loss makes the latent space embedding of style labels more independent, which helps to generate signature samples with clear style features. Through the style-decoupled latent space, the model can maintain efficient style control during the multi-style generation process, while avoiding feature interference between different styles, thereby improving the controllability and stability of signature generation. The independence of the embedding vector of each style in the latent space ensures the accuracy and consistency of the model when dealing with style conversion, making the generated signatures show clear feature differences under different styles.
[0035] In step 302, the attention gating unit is used to concatenate the style embedding vector with the intermediate feature map , perform a convolution operation on the concatenated result, and then use the sigmoid function to perform non-linear activation on the result of the convolution to generate a spatial attention weight map Among them, represents the sigmoid function, represents addition. That is, after adding the style vector and the intermediate feature map, convolution extraction is performed, and finally non-linear activation is performed using the sigmoid function.
[0036] In step 302, the residual connection unit is used to output a style-modulated feature map according to the spatial attention weight map and the intermediate feature map ; among them, the intermediate feature map is generated by processing the noise-added feature map by the U-Net encoder in the reverse processing module. is a convolution operation.
[0037] This embodiment effectively guides the style information to each layer of the network, thereby achieving precise control over the signature pen tip strength, connected stroke form, and writing tool characteristics, ensuring that the generated signature has rich style variations.
[0038] In an optional implementation manner, the reverse processing module includes a U-Net encoder and a U-Net decoder with skip connections; the reverse processing module uses the style modulation output features to perform reverse denoising on the noise-added feature map to generate a handwritten signature with the corresponding style, as Figure 3 shown, specifically including: In step 401, the U-Net encoder extracts the intermediate feature map of the i-th level from the output feature map of the (i - 1)-th level, and further extracts the low-dimensional feature of the i-th level from the intermediate feature map of the i-th level.
[0039] In step 402, the U-Net decoder is used to upsample the low-dimensional feature of the i-th level to obtain the upsampling result of the i-th level, and superimpose the style modulation feature map on the upsampling result of the i-th level to obtain the output feature map of the i-th level; wherein, the noise-added feature map is used as the output feature map of the 0-th level; the superimposing the style modulation feature map on the upsampling result of the i-th level can be: , where is the output feature map of the i-th level, is the upsampling result of the i-th level, is the style modulation feature map.
[0040] In step 403, the output feature map of the last level is used as the handwritten signature with the corresponding style.
[0041] It should be noted here that the processing process of the forward processing module and the processing process of the reverse processing module are both processes that loop. As Figure 4 shown, for the forward processing module, it performs the i-th noise addition on the output result of the (i - 1)-th time of the forward processing module to obtain the output result of the i-th time of the forward processing module, that is, Figure 4 the forward process shown, so as to finally obtain the noise-added feature map through continuous noise addition; for the reverse processing module, it performs denoising on the output result of the (i - 1)-th time of the reverse processing module (that is, the output feature map of the above (i - 1)-th level) to obtain the output result of the i-th time of the reverse processing module (that is, the output feature map of the above i-th level), and moreover, the first denoising of the reverse processing module is performed on the noise-added feature map, that is, the above-mentioned use of the noise-added feature map as the output feature map of the 0-th level. Among them, each denoising process of the reverse processing module is the encoding and decoding process in the above steps 401 - step 402.
[0042] The U-Net encoder is mainly composed of convolutional layers, activation functions, and pooling layers, and is responsible for feature extraction and downsampling operations on the input image. The intermediate feature map is obtained through convolutional operations. It is the output after the convolutional layer processes and contains the feature information of the image, with a relatively high spatial resolution. The low-dimensional features are the results after convolutional and pooling operations. The pooling layer further compresses the spatial resolution of the feature map while increasing the number of channels of the feature map. Simply put, the intermediate feature map is the feature map obtained after convolutional operations, while the low-dimensional features are the results obtained under the combined action of convolutional and pooling layers, with a lower spatial resolution and a higher abstract feature representation. Therefore, the relationship between the two can be summarized as: the intermediate feature map is the output of convolutional operations, while the low-dimensional features are the compressed features obtained through convolutional and pooling operations. The above step 302 can actually be understood as: the intermediate feature map of the i-th level is concatenated with the style embedding vector in the channel dimension to generate the spatial attention weight map of the i-th level. The spatial attention weight map of the i-th level is processed through step 303 to obtain the style modulation feature map of the i-th level. In step 402, the style modulation feature map of the i-th level is superimposed with the upsampling result of the i-th level to obtain the output feature map of the i-th level.
[0043] Among them, the forward processing module adds noise to the original signature in the forward direction to obtain a noisy feature map, specifically including: performing time-step adaptive noise addition on the original signature, injecting hierarchical controllable noise based on a dynamic noise coefficient, and recording the noise distribution parameters at each time step until the signature features spread into a Gaussian distribution to obtain the noisy feature map; among them, at each time step , the noise coefficient is jointly determined by the style label and the current time step, specifically including: the noise coefficient ; among them, and are the minimum and maximum noise intensities respectively, and are learning parameters, is the style label, is the total number of time steps; the noise type adopts a uniform noise distribution, and the noise at each time step is ; among them, is the noise, is the noise amplitude. Through this way of adding noise, the model can gradually increase the noise intensity while maintaining the style features, enabling signatures of different styles to show their respective characteristics during the noise addition process.
[0044] In actual use, the preset model is pre-trained, and the loss function used in the training process is ; among them, is the noise prediction network, t is the time step, The discriminator is for pre-training the style, and MMD is the maximum mean discrepancy metric. and are both preset weights, which are obtained by those skilled in the art through empirical analysis. is the signature noise image corresponding to time step t. is the style generation signature image. is the expectation of the random variable. By adjusting these two weights, the model can balance the relationship between noise removal and style consistency during the training process, further improving the quality and accuracy of the generated signature. In another alternative implementation, multiple loss functions can also be used for training. For example, is used to constrain the noise prediction error to measure the difference between the generated signature and the target signature, optimizing the noise removal effect. And is used to constrain the style consistency loss to maintain the difference between different styles and ensure that the generated signature conforms to the target style. Among them, represents the style conversion function. and respectively represent the style features of the generated signature and the target signature. By jointly optimizing these two loss functions, the model can maintain both the accuracy of details and the consistency of style when generating signatures.
[0045] The training process of the preset model, as shown in Figure 5 , specifically includes: In step 501, an original dataset containing handwritten signature samples and multi-style labels is constructed, and each handwritten signature sample in the original dataset is preprocessed to obtain preprocessed signature samples. The preprocessed signature samples and multi-style labels form a training dataset.
[0046] In step 502, the initial model is trained using the training dataset. Among them, the training includes the first-stage training and the second-stage training. In the first-stage training, the parameters of the style control module are kept fixed, and the parameters of the forward processing module and the reverse processing module are adjusted. In the second-stage training, the parameters of the forward processing module and the reverse processing module are frozen, and the style control module is fine-tuned.
[0047] In actual use, when to end the first-stage training and enter the second-stage training is obtained by those skilled in the art through empirical analysis. For example, it can be: when the number of input preprocessed signature samples is greater than a preset number, or the first-stage training is until the loss function converges to a preset value.
[0048] During the training process, 1000 training epochs were set, and the batch size was 16. The learning rate lr was set to 1e-4, the first momentum beta_1 of the Adam optimizer was 1e-4, the hyperparameter beta_T in the diffusion model was set to 0.02, and the gradient clipping value was 1.0. In terms of the network structure, the time step T was 100, the initial number of channels channel was 128, the channel doubling was set to [1, 2, 3, 4], 2 layers of attention mechanism attn were used, the number of residual blocks was 2, the Dropout probability was set to 0.15, and the image input size was set to 128.
[0049] Among them, the preprocessing of each handwritten signature sample in the original dataset specifically includes: performing a fast Fourier transform on the handwritten signature sample to filter out the high-frequency background noise component to obtain the filtered sample; using an adaptive threshold to segment and extract the stroke connected regions from the filtered sample; using a generative adversarial network to reconstruct the incomplete stroke edges of the stroke connected regions to restore the original writing dynamic features to obtain an intermediate sample; and performing data augmentation on the intermediate sample to obtain the preprocessed signature sample.
[0050] The preprocessing can also be understood as the following operations: (1) Frequency domain filtering: Using Fourier transform to suppress high-frequency noise in the frequency domain, thereby removing the interference noise that may be introduced during writing.
[0051] (2) Spatial morphological operation: Using an adaptive threshold method to segment the image to further eliminate interference and obtain a clear signature contour.
[0052] Through the above preprocessing, the quality of the dataset and the training effect of the model can be significantly improved.
[0053] In addition, in order to enhance the generalization ability of the model, a data augmentation strategy is adopted, including applying geometric transformations such as rotation, translation, and scaling to the signature samples, thereby increasing the diversity of the training data.
[0054] In a specific application scenario, the handwritten signature samples include real handwritten signature samples and forged signature samples generated by perturbing the style and writing method of the real handwritten signature samples; the perturbations include one or more of the perturbations of writing speed, stroke pressure, and rotation angle; the perturbation of writing speed is expressed as ; where , is the original speed, is the perturbation range; the perturbation of stroke pressure is expressed as ; wherein, , is the original pressure, is the perturbation value, is the perturbation range; The perturbation of the rotation angle is manifested as ; wherein, represents the rotation matrix, is the rotation angle .
[0055] Through the generation of these forged signature samples, the model can not only identify real signatures but also effectively distinguish forged signatures, thereby enhancing the security and robustness of the model.
[0056] The method further includes: deploying a style discriminator, which is used to calculate a style similarity score according to the handwritten signature and the corresponding real signature , where is the sample format, is the i-th real signature, is the i-th generated forged signature; When the style similarity score is lower than a preset threshold, the preset model is retrained.
[0057] The preset threshold is obtained by those skilled in the art through empirical analysis. In an actual application scenario, the preset threshold can be 0.85, and the style discriminator can be a lightweight style discriminator composed of a 5-layer convolutional application network. What can be input into the style discriminator can be the histogram of gradient directions of the handwritten signature (i.e., the generated signature) and the real signature.
[0058] It should be noted here that the arrow leading to the orientation style control module in the forward process of the Figure 4 is only shown schematically and does not represent that the processing result of the forward processing module is input into the style control module.
[0059] In this embodiment, by introducing a diffusion model to model the generation process of handwritten signatures, and through the forward noise addition and reverse denoising processes, effectively combining the style control mechanism, it can realize the high-fidelity generation of multi-style offline handwritten signatures according to user input and pre-set style tags, so as to generate diverse and high-quality handwritten signatures, which are applicable to fields such as finance, law, and electronic contracts that require personalized signature verification and generation. This embodiment not only improves the diversity of the generated signatures but also ensures the style expressiveness and detail fidelity.
[0060] Generally speaking, the present invention provides an efficient, accurate and diverse offline handwritten signature generation scheme through a multi-style generation method based on a diffusion model. Compared with the prior art, it has the following beneficial effects: (1) The present invention can utilize a diffusion model and a style control mechanism to generate handwritten signatures that not only ensure quality but also exhibit rich style expressions and detailed fidelity.
[0061] (2) Through the guidance of style tags and fine-tuning of the noise addition process, handwritten signatures of various styles can be generated, including simulating characteristics such as different writing tools, writing speeds, and pressures.
[0062] Example 2: Based on the method described in Example 1, the present invention combines specific application scenarios and uses technical expressions in relevant scenarios to elaborate on the implementation process in the characteristic scenarios of the present invention.
[0063] A multi-style offline handwritten signature generation method based on a diffusion model provided in this embodiment specifically includes: As Figure 4 shown, the multi-style offline handwritten signature generation method based on a diffusion model proposed by the present invention specifically includes: Step 1: Data preprocessing: Construct an original dataset containing real handwritten signatures (i.e., real handwritten signature samples) and various style tags. Each signature sample is associated with a style tag, and the style tags include information such as writing tools, writing pressures, and writing speeds. The dataset also includes forged signature samples to enhance the robustness of the model.
[0064] Specifically, for each signature sample, noise suppression processing is performed. The noise suppression process includes the following two steps: (1) Frequency domain filtering: Use Fourier transform to suppress high-frequency noise in the frequency domain, thereby removing interference noise that may be introduced during the writing process.
[0065] (2) Spatial morphological operation: Use an adaptive threshold method to segment the image and further eliminate interference to obtain a clear signature contour.
[0066] Through the above preprocessing, the quality of the dataset and the training effect of the model can be significantly improved.
[0067] In addition, to enhance the generalization ability of the model, a data augmentation strategy is adopted, including applying geometric transformations such as rotation, translation, and scaling to the signature samples, thereby increasing the diversity of training data.
[0068] Step 2: Hierarchical controllable forward noise addition: In this stage, a diffusion model is used to perform a forward noise addition process on each signature sample. This process gradually blurs the original signature information in multiple stages, adding noise at each time step until the signature image gradually becomes unrecognizable.
[0069] The noise intensity in the noise addition process is closely related to the style tag and the current time step. Specifically, the noise coefficient is given by the following formula:
[0070] where and are the minimum and maximum noise intensities respectively, and are learning parameters, is the style label, is the total number of time steps, is a control function for the style label and the noise intensity.
[0071] Through this way of adding noise, the model can gradually increase the noise intensity while maintaining the style features, enabling signatures of different styles to exhibit their respective characteristics during the noise addition process.
[0072] Step 3: Style-guided reverse denoising generation: In the reverse denoising stage, an improved U-Net architecture is adopted to gradually remove the noise in the signature image. At each stage, the style label and the noise information are jointly input into the network so that the style features can be retained during the denoising process.
[0073] The specific process is as follows: The encoder part of the U-Net is responsible for extracting low-dimensional features from the noisy signature image.
[0074] The decoder then gradually reconstructs the denoised image, and the style control module adjusts the denoising strategy according to the specific features of the style label.
[0075] This process ensures that the generated signatures are consistent in style performance and avoids style distortion by introducing the attention mechanism and residual connections. The style adjustment at each stage is based on the features of the current signature, and the denoising effect is optimized through an adaptive mechanism.
[0076] Step 4: Dual-path joint parameter optimization: During the model training process, a dual-path joint optimization strategy is adopted, combining the forward noise addition error and the reverse denoising error for multi-objective optimization. This strategy is optimized through the following two loss functions: Noise prediction error: Measures the difference between the generated signature and the target signature, and optimizes the noise removal effect.
[0077] The loss function is defined as:
[0078] Style consistency loss: Used to maintain the differences between different styles and ensure that the generated signatures conform to the target style.
[0079] The style consistency loss is:
[0080] Among them, represents the style conversion function, and respectively represent the style features of the generated signature and the target signature. By jointly optimizing these two loss functions, the model can maintain both the accuracy of details and the consistency of style when generating signatures.
[0081] Step 5: Construction and optimization of the style-disentangled latent space: To further improve the stability and controllability of multi-style generation, a style-disentangled latent space is constructed. Specifically, an orthogonal constraint loss is used to minimize the correlation between style labels, ensuring the independence of embedding vectors of different styles in the latent space. The definition of the orthogonal constraint loss is:
[0082] Among them, is the self-inner product of the style embedding matrix, is the identity matrix, is the Frobenius norm. This constraint ensures the independence of style labels in the latent space, thus avoiding style confusion and interference, and making the generation results of each style more accurate.
[0083] To further improve the robustness of the model, forged signature samples are also included in the dataset. Forged signatures are generated by perturbing the style and writing method of real signatures. The specific operations include: (1)Perturbation of writing speed: By increasing or decreasing the writing speed, the change in the density of signature strokes is simulated. The speed perturbation formula is:
[0084] Among them , is the original speed, is the perturbation range.
[0085] (2)Perturbation of stroke pressure: By adjusting the writing pressure, the thickness of the strokes of the signature is made different. The pressure perturbation formula is: By introducing the change in writing pressure, different effects of the thickness of the signature strokes are presented. The pressure perturbation formula is:
[0086] Among them , is the original pressure, is the perturbation value, is the perturbation range.
[0087] (3)Rotation angle perturbation: By slightly rotating the signature, the angle changes during writing are simulated. The rotation operation is achieved through affine transformation, and the formula is:
[0088] where represents the rotation matrix, is the rotation angle .
[0089] Through the generation of these forged signature samples, the model can not only identify real signatures but also effectively distinguish forged signatures, thereby enhancing the security and robustness of the model.
[0090] During the training process, 1000 training epochs and a batch size of 16 were set. The learning rate lr was set to 1e-4, the first momentum beta_1 of the Adam optimizer was 1e-4, the hyperparameter beta_T in the diffusion model was set to 0.02, and the gradient clipping value was 1.0 grad_clip. In terms of the network structure, the time step T was 100, the initial number of channels channel was 128, the channel doubling was set to [1, 2, 3, 4], 2 layers of attention mechanism attn were used, the number of residual blocks was 2 num_res_blocks, and the Dropout probability was set to 0.15 dropout. The image input size was set to 128 img_size.
[0091] In another implementation manner, the multi-style offline handwritten signature generation method based on the diffusion model proposed by the present invention specifically includes: Step 1: Construction and preprocessing of the multi-style dataset. First, an original dataset containing real handwritten signatures and various style labels is constructed. Each signature sample is associated with a style label, and the style label can be marked according to the writing tool, writing pressure, speed, etc. of the signature. The dataset also includes forged signature samples to enhance the robustness of the model. For each sample, noise suppression processing is performed to remove the noise that may be introduced during writing. The noise suppression process includes frequency-domain filtering and spatial morphological operations. The Fourier transform is used to suppress high-frequency noise in the frequency domain, and the image is segmented by an adaptive threshold method to eliminate interference. The data augmentation strategy further expands the dataset, including applying geometric transformations such as rotation, translation, and scaling to the signature samples to enhance the generalization ability of the model.
[0092] Step 2: Layered Controllable Forward Noise Addition Apply the noise addition process to each signature sample. Gradually add noise to the signature image through the diffusion model. The noise addition process is multi-stage, gradually blurring the original signature information. During this process, the noise intensity is jointly controlled by the time step and the style label. Specifically, at each time step , the noise coefficient is jointly determined by the style label and the current time step to regulate the intensity of noise addition, so that signatures of different styles retain their respective style characteristics during the noise addition process. The noise coefficient has the following formula:
[0093] where and are the minimum and maximum noise intensities respectively, and are learning parameters, is the style label, is the total number of time steps. Through this noise addition method, the model can balance detail retention and style change during the transformation of each style.
[0094] Step 3: Style-Guided Reverse Denoising Generation During the reverse denoising process, use the improved U-Net architecture to gradually denoise the signature image with added noise. During this process, the style label and the noise information are jointly input into the network. The style control module learns to obtain specific features of each style and adjusts the features of the denoising process at each stage. The encoder in the U-Net structure gradually extracts the low-dimensional features of the image, and the decoder gradually reconstructs the denoised image. At the same time, the style control module adjusts features such as the stroke shape, pressure, and speed of different styles through the attention mechanism and residual connections, so that the generated signatures are consistent in style and avoid style distortion.
[0095] Step 4: Dual-Path Joint Parameter Optimization During the model training process, adopt the dual-path joint optimization strategy, combining the forward noise addition error and the reverse denoising error for multi-objective optimization. The loss function consists of two parts: one is the noise prediction error, and the other is the style consistency loss. The noise prediction error measures the difference between the generated signature and the target signature, and the style consistency loss is used to maintain the difference between different styles. Specifically, the style consistency loss is given by the following formula:
[0096] where represents the style conversion function, and represent the style features of the generated signature and the target signature respectively. Through joint optimization, ensure that the generated signature not only meets the style requirements but also can maintain details under noise conditions.
[0097] Step 5: Construction and Optimization of Style-Decoupled Latent Space To further improve the stability and controllability of multi-style generation, a style-decoupled latent space is constructed. Specifically, the orthogonality constraint loss is used to minimize the correlation between style labels, ensuring that the embedding vectors of different styles are decoupled from each other in the latent space. The form of the orthogonality constraint is as follows:
[0098] where represents the style embedding matrix, is the identity matrix, represents the Frobenius norm. This constraint ensures the independence of style labels in the latent space, avoiding style confusion and interference with each other.
[0099] In an embodiment of the present invention, the dataset in Step 1 not only contains real signature samples, but also enhances the diversity of the dataset through forged signatures. The forged signatures are generated by perturbing the style and writing manner of real signatures. Specifically, the forged signatures introduce perturbations in aspects such as writing speed, stroke pressure, and rotation angle on the basis of real signatures to simulate forged samples with different writing styles. By changing the writing speed of the signature, the density of the signature strokes changes. The speed change formula is:
[0100] where , is the original speed, is the perturbation value, is the perturbation range.
[0101] By introducing changes in writing pressure, the thickness of the signature strokes presents different effects. The pressure perturbation formula is:
[0102] where , is the original pressure, is the perturbation value, is the perturbation range.
[0103] The signature is slightly rotated to simulate the change in angle during writing. The rotation operation can be achieved through affine transformation, and the formula is:
[0104] where represents the rotation matrix, is the rotation angle .
[0105] After the forged signatures are generated through these perturbations, they are added to the training dataset, further enhancing the model's adaptability to multiple styles and enabling the model to maintain high accuracy and robustness when processing signatures of different styles.
[0106] In one embodiment of the present invention, the noise addition process in step two adds noise to each time step of the signature image, making the noise addition intensity associated with specific features of the style label, thereby achieving fine control of signatures of different styles. The noise addition process uses the forward noise addition stage of the diffusion model to achieve style control by gradually adding noise to the signature image at each time step. The specific process is as follows: Jointly control the intensity of noise addition through the style label and the current time step. Each style label corresponds to a specific noise intensity curve, and the intensity of the noise will vary according to the style label. The noise intensity
[0107] where, and respectively represent the minimum and maximum values of the noise intensity, and are learning parameters, is the style label.
[0108] The noise type adopts a uniform noise distribution to ensure the diversity of the generated samples. The noise at each time step ttt can be expressed as:
[0109] where, is the noise, is the noise amplitude, which is adjusted by the style label to ensure that signatures of different styles exhibit different detailed features during the noise addition process.
[0110] By introducing the interaction between the style label and the noise, the model can incorporate style features into the generation process during the noise addition process, ensuring that each signature has a unique style and rich details. The style features play a guiding role in the noise process, ensuring the consistency of the generated results and the accuracy of the style.
[0111] In one embodiment of the present invention, the U-Net architecture in step three introduces an improved style control module, enabling the reverse denoising process to adjust the denoising strategy according to the characteristics of the style label, effectively avoiding style distortion and ensuring that each generated signature is consistent with the target style. As a classic image generation model, the U-Net architecture performs feature extraction and reconstruction through a multi-layer encoding and decoding structure. In the present invention, the U-Net is improved to handle the influence of the style label on the denoising process. The improvement method includes: during the decoding process, a style control module is added between the encoder and decoder of the U-Net, and this module adjusts the weights of each layer of the network through the input of the style label, so as to fuse style features during the denoising process. The style control module affects the denoising process in the following way: style embedding vector .
[0112] Wherein, is the input feature, and are the learned style weight parameters, is the feature after style adjustment.
[0113] In each decoding stage, the style feature guides the reconstruction process, ensuring that the generated signature can not only remove noise but also maintain the stroke form, writing speed, and pressure characteristics of the original style.
[0114] The style control module avoids mutual interference between different styles through an adaptive learning mechanism, enabling the signature of each style to accurately reflect the target writing characteristics.
[0115] In one embodiment of the present invention, the dual-path joint optimization strategy in step four can optimize both noise removal and style consistency, thereby improving the quality of the generated signature. During the training process, the model is optimized through the dual-path joint optimization strategy, and the core of this strategy lies in optimizing two objectives simultaneously: noise removal and style consistency. Specifically, the optimization objective consists of the following two parts: By calculating the L2 loss between the generated signature and the original signature, the noise removal effect is optimized. The loss function is
[0116] By calculating the style difference between the generated signature and the target style signature, the style consistency is optimized. The style consistency loss uses a style matching loss function, which quantifies the style consistency by calculating the cosine similarity of the signature features:
[0117] Wherein, is the feature vector of the target style.
[0118] The combined optimization loss function consists of a noise removal loss and a style consistency loss, ensuring that the generated signature is close to the original signature in details while maintaining consistency with the target style in style. The final loss function is:
[0119] where and are weight parameters. By adjusting these two weights, the model can balance the relationship between noise removal and style consistency during training, further improving the quality and accuracy of the generated signature.
[0120] In an example of the present invention, in the construction and optimization of the style decoupled latent space in step five, a style decoupled latent space is constructed. Specifically, the style decoupled latent space minimizes the correlation between style labels by introducing an orthogonal constraint loss, thereby ensuring that the embedding vectors of different styles are decoupled from each other in the latent space and avoiding interference between styles. This process is achieved through the following steps: Style labels are mapped to a low-dimensional latent space, where each style label corresponds to an embedding vector. Suppose there is a style embedding matrix where is the dimension of the latent space, is the number of types of style labels, and each column corresponds to an embedding vector of a style label. By optimizing this matrix, the embedding vectors of different styles are made as independent as possible in the latent space.
[0121] To ensure the independence of different style labels in the latent space, we introduce an orthogonal constraint loss. This loss function promotes the embedding vectors to be orthogonal to each other by calculating the product of the transpose of the style embedding matrix and itself, that is:
[0122] where is the self-inner product of the style embedding matrix, is the identity matrix, is the Frobenius norm. This constraint ensures that the embedding vectors of style labels are orthogonal to each other in the latent space, thus effectively avoiding confusion and mutual interference between styles. During the optimization process, minimizing this loss makes the latent space embedding of style labels more independent, which helps to generate signature samples with clear style features.
[0123] By decoupling the latent space by style, the model can maintain efficient style control during the multi-style generation process, while avoiding feature interference between different styles, thereby improving the controllability and stability of the generated signature. The independence of the embedding vectors of each style in the latent space ensures the accuracy and consistency of the model when dealing with style conversion, enabling the generated signature to exhibit distinct feature differences under different styles.
[0124] Embodiment 3: As Figure 6 shown, it is a schematic diagram of the architecture of the multi-style offline handwritten signature generation device based on the diffusion model according to the embodiment of the present invention. The multi-style offline handwritten signature generation device based on the diffusion model in this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 6 One processor 21 is taken as an example here.
[0125] The processor 21 and the memory 22 can be connected through a bus or other means, Figure 6 Taking the connection through the bus as an example here.
[0126] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the multi-style offline handwritten signature generation method based on the diffusion model in Embodiment 1. The processor 21 executes the multi-style offline handwritten signature generation method based on the diffusion model by running the non-volatile software programs and instructions stored in the memory 22.
[0127] The memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely set relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0128] The program instructions / modules are stored in the memory 22 and, when executed by the one or more processors 21, execute the multi-style offline handwritten signature generation method based on the diffusion model in Embodiment 1 above.
[0129] It should be noted that the content such as information interaction and execution process between the modules and units within the above device and system, due to being based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.
[0130] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, optical discs, etc.
[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-style offline handwritten signature generation method based on a diffusion model, characterized in that, Generate multi-style handwritten signatures using a preset model, where the preset model includes a forward processing module, a reverse processing module, and a style control module; the method includes: The forward processing module performs forward noise addition on the original signature to obtain a noise-added feature map; The style control module generates a style modulation feature map according to the style label; The reverse processing module uses the style modulation output feature to perform reverse denoising on the noise-added feature map to generate a handwritten signature of the corresponding style.
2. The method for generating multi-style offline handwritten signatures based on a diffusion model according to claim 1, wherein The style control module includes a label processing unit, an attention gating unit, and a residual connection unit; The feature processing unit is used to perform a convolution operation on the style tags to extract style features, and minimize the correlation of the style features to obtain a style embedding vector ; The attention gating unit is used to combine the style embedding vector with the intermediate feature map for channel concatenation, perform a convolution operation on the concatenated result, and then non-linearly activate the result of the convolution using the sigmoid function to generate a spatial attention weight map , where represents the sigmoid function represents addition; The residual connection unit is used to output a style modulation feature map according to the spatial attention weight map and the intermediate feature map ; wherein, the intermediate feature map is generated by processing the noise-added feature map by a U-Net encoder in the reverse processing module, and is a convolution operation.
3. The multi-style offline handwritten signature generation method based on a diffusion model according to claim 2, wherein, The minimization of the correlation of the style features is achieved by minimizing the correlation between different style features using an orthogonal constraint loss function, and the orthogonal constraint loss function is ; Among them, is the style embedding matrix, represents the Frobenius norm.
4. The multi-style offline handwritten signature generation method based on the diffusion model according to claim 1, wherein, The reverse processing module includes a U-Net encoder and a U-Net decoder with skip connections; The reverse processing module uses the style modulation output feature to perform reverse denoising on the noise-added feature map to generate a handwritten signature of the corresponding style, specifically including: The U-Net encoder extracts the intermediate feature map of the i-th level from the output feature map of the (i - 1)-th level, and further extracts the low-dimensional feature of the i-th level from the intermediate feature map of the i-th level; The U-Net decoder is used to upsample the low-dimensional feature of the i-th level to obtain the upsampling result of the i-th level, and superimpose the style modulation feature map on the upsampling result of the i-th level to obtain the output feature map of the i-th level; among them, the noise-added feature map is used as the output feature map of the 0-th level; Use the output feature map of the last level as the handwritten signature of the corresponding style.
5. The method for generating multi-style offline handwritten signatures based on a diffusion model according to claim 1, wherein The forward processing module performs forward noise addition on the original signature to obtain a noise-added feature map, specifically including: Perform time-step adaptive noise addition on the original signature, inject hierarchical controllable noise based on the dynamic noise coefficient, record the noise distribution parameters of each time step until the signature features diffuse into a Gaussian distribution to obtain the noise-added feature map; Among them, at each time step , the noise factor is jointly determined by the style label and the current time step, specifically including: Noise coefficient ; wherein and are the minimum and maximum noise intensities respectively, and are learning parameters, is the style label, is the total number of time steps; The noise type adopts a uniform noise distribution, and for each time step the noise is ; where is the noise, is the noise amplitude, and U represents a uniform distribution.
6. The multi-style offline handwritten signature generation method based on a diffusion model according to claim 1, characterized in that The preset model is obtained by pre-training, and the loss function used in the training process is ; Among them, is the noise prediction network, t is the time step, is the pre-trained style discriminator, MMD is the maximum mean discrepancy metric, and are both preset weights, is the signature noise image corresponding to the time step t, is the style generation signature image, is the expectation of the random variable.
7. The method for generating multi-style offline handwritten signatures based on a diffusion model according to claim 6, characterized in that, The training process of the preset model specifically includes: Construct an original dataset containing handwritten signature samples and multi-style labels, preprocess each handwritten signature sample in the original dataset to obtain a preprocessed signature sample, and form a training dataset from the preprocessed signature sample and the multi-style labels; Use the training dataset to train the initial model; among them, the training includes the first-stage training and the second-stage training; In the first-stage training, keep the parameters of the style control module fixed and adjust the parameters of the forward processing module and the reverse processing module; In the second-stage training, freeze the parameters of the forward processing module and the reverse processing module and fine-tune the style control module.
8. The method for generating multi-style offline handwritten signatures based on a diffusion model according to claim 7, characterized in that, The preprocessing of each handwritten signature sample in the original dataset specifically includes: Perform a fast Fourier transform on the handwritten signature sample to filter out the high-frequency background noise component to obtain a filtered sample; Use an adaptive threshold to segment and extract the stroke connected domain from the filtered sample; Use a generative adversarial network to reconstruct the incomplete stroke edges of the stroke connected domain to restore the original writing dynamic features to obtain an intermediate sample; Perform data augmentation on the intermediate sample to obtain a preprocessed signature sample.
9. The method for generating multi-style offline handwritten signatures based on a diffusion model according to claim 7, wherein The handwritten signature samples include real handwritten signature samples and forged signature samples generated by perturbing the style and writing manner of the real handwritten signature samples; the perturbation includes one or more of perturbation of writing speed, perturbation of stroke pressure, and perturbation of rotation angle; The disturbance of the writing speed is manifested as ; where , is the original speed, is the disturbance range; The disturbance of the stroke pressure is manifested as ; among which , is the original pressure, is the disturbance value, is the disturbance range; The perturbation of the rotation angle is manifested as ; where represents the rotation matrix, is the rotation angle .
10. The method for generating multi-style offline handwritten signatures based on a diffusion model according to any one of claims 1-9, characterized in that, The method further includes: Deploy a style discriminator, which is used to calculate a style similarity score based on the handwritten signature and the corresponding genuine signature , where is the sample format, is the i-th genuine signature, is the i-th generated forged signature; When the style similarity score is lower than a preset threshold, retraining the preset model.