Diffusion model character repairing method and device based on dual-condition guidance

Through the dual-condition-guided diffusion model, the text skeleton topology and stroke timing feature matrix are used to solve the problem of insufficient accuracy and adaptability of Chu-based text repair in the existing technology, and high-quality text repair results are achieved.

CN120298259AActive Publication Date: 2025-07-11YANGTZE UNIVERSITY

Patent Information

Application Number
CN202510313957.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-11
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

In the field of text repair, especially when repairing Chu-based texts with complex backgrounds and severely degraded images, there are problems of insufficient repair accuracy, structural distortion and missing stroke details. It is difficult for existing methods to accurately control the structural details and stroke direction of text.

Method used

A diffusion model based on dual condition guidance is adopted to obtain text skeleton topology and stroke timing feature matrix, combining cross-attention fusion and dynamic weight scheduling to generate high-quality repair results.

Benefits of technology

High-precision text repair in complex scenarios is realized, ensuring the structural rationality and visual coherence of the generated results, improving the structural rationality and visual coherence of the repair effect, and reducing data dependence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298259A_ABST
    Figure CN120298259A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion model character restoration method and device based on dual-condition guidance, and belongs to the field of artificial intelligence of character restoration technology, and the method comprises the steps: obtaining a first feature matrix and a second feature matrix; taking the first feature matrix as a query vector, and carrying out cross-attention fusion on the query vector and a key vector and a value vector of the second feature matrix to generate a target condition vector; injecting the target condition vector into a noise preset network of a diffusion model of which the initial input is the to-be-repaired character image, and dynamically fusing the local feature of the jump connection with the target condition vector through a cross attention block to generate an enhanced feature; and generating a repaired character image by combining weight scheduling of skeleton priori and stroke priori. According to the method, the structure and stroke features are obtained through the two feature matrixes, and character constraint conditions are dynamically introduced in the reverse denoising process of the diffusion model, so that the model can retain details of strokes and character style features while generating a high-quality repair result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence for text restoration technology, and more specifically, to a method and device for text restoration based on a diffusion model guided by dual conditions. Background Art

[0002] As an important carrier of ancient Chu culture, Chu-style characters contain rich historical, cultural, and academic values. However, having been passed down for thousands of years, these characters are often damaged to varying degrees, such as stroke loss and blurred text, due to erosion by natural factors such as water and wind. In such cases, restoration usually involves highly complex, time-consuming, and professional workflows.

[0003] Existing technologies in the field of text restoration mainly include traditional text restoration methods, deep learning-based image generation methods, and diffusion models.

[0004] Traditional text restoration methods achieve restoration through techniques such as stroke feature extraction, edge detection, and image segmentation. They have limited effects when dealing with complex backgrounds, severely degraded images, and missing detailed strokes, are prone to structural distortions (such as broken or adhered strokes) or semantic errors (such as incorrect combinations of Chinese character radicals), and the restoration process lacks flexibility. Existing conditional control methods based on text prompts are difficult to precisely constrain the structural details of text and have insufficient robustness to complex glyphs or low-quality inputs.

[0005] Among deep learning-based image generation methods, the VAE-based restoration method can lack detail control; the GAN-based restoration method is vulnerable to the instability of adversarial training, relies on a large amount of labeled data, and the restored strokes are prone to blurring or artifacts and are difficult to constrain the global structure; the CNN-based restoration method relies on local textures and lacks explicit modeling of the global skeleton and stroke order of text.

[0006] Diffusion models perform outstandingly in the field of image generation. However, in existing technologies, the resolution of the prior mask limits the fine restoration ability, the restoration process lacks dynamic adjustment, and it is impossible to precisely control the stroke direction and skeleton topology, resulting in distorted generated text structures, and there is also a certain dependence on training data. Summary of the Invention

[0007] Embodiments of the present invention provide a method and device for text restoration based on a diffusion model guided by dual conditions, aiming to solve the problems of insufficient restoration accuracy and adaptability to complex damage scenarios in the existing text restoration process.

[0008] A method for text restoration based on a diffusion model guided by dual conditions includes:

[0009] Obtain the first feature matrix and the second feature matrix; the first feature matrix is a global skeleton feature matrix modeled based on the text skeleton topology of a preset complete character image; the second feature matrix is a stroke timing feature matrix modeled based on the timing features and morphological details of a preset complete character image.

[0010] Use the first feature matrix as a query vector, and perform cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector.

[0011] Inject the target conditional vector into the noise preset network of a diffusion model with the initial input being the character image to be repaired, and dynamically fuse the local features of the skip connection with the target conditional vector through a cross-attention block to generate enhanced features.

[0012] Perform a linear combination of the enhanced features with the first feature matrix and the second feature matrix, and combine the weight scheduling of the skeleton prior and the stroke prior to generate the repaired character image.

[0013] Further, the step of using the first feature matrix as a query vector and performing cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector includes:

[0014] Project the first feature matrix into the query space, and project the second feature matrix into the key space and value space respectively.

[0015] Calculate the similarity matrix between the query space and the key space, and normalize it using the Softmax function to obtain the attention weights.

[0016] Perform weighted summation on the value vectors based on the attention weights to obtain the target conditional vector.

[0017] Further, the step of injecting the target conditional vector into the noise preset network of a diffusion model with the initial input being the character image to be repaired and dynamically fusing the local features of the skip connection with the target conditional vector through a cross-attention block to generate enhanced features includes:

[0018] At the skip connection of the noise preset network of the diffusion model, use the first feature matrix as a query to match the key vector and value vector of the second feature matrix.

[0019] Calculate the context vector through scaled dot-product attention, splice it with the preset complete character image, and then generate enhanced features through convolution fusion.

[0020] Further, before the step of combining the weight scheduling of the skeleton prior and the stroke prior to generate the repaired character image, it also includes:

[0021] During the reverse denoising process of the diffusion model, a dynamic stroke generation mechanism is introduced to activate the stroke regions to be repaired in stages through a noise mask;

[0022] The constraint formula of the noise mask is:

[0023] The constraint formula of the noise mask is:

[0024]

[0025] where M stroke ∈ {0, 1} H×W is the mask of the stroke region to be repaired in the current stage; through the Hadamard product ⊙, denoising is enhanced in the masked region, and the original prediction is maintained in the non-masked region;

[0026] ∈0: that is, ∈ θ (x t , t, C fusion ) represents the new noise prediction, which is the updated result obtained after the network performs stronger repair on the stroke region at the current step.

[0027] ∈0(original): represents the original noise prediction, that is, the noise estimate output by the model before special processing of the stroke region.

[0028] ∈0(masked): represents the final output noise prediction, which adopts the new prediction value in the stroke region and maintains the original prediction value in the non-stroke region.

[0029] When M stroke (i, j) = 1, it represents the stroke part that needs to be repaired, and forces the model to apply the newly predicted noise ∈ θ (x t , t , C fusion ) to denoise the specified stroke region;

[0030] When M stroke (i, j) = 0, for the background part outside the text, the noise state of the previous time step is retained to avoid interfering with the repaired or un-repaired regions; The staged activation is to perform staged conditional weight scheduling through a dynamic weight formula;

[0031] The specific dynamic weight formula is:

[0032] The specific dynamic weight formula is:

[0033]

[0034] Early stage (t → T, high noise stage):

[0035] α(t) ≈ 1; The skeleton features are dominant, and the model first restores the global topological structure of Chinese characters to ensure the glyph standardization.

[0036] Mid - term (t = T / 2, transition stage):

[0037] α(t) ≈ 0.4; The weight of the skeleton features gradually decreases, and the stroke features begin to participate in the local morphology optimization.

[0038] Late - term (t → 0, low - noise stage):

[0039] α(t) ≈ 0; The stroke features are dominant, and the model focuses on the details reconstruction such as pen - tip sharpening and ink density to enhance the visual authenticity.

[0040] Furthermore, the steps for obtaining the first feature matrix include:

[0041] Input the preset complete character image into the skeleton graph encoder, and output the node feature matrix representing the global skeleton topology to obtain the first feature matrix; the skeleton graph encoder is constructed based on the graph attention network.

[0042] Furthermore, the implementation of the skeleton graph encoder includes:

[0043] Extract the binary skeleton image based on the preset complete character image;

[0044] Model the binary skeleton image as a graph structure G=(V, E);

[0045] Among them, the node set V represents the stroke intersections, and the edge set E represents the stroke connection relationships;

[0046] Aggregate the node neighborhood features through a multi - layer graph attention network to update the node features to the global skeleton feature matrix.

[0047] Furthermore, the steps for obtaining the second feature matrix include:

[0048] Input the preset complete character image into the stroke set encoder, and output the feature matrix encoding the local stroke morphology and temporal sequence dependencies to obtain the second feature matrix; the stroke set encoder is constructed by a Transformer encoder.

[0049] Furthermore, the implementation of the stroke set encoder includes:

[0050] Extract the stroke direction encoding sequence arranged in the writing order based on the preset complete character image to obtain a normalized stroke direction vector sequence;

[0051] Input the normalized stroke direction vector sequence, and introduce the sine position encoding to retain the writing temporal information;

[0052] Calculate the self-attention weights through a multi-layer Transformer encoder, and output the stroke timing feature matrix.

[0053] According to another aspect of the present application, there is provided an electronic device, including: a processor; and a memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the double-condition-guided diffusion model text repair method as described above.

[0054] According to still another aspect of the present application, there is provided a computer-readable storage medium on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the double-condition-guided diffusion model text repair method as described above.

[0055] An embodiment of the present invention provides a double-condition-guided diffusion model text repair method and apparatus. The method includes: obtaining a first feature matrix and a second feature matrix; wherein the first feature matrix is a global skeleton feature matrix modeled based on the text skeleton topology of a preset complete character image; the second feature matrix is a stroke timing feature matrix modeled based on the timing features and morphological details of a preset complete character image. Using the first feature matrix as a query vector, perform cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector; inject the target conditional vector into the noise preset network of a diffusion model with the initial input being the character image to be repaired, and dynamically fuse the local features of the skip connection with the target conditional vector through a cross-attention block to generate an enhanced feature; linearly combine the enhanced feature with the first feature matrix and the second feature matrix, and combine the weight scheduling of the skeleton prior and the stroke prior to generate the repaired character image.

[0056] The present invention respectively obtains the structural and stroke features through the global skeleton feature matrix and the stroke timing feature matrix, adopts cross-attention fusion of two priors at the skip connection to obtain the target conditional vector, replaces the original features of the network with the target conditional vector, and then injects the text constraint conditions. Then, through a step-by-step iterative process, while completing image denoising, the missing strokes are restored until a clear and complete character image is generated, so that the model can retain the details of the strokes and the text style features while generating high-quality repair results. Description of the Drawings

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 It is a schematic flow chart of the text repair method based on the dual-condition guidance diffusion model provided by the embodiment of the present invention;

[0059] Figure 2 It is a schematic structural diagram of an electronic device provided by the embodiment of the present invention. Specific embodiments

[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0062] Next, example embodiments according to the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0063] Application overview

[0064] The existing technologies in the field of text repair mainly include traditional text repair methods, deep learning-based image generation methods, and diffusion models.

[0065] Traditional text repair methods achieve repair through technologies such as stroke feature extraction, edge detection, and image segmentation, and perform well in slightly damaged images. However, they have limited effects when dealing with complex backgrounds, severely degraded images, and missing detailed strokes, and are prone to structural distortion (such as stroke breakage and adhesion) or semantic errors (such as incorrect combinations of Chinese character radicals), and the repair process lacks flexibility. The existing text prompt-based conditional control methods are difficult to accurately constrain the structural details of text and have insufficient robustness to complex glyphs or low-quality inputs.

[0066] In the image generation method based on deep learning, the VAE-based restoration method can generate natural restoration samples but lacks detail control. The GAN-based restoration method can generate realistic restoration images, but the generated results are vulnerable to the instability of adversarial training, rely on a large amount of labeled data, and the restored strokes are prone to blurring or artifacts, making it difficult to constrain the global structure. The CNN-based restoration method relies on local textures and lacks explicit modeling of the global skeleton and stroke order of text.

[0067] Diffusion models have shown outstanding performance in the field of image generation. However, in the existing technology, the resolution of the prior mask limits the fine restoration ability, and the restoration process lacks dynamic adjustment, making it impossible to precisely control the stroke direction and skeleton topology, resulting in distorted text structures. At the same time, there is a certain dependence on training data.

[0068] To address the above technical problems, the concept of this application is to propose a text restoration method and device based on a diffusion model with bi-conditional guidance, aiming to solve the problems of insufficient detail restoration, poor semantic consistency in complex scenarios, and structural distortion and lack of stroke details in degraded text restoration in traditional text restoration technologies, and to improve the structural rationality and visual coherence of the generated results. The core of the present invention lies in constraining the reverse denoising process of the diffusion model through a bi-conditional guidance mechanism (skeleton topology and stroke timing), so as to take into account the global structural rationality and the accuracy of local stroke details when restoring text images.

[0069] After introducing the basic principle of this application, various non-limiting embodiments of this application will be specifically introduced with reference to the accompanying drawings.

[0070] Exemplary Method

[0071] Figure 1 The flowchart of the text restoration method based on a diffusion model with bi-conditional guidance according to an embodiment of this application is illustrated.

[0072] As Figure 1 shown, the text restoration method based on a diffusion model with bi-conditional guidance according to an embodiment of this application includes:

[0073] S110, obtaining a first feature matrix and a second feature matrix;

[0074] Among them, the first feature matrix is a global skeleton feature matrix modeled based on the text skeleton topology of a preset complete character image; the second feature matrix is a stroke timing feature matrix modeled based on the timing features and morphological details of a preset complete character image;

[0075] S120, using the first feature matrix as a query vector, performing cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector;

[0076] S130. Inject the target conditional vector into the noise preset network of the diffusion model with the initial input being the character image to be repaired, and dynamically fuse the local features of the skip connection with the target conditional vector through the cross-attention block to generate enhanced features;

[0077] S140. Linearly combine the enhanced features with the first feature matrix and the second feature matrix, and combine the weight scheduling of the skeleton prior and the stroke prior to generate the repaired character image.

[0078] The extraction of skeleton and stroke features can model the text skeleton topology through a skeleton graph encoder (SGE) and model the stroke temporal features through a stroke set encoder (SSE).

[0079] Cross-modal conditional fusion can fuse the skeleton features and the stroke features into a conditional vector through a cross-attention mechanism.

[0080] Condition injection and dynamic generation are to embed the conditional vector into the U-Net architecture of the diffusion model and repair the missing strokes in stages by combining dynamic weight scheduling.

[0081] Next, each step will be described in detail.

[0082] In step S110, obtain the first feature matrix and the second feature matrix.

[0083] Among them, the first feature matrix is a global skeleton feature matrix modeled based on the text skeleton topology of the preset complete character image; the second feature matrix is a stroke temporal feature matrix modeled based on the temporal features and morphological details of the preset complete character image;

[0084] For example, input the preset complete character image into the skeleton graph encoder, and output the node feature matrix representing the global skeleton topology to obtain the first feature matrix; the skeleton graph encoder is constructed based on the graph attention network.

[0085] Furthermore, the implementation of the skeleton graph encoder includes:

[0086] Extract the binary skeleton image based on the preset complete character image;

[0087] Model the binary skeleton image as a graph structure G=(V, E);

[0088] Among them, the node set V represents the stroke intersection points, and the edge set E represents the stroke connection relationship;

[0089] Aggregate the node neighborhood features through a multi-layer graph attention network and update the node features to the global skeleton feature matrix.

[0090] Specifically, the Canny edge detection algorithm can be used to extract the text contour and generate an initial binary image; apply a thinning algorithm (such as the Zhang-Suen algorithm) to the binary image to obtain a single-pixel-width skeleton image.

[0091] Model the skeleton image as a graph structure G=(V,E), where the node set V represents the stroke intersection points and the edge set E represents the stroke connection relationship. The network uses a graph attention network to capture multi-scale spatial dependencies between nodes through multi-layer feature aggregation.

[0092] First, initialize the node features. The initial feature of each node υ i ∈V is generated by linear projection of its spatial coordinates (x,y).

[0093] At the L-th layer, aggregate the neighborhood features, and update the feature of node vi to The formula is as follows:

[0094]

[0095] where N(v) is the set of neighbor nodes, σ is the LeakyReLU activation function, is the learnable parameter matrix.

[0096] Finally, after passing through the L-layer graph attention network, the features of all nodes are updated, and the first feature matrix is finally output, which is also the global skeleton feature matrix H G ∈R N×128 (N represents the number of skeleton nodes), which characterizes the global skeleton topology for subsequent cross-modal fusion.

[0097] In an alternative embodiment, the implementation of the skeleton graph encoder (SGE) includes a network architecture and hierarchical division of labor.

[0098] In this embodiment, the skeleton graph encoder uses a four-layer graph attention network (GAT), and the number of neurons in each layer is {64, 64, 128, 128} in turn. Layers 1-2: Focus on local feature aggregation and capture the spatial relationship of stroke intersection points through neighborhood node feature interaction. Layers 3-4: Expand the scope of action and capture the global topological relationship across nodes. For example, the third layer aggregates the dependencies between distant nodes through a 128-dimensional feature map to ensure the global consistency of the skeleton topology.

[0099] After the global skeleton feature matrix is output, it can be mapped to a channel dimension (such as 512 dimensions) compatible with the diffusion model U-Net through a linear projection layer as the query input for cross-attention fusion.

[0100] The interaction between the skeleton features and the diffusion model can be carried out by obtaining the binary skeleton graph, then performing graph structure modeling to initialize the node features, and then performing feature aggregation in the GAT layer. After obtaining the global skeleton feature matrix (for example, N×128), it is used as the query vector for cross-attention fusion.

[0101] Optionally, the steps for obtaining the second feature matrix include:

[0102] Input the preset complete character image into the stroke set encoder, and output the feature matrix encoding the local stroke morphology and temporal dependence relationship to obtain the second feature matrix; the stroke set encoder is constructed by a Transformer encoder.

[0103] Furthermore, the implementation of the stroke set encoder includes:

[0104] Extract the stroke direction encoding sequence arranged in the writing order based on the preset complete character image to obtain the normalized stroke direction vector sequence;

[0105] Input the normalized stroke direction vector sequence, and introduce the sine position encoding to retain the writing timing information;

[0106] Calculate the self-attention weights through multiple layers of Transformer encoders and output the stroke temporal feature matrix.

[0107] Specifically, first, decompose the text into strokes based on the skeleton graph, and arrange them in the writing order as S={s1, s2,..., sM}, where M is the number of strokes;

[0108] Secondly, for each stroke sk, calculate its main direction (such as horizontal, vertical, left-falling, right-falling) and normalize it to the direction vector dk∈R4dk∈R4. For example:

[0109] Horizontal stroke: dk = [1, 0, 0, 0];

[0110] Vertical stroke: dk = [0, 1, 0, 0];

[0111] And so on for the other directions.

[0112] Then, the network is constructed based on six layers of Transformer encoders. To retain the stroke writing timing information, the sine position encoding s i ∈R 4 is introduced. After the stroke features are concatenated, they are input into the encoder. The attention weight calculation for the k-th layer is:

[0113]

[0114] where, Q (k) , K (h) , V (k)They are generated by linear projection of stroke features, and the attention dimension is d k =64.

[0115] After attention calculation and full connection layer, the second feature matrix is ​​finally output, which is the stroke time sequence feature matrix H S ∈R M×128 (M represents the number of strokes), encoding the local stroke shape and temporal dependency.

[0116] In an optional embodiment, the stroke set encoder aims to convert the stroke sequence arranged in writing order into a high-dimensional feature representation by modeling the temporal characteristics and morphological details of the Chinese character strokes, and its input is the stroke direction encoding sequence S={s1, s2, ..., s M}, where each stroke represents the normalized direction vector of horizontal, left-falling, vertical, and right-falling strokes, and the initial stroke feature matrix P∈R with a dimension of 64 is generated by linear projection M×64 .

[0117] In this embodiment, the stroke set encoder is built based on a six-layer Transformer encoder. The query (Q) of each stroke is calculated with the key (K) of all strokes, and the attention weight is generated through Softmax.

[0118] The six-layer transformer is used to capture the stroke series dependency and avoid overfitting, including 1-2 shallow layers: learning the connection direction of adjacent strokes; 3-4 transition layers: temporal dependency across strokes; 5-6 deep layers: learning the global sequence (the writing rules of the entire Chinese character).

[0119] The six-layer Transformer encoder also contains four self-attention heads, each of which has the following functions:

[0120] Head 1 (directional similarity): calculate the cosine similarity of the stroke direction vectors, for example, the attention weight of horizontal strokes is higher between strokes with similar directions; Head 2 (temporal proximity): capture the temporal dependency of adjacent strokes, such as the writing order of "horizontal" followed by "vertical"; Head 3 (morphological complementarity): analyze the morphological complementarity features of symmetrical structures (such as the left and right strokes of the character "口"); Head 4 (cross-stroke structural constraints): model the spatial relationship of enclosing and semi-enclosing structures, such as the outer frame and internal strokes of the character "国".

[0121] It should be noted that the attention dimension d of each head k = 64, gradually enhancing the context representation of each stroke so that it contains both its own morphological information and the global sequence features. The self-attention mechanism cannot capture the positional relationship between strokes, so we introduce the sinusoidal position encoding s i ∈R 4, mainly responsible for stroke positioning. The periodicity of the sine function can capture both short-distance and long-distance temporal relationships. To retain the stroke writing temporal information, the stroke features are concatenated and then input into the encoder. The attention weight calculation for the k-th layer is as follows:

[0122]

[0123] where Q (k) , K (k) , V (k) P ∈ R M×64 are respectively generated by linear projection of the stroke features.

[0124] After each layer of self-attention, two fully connected networks are connected, with dimensions of 256 and 128 respectively. The first layer: dimension 64 - 256, expansion, learning more complex text structures; the second layer: 256 - 128, retaining key information and then facilitating the features for the final output to align with H G . The activation function is GELU, enhancing the non-linear expression ability.

[0125] After the attention calculation and the fully connected layer, the final output is the stroke temporal feature matrix H S ∈ R M×128 (M represents the number of strokes), encoding the local stroke morphology and temporal dependence relationship.

[0126] After the stroke temporal feature matrix is output, the dimension can be adjusted through a fully connected layer (e.g., 128 → 512), serving as the key (K) and value (V) of the cross-attention module.

[0127] It should be noted that to eliminate the modal differences between the skeleton and the strokes, in the cross-attention fusion, strategies such as spatial alignment can be adopted, mapping the skeleton node coordinates and the stroke direction vectors to the same high-dimensional space. Or temporal alignment, retaining the stroke writing order through sine position encoding and dynamically matching it with the static structure of the skeleton topology.

[0128] In step S120, the first feature matrix is used as the query vector, and cross-attention fusion is performed with the key vector and value vector of the second feature matrix to generate the target conditional vector.

[0129] To better capture the complementary features of the skeleton and the strokes in terms of spatial topology and temporal morphology, and eliminate the modal differences between the skeleton and stroke features, a cross-attention strategy can be adopted. The first feature matrix and the second feature matrix are regarded as input features of different modalities, and are respectively projected into the query (Q), key (K), and value (V) spaces. The skeleton features are used as the query, and are matched with the key and value of the stroke features to capture the correlation between the two:

[0130]

[0131] Among them, W Q , W K , W V ∈R 128×128 is a learnable projection matrix.

[0132] Calculate the correlation through dot product and normalize it to obtain the attention weights, and use the attention weights to weight the stroke value vectors to generate the fused target conditional vector.

[0133] Optionally, in the method for text repair of the diffusion model guided by dual conditions provided in the embodiments of the present application, the step of using the first feature matrix as a query vector to perform cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector includes:

[0134] Project the first feature matrix into the query space, and project the second feature matrix into the key space and value space respectively;

[0135] Calculate the similarity matrix between the query space and the key space, and normalize it using the Softmax function to obtain the attention weights;

[0136] Weight and sum the value vectors based on the attention weights to obtain the target conditional vector.

[0137] Specifically, take the skeleton feature matrix as the query (Query), the stroke feature matrix as the key (Key) and value (Value), and fuse the two types of features through the cross-attention mechanism. Among them, the skeleton features are linearly projected to generate the query vector Q, and the stroke features are projected into the key vector K and value vector V respectively. After calculating the dot product similarity matrix between Q and K, the attention weights are generated through Softmax normalization. For example, if the number of skeleton nodes is 50 and the number of strokes is 8, the size of the attention weight matrix is 50×8, indicating the attention degree of each skeleton node to each stroke.

[0138] Subsequently, use the attention weights to weight and sum the value vector V to generate the fused target conditional vector. For example, a certain skeleton node may have a higher weight (such as 0.8) for the key vector of the 3rd stroke, then the conditional vector of this node will mainly fuse the morphological information of the 3rd stroke. This conditional vector contains both the topological constraints of the skeleton and the temporal details of the strokes, providing a structured prior for the subsequent diffusion model.

[0139] In an alternative embodiment, using the first feature matrix as a query vector to perform cross-attention fusion with the key vector and value vector of the second feature matrix includes:

[0140] First, construct a linear projection layer to project the skeleton features (Q) and stroke features (K / V) into a 512-dimensional space respectively.

[0141] Next is to perform scaled dot-product attention calculation, for example:

[0142]

[0143] where dk = 64 is the attention dimension, and after Softmax normalization, a 50×8 weight matrix (number of skeleton nodes × number of strokes) is generated;

[0144] Then is to perform context vector generation, weighted sum the value vectors, and output the conditional vector

[0145] In addition, the conditional vector can be, for example, through feature replacement. At the third layer skip connection of the U-Net, the original feature map F(64×64×256) is concatenated with the conditional vector, and fused into an enhanced feature map F_fused(64×64×512) through 1×1 convolution; and for example, through dynamic weight scheduling, using, for example, a cosine function to adjust the weight ratio of the skeleton and strokes, and then obtaining the fused target conditional vector.

[0146] In step S130, inject the target conditional vector into the noise preset network of the diffusion model with the initial input being the character image to be repaired, and dynamically fuse the local features of the skip connection with the target conditional vector through the cross-attention block to generate enhanced features.

[0147] The image to be repaired is used as the initial input of the diffusion model, gradually denoised in the reverse process, and combined with the corresponding conditional vector to generate the repair result.

[0148] In this embodiment, the skip connection of the noise preset network U-Net of the diffusion model can directly transfer the feature map of a certain layer of the encoder (for example, a 4-layer downsampling module, with the number of channels in each layer being {64, 128, 256, 512} in turn, using stride 2 convolution and LeakyReLU activation) to the corresponding layer of the decoder (for example, a 4-layer upsampling module, restoring the resolution through bilinear interpolation and convolution, and the number of channels in each layer is symmetric with the encoder), retaining multi-scale spatial information to retain high-resolution information. It should be noted that in this embodiment, the fusion condition C (external condition information) can provide high-level conditions regarding the overall structure and detailed style of the target image.

[0149] Therefore, a cross-attention block is introduced at the skip connection, using C as the key (Key) and value (Value), and the local feature F of the skip connection as the query (Query) to achieve dynamic fusion of information, enabling the decoder to adaptively adjust the feature weights according to the prior information and better guide the reconstruction process.

[0150] First, the feature of the current jump connection is F ∈ R H×W×d , with a spatial resolution of H×W and the number of channels being D. At the same time, the external prior condition vector C ∈ R N×d , with a sequence length of N and a feature dimension of d. Next, three learnable linear mappings are used to map the encoder features and the condition vector to the Query, Key, and Value spaces respectively. Using the scaled dot-product attention formula, the similarity between the encoder features (queries) and the condition vector (keys) is calculated, and the attention weights A are generated through Softmax normalization:

[0151]

[0152] where the weight A h,w,n represents the dependence degree of the encoder feature at position (h, w) on the nth element of the condition vector.

[0153] Then, the values (Values) of the condition vector are weighted and summed using the attention weights to generate the context vector Context. Among them, this context vector carries the global prior information related to the original features. The context vector is fused with the original feature map. Next, the original feature map and the reshaped context tensor are concatenated in the channel dimension to ensure that both parts of the information (local features and global prior) are retained. To automatically learn how to integrate the concatenated feature information, a 1*1 convolutional layer can be used on the concatenated result for fusion, and finally a new feature map F fused is generated, and the specific formula is as follows:

[0154] F fused = σ(Concat(F, Context reshaped ) * W 1×1 + b) (5);

[0155] The generated F fused not only contains the local detailed features of F but also fully embeds the global context information obtained through the attention mechanism. The fused feature F fused is passed to the corresponding layer of the decoder through the skip connection and concatenated or added to the upsampled features of the decoder to drive the detail recovery.

[0156] Optionally, in the method for text repair of the diffusion model based on dual-condition guidance provided in the embodiments of the present application, when injecting the target condition vector into the noise preset network of the diffusion model with the initial input being the character image to be repaired, the local features of the skip connection and the target condition vector are dynamically fused through the cross-attention block to generate enhanced features, including:

[0157] At the skip connection of the noise prediction network of the diffusion model, the first feature matrix is used as a query to match the key vectors and value vectors of the second feature matrix;

[0158] For example, the context vector is calculated through scaled dot product attention, and after being concatenated with the preset complete character image, it is fused through convolution to generate enhanced features.

[0159] Specifically, in the noise prediction network of the diffusion model, the target conditional vector is dynamically injected into the skip connection of the U-Net. Taking the third-layer skip connection of the U-Net as an example, the size of the local feature map output by the encoder is 64×64×256. After the conditional vector C is reshaped into the same spatial size, it is fused with the local feature through the cross-attention block. Specifically, the local feature is used as a query, and the conditional vector is used as the key and value to calculate the scaled dot product attention. For example, the query vector at a certain position (h, w) in the local feature calculates the similarity with all the key vectors of the conditional vector to generate the attention weight matrix, and then the value vectors are weighted and summed to obtain the context vector.

[0160] Finally, the context vector is concatenated with the original local feature, and the channel information is fused through a 1×1 convolutional layer. For example, the number of channels of the original feature is 256, and the number of channels of the context vector is 128. After concatenation, it is reduced to 256 channels through convolution to generate an enhanced feature map, obtaining enhanced features to guide the detailed reconstruction of the missing area.

[0161] In step S140, the enhanced feature is linearly combined with the first feature matrix and the second feature matrix, and combined with the weight scheduling of the skeleton prior and the stroke prior to generate the repaired character image.

[0162] In this embodiment, by integrating the enhanced feature into the noise prediction network, the global skeleton feature matrix H G and the stroke timing feature matrix H S of the conditional information.

[0163] The improved Gaussian distribution can be:

[0164]

[0165] Correspondingly, the mean formula can be changed to:

[0166]

[0167] The variance remains unchanged:

[0168]

[0169] Where: the conditional reverse distribution p θ (x t-1 |x t, the mean and variance of C) are determined by ∈ θ (x t , t, C fusion ) are determined; the condition C = C fueion = Fusion(H G , H S ):

[0170] The sampling formula for each step of the reverse process can be:

[0171]

[0172] Substituting into the mean expression, x t-1 can be updated:

[0173]

[0174] Through a step-by-step iterative process, while completing image denoising, the missing strokes are restored until a clear and complete repaired character image is generated.

[0175] Optionally, before generating the repaired character image in the weight scheduling of combining the skeleton prior and the stroke prior in the text repair method based on the bi-conditional guided diffusion model provided in the embodiments of the present application, it further includes:

[0176] In the reverse denoising process of the diffusion model, the present invention introduces a noise mask mechanism to enhance the refined control of the Chinese character stroke order by dynamically constraining the noise prediction range.

[0177] The constraint formula of the noise mask is:

[0178]

[0179] Among them, M stroke ∈ {0, 1} H×W is the stroke area mask to be repaired in the current stage; through the Hadamard product ⊙, denoising is enhanced in the masked area, and the original prediction is maintained in the non-masked area;

[0180] ∈0: That is, ∈ θ (x t , t, C fusion ) represents the new noise prediction, which is the updated result obtained after the network performs stronger repair on the stroke area at the current step.

[0181] ∈0(original): Represents the original noise prediction, that is, the noise estimate output by the model before special processing of the stroke area.

[0182] ∈0(masked): Represents the final output noise prediction, which adopts the new prediction value in the stroke area and maintains the original prediction value in the non-stroke area.

[0183] When M stroke (i, j) = 1, it indicates the stroke part to be repaired, and the forced model applies the newly predicted noise ∈ θ (x t , t, C fusion ) to denoise the specified stroke area;

[0184] When M stroke (i, j) = 0, for the background part outside the text, retain the noise state of the previous time step to avoid interfering with the repaired or unrepaired areas.

[0185] The phased activation is to perform phased conditional weight scheduling through a dynamic weight formula;

[0186] The specific dynamic weight formula is:

[0187]

[0188] Early stage (t → T, high noise stage):

[0189] α(t) ≈ 1; The skeleton features dominate, and the model preferentially restores the global topological structure of Chinese characters to ensure the glyph standardization.

[0190] Middle stage (t = T / 2, transition stage):

[0191] α(t) ≈ 0.4; The weight of the skeleton features gradually decreases, and the stroke features begin to participate in the local morphology optimization.

[0192] Late stage (t → 0, low noise stage):

[0193] α(t) ≈ 0; The stroke features dominate, and the model focuses on details reconstruction such as pen tip sharpening and ink concentration to improve visual authenticity.

[0194] Optionally, use the cosine scheduling function α(t) to control the weight of the skeleton features, and 1 - α(t) is used to adjust the weight of the stroke features:

[0195]

[0196] where: t ∈ {T, T - 1,..., 0} represents the current time step of reverse denoising, and T is the maximum time step;

[0197] Initial stage (t ≈ 0): The weight decreases slowly, and the skeleton features are given priority.

[0198] Late stage (t ≈ T): The weight decreases rapidly, and the detail features dominate.

[0199] Specifically, in the character restoration task, the stroke order often has an important impact on the final visual effect. If denoising is directly performed on the entire character image simultaneously, it may cause strokes to cover each other or the order to be chaotic. Therefore, this application introduces a "stroke-order-guided noise mask" in the original diffusion denoising process, enabling the network to focus on restoring the corresponding stroke regions at different stages, while keeping other completed or untimely stroke regions relatively stable.

[0200] In the reverse denoising process of the diffusion model, to achieve smoother feature fusion, this application adopts a dynamic weight scheduling mechanism, aiming to coordinate the contribution ratios of the skeleton features (global structure) and stroke features (local details), ensuring that the restoration process conforms to the "structure-first, details-progressive" logic of Chinese character writing. For example, when restoring a damaged character "horse", first generate a noise mask and activate the stroke regions to be restored in stages according to the writing order. At the early time steps (t approaching the total number of steps T), the skeleton prior weight is relatively high (e.g., λ(t) = 0.9), and the model first restores the overall structure of the character; at the later time steps (t approaching 0), the stroke prior weight gradually increases (e.g., λ(t) = 0.1), and the model focuses on refining the stroke edges.

[0201] Through iterative sampling, the noise is gradually removed and the strokes are restored. For example, at the 100th reverse sampling, the model reconstructs the structure of the horizontal fold hook of the character "horse" according to the conditional vector; at the 20th sampling, the pen tip details of the hook are refined, and finally a complete character image is generated.

[0202] This application uses the noise mask to enhance denoising in the stroke stages in the spatial dimension and smoothly transition from the skeleton prior to the stroke prior in the time dimension. During the diffusion reverse sampling process, each stroke is sequentially activated and restored, and the overall structure is ensured to take shape in the early stage and the stroke details are polished in the later stage. This "dynamic stroke generation mechanism" not only makes the restoration process more interpretable and controllable (able to simulate the human writing order), but also effectively improves the layering and detail accuracy of the restoration effect, achieving significant improvements in both subjective visual evaluation and quantitative metrics.

[0203] In this way, in the text restoration method based on the dual-conditional-guided diffusion model according to the embodiments of this application, the global skeleton feature matrix obtained by modeling based on the text skeleton topology of the preset complete character image accurately reconstructs the symmetric topology structure of the tripod body; the stroke temporal feature matrix obtained by modeling based on the temporal sequence features and morphological details of the preset complete character image captures the starting and ending directions of the existing strokes in the left half; furthermore, cross-attention fusion is used to ensure that the newly generated strokes in the right half are consistent with the style of the left half; at the same time, a dynamic generation mechanism is adopted to repair the tripod feet and tripod ears in stages, avoiding stroke adhesion.

[0204] In the Chinese character test set, the structural error rate is significantly reduced, and the success rate of reconstructing broken strokes is significantly increased, ensuring the coherence of the pen tips of calligraphic characters. The experimental results show that the repaired results are highly consistent with the distribution of real characters, and the edge sharpness is significantly improved. Compared with the original diffusion model, the inference speed is increased, the convergence is accelerated due to conditional guidance, and the data and computing power requirements are reduced compared with the original diffusion model, achieving efficiency and resource optimization.

[0205] Through the skeleton-stroke dual-conditional guidance mechanism, the present invention is superior to the prior art in terms of structural accuracy, image quality, efficiency, and resource optimization, providing key technical support for cultural heritage protection and industrial intelligence.

[0206] Exemplary electronic device

[0207] Next, refer to Figure 2 to describe the electronic device according to an embodiment of the present application. The electronic device may be an electronic device integrated with a first imaging device, or a stand-alone device independent of the first imaging device, and the stand-alone device may communicate with the first imaging device to receive the input signal collected therefrom.

[0208] Figure 2 The block diagram of the electronic device according to an embodiment of the present application is illustrated.

[0209] As Figure 2 shown, the electronic device 10 includes one or more processors 11 and a memory 12.

[0210] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0211] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run the program instructions to implement the dual-conditional guidance-based diffusion model text repair method of the various embodiments of the present application described above and / or other desired functions.

[0212] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0213] In addition, the input device 13 may further include, for example, a keyboard, a mouse, and so on.

[0214] The output device 14 may output various information to the outside. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and a remote output device connected thereto, and so on.

[0215] Of course, for simplicity, Figure 2 only some of the components related to the present application in the electronic device 10 are shown in [the figure], and components such as a bus, an input / output interface, and so on are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.

[0216] Exemplary computer program products and computer-readable storage media.

[0217] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the diffusion model text repair method based on a biconditional guidance described in the above "Exemplary Method" section of this specification according to various embodiments of the present application.

[0218] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0219] In addition, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when run by a processor, cause the processor to execute the steps in the moving object tracking method described in the above "Exemplary Method" section of this specification according to various embodiments of the present application.

[0220] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0221] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. Additionally, the above-disclosed specific details are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.

[0222] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.

[0223] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0224] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0225] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present application to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A text repair method based on a double - conditional guidance diffusion model, characterized in that, Including: Obtain a first feature matrix and a second feature matrix; The first feature matrix is a global skeleton feature matrix modeled based on the text skeleton topology of a preset complete character image; The second feature matrix is a stroke timing feature matrix modeled based on the timing features and morphological details of a preset complete character image; Use the first feature matrix as a query vector, and perform cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector; Inject the target conditional vector into the noise preset network of a diffusion model with an initial input of a character image to be repaired, and dynamically fuse the local features of the skip connection with the target conditional vector through a cross-attention block to generate enhanced features; Linearly combine the enhanced features with the first feature matrix and the second feature matrix, and combine the weight scheduling of the skeleton prior and the stroke prior to generate a repaired character image.

2. The method for text repair of the diffusion model guided by dual conditions as described in claim 1, wherein, The step of using the first feature matrix as a query vector, and performing cross-attention fusion with the key vector and value vector of the second feature matrix to generate a target conditional vector includes: Project the first feature matrix into a query space, and project the second feature matrix into a key space and a value space respectively; Calculate the similarity matrix between the query space and the key space, and normalize it using the Softmax function to obtain attention weights; Weight-sum the value vectors based on the attention weights to obtain a target conditional vector.

3. The method for text repair of a diffusion model based on dual-condition guidance according to claim 1, wherein, The step of injecting the target conditional vector into the noise preset network of a diffusion model with an initial input of a character image to be repaired, and dynamically fusing the local features of the skip connection with the target conditional vector through a cross-attention block to generate enhanced features includes: At the skip connection of the noise preset network of the diffusion model, use the first feature matrix as a query to match the key vector and value vector of the second feature matrix; Calculate the context vector through scaled dot-product attention, splice it with a preset complete character image, and generate enhanced features through convolution fusion.

4. The method for text repair of a diffusion model based on dual-condition guidance according to claim 1, wherein, Before the step of combining the weight scheduling of the skeleton prior and the stroke prior to generate a repaired character image, it further includes: During the reverse denoising process of the diffusion model, introduce a dynamic stroke generation mechanism, and activate the stroke area to be repaired in stages through a noise mask; The constraint formula of the noise mask is: The constraint formula of the noise mask is: Among them, M stroke ∈ {0, 1} H×W is the stroke area mask to be repaired in the current stage; through the Hadamard product ⊙, denoising is enhanced in the masked area, and the original prediction is maintained in the non-masked area; ∈0: namely, ∈ θ (x t , t, C fusion ), representing the new noise prediction, which is the updated result obtained after the network performs stronger repair on the stroke area at the current step; ∈0(original): Represents the original noise prediction, that is, the noise estimate output by the model before special processing of the non-stroke area; ∈0(masked): Represents the final output noise prediction, which uses a new predicted value in the stroke area and maintains the original predicted value in the non-stroke area; When M stroke (i, j) = 1, it indicates the stroke part to be repaired, and forces the model to apply the newly predicted noise ∈ θ (x t , t, C fusion ) to denoise the specified stroke area; When M stroke (i, j) = 0, the noise state of the background part other than the text at the previous time step is retained to avoid interfering with the repaired or unrepaired areas; ​ The staged activation is to perform staged conditional weight scheduling through a dynamic weight formula; The specific dynamic weight formula is: In the early stage (t → T, high-noise stage): α(t) ≈ 1; The skeleton feature dominates, and the model first restores the global topology of Chinese characters to ensure the normality of the glyph; In the middle stage (t = T / 2, transition stage): α(t) ≈ 0.4; The weight of the skeleton feature gradually decreases, and the stroke feature begins to participate in local morphology optimization; In the late stage (t → 0, low-noise stage): α(t)≈0; The stroke features are dominant, and the model focuses on reconstructing details such as pen tip sharpening and ink shading to enhance visual authenticity.

5. The text repair method based on the diffusion model with dual-condition guidance as claimed in claim 1, wherein, The steps for obtaining the first feature matrix include: Input the preset complete character image into the skeleton graph encoder to output a node feature matrix representing the global skeleton topology, thereby obtaining the first feature matrix; the skeleton graph encoder is constructed based on the graph attention network.

6. The text repair method based on the diffusion model guided by dual conditions according to claim 5, wherein The implementation of the skeleton graph encoder includes: Extract the binary skeleton image based on the preset complete character image; Model the binary skeleton image as a graph structure G=(V, E); Among them, the node set V represents the stroke intersection points, and the edge set E represents the stroke connection relationship; Aggregate the node neighborhood features through a multi-layer graph attention network to update the node features to the global skeleton feature matrix.

7. The text repair method based on the diffusion model guided by double conditions according to claim 1, characterized in that, The steps for obtaining the second feature matrix include: Input the preset complete character image into the stroke set encoder to output a feature matrix encoding the local stroke morphology and temporal dependence relationship, thereby obtaining the second feature matrix; the stroke set encoder is constructed by the Transformer encoder.

8. The method for text restoration of a diffusion model based on dual-condition guidance according to claim 7, wherein The implementation of the stroke set encoder includes: Extract the stroke direction encoding sequence arranged in the writing order based on the preset complete character image to obtain a normalized stroke direction vector sequence; Input the normalized stroke direction vector sequence and introduce sinusoidal position encoding to retain the writing temporal information; Calculate the self-attention weights through a multi-layer Transformer encoder and output the stroke temporal feature matrix.

9. An electronic device, comprising: Processor; And a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the text repair method based on the dual-condition guided diffusion model according to any one of claims 1-8.

10. A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor executes the text repair method based on the dual-condition guided diffusion model according to any one of claims 1-8.

Citation Information

Patent Citations

  • Chinese character image restoration algorithm based on skeleton extraction and adversarial learning

    CN114742714A

  • Text image restoration model and method based on structural attention and text perception

    CN116258652A

  • Skeleton extraction-based two-stage damaged individual character repairing method, system and equipment

    CN118261831A

  • Apparatus and method for image recognition processing

    KR101959831B1

Cited By

  • Super-resolution remote sensing image reconstruction method, system and device, and storage medium

    CN120765466A

  • Industrial image automatic labeling method and device, equipment and storage medium

    CN120997834A

  • Vectorization Chinese character graph generation method based on large model

    CN121010668A

  • A large model-based vectorized Chinese character pattern generation method

    CN121010668B