Text rendering method and device based on proxy confrontation, equipment and storage medium
Through the proxy adversarial-based text rendering method, combined with adversarial loss, perceptual consistency loss and reconstruction loss, the problems of efficiency and detail preservation in text rendering are solved, and efficient and complete text rendering effects are achieved.
Patent Information
- Application Number
- CN202510718185.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-16
AI Technical Summary
Existing text rendering technologies find it difficult to strike a balance between efficiency and detail preservation, especially the insufficient preservation of geometric features at the stroke connections and intersections of the glyph structure. Traditional methods also fail to effectively adapt to the requirements of multi-scale features in text rendering.
A proxy adversarial text rendering method is adopted. The adversarial loss, perceptual consistency loss and reconstruction loss are calculated in parallel through the rendering model. Combined with the dynamic weight adjustment mechanism, the target rendering model is generated to reduce the computational complexity and maintain the integrity of the glyph structure.
It improves the efficiency and quality of text rendering, ensures the structural integrity and style consistency of stroke joints, adapts to multi-scale rendering scenarios, and reduces model complexity.
Smart Images

Figure CN120655769A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology and can be applied to the fields of medical health and financial payment, and in particular to a text rendering method, apparatus, device and storage medium based on proxy confrontation. Background Art
[0002] Text rendering technology is an active research area, aiming to transform discrete text symbols into high-quality visual representations. This technology can be applied in the fields of fintech and healthcare, for example, in electronic health record systems or in mobile banking and payment scenarios. Existing research has primarily explored two approaches to address these challenges. On the one hand, attempts have been made to introduce knowledge distillation techniques, simulating the behavior of complex discriminators through lightweight models. However, such approaches often struggle to balance efficiency and detail preservation in text rendering tasks, particularly for critical geometric features such as stroke connections and intersections in glyph structures. On the other hand, perceptual losses have been employed to replace some adversarial losses, leveraging pre-trained convolutional neural networks to extract high-level visual features. While these approaches reduce computational complexity, they suffer from significant limitations in glyph structural consistency, such as blurring or discontinuity at stroke connections. Recent research has proposed using autoencoders to construct compact feature spaces to enhance local details. However, these loss functions fail to consider the dynamic nature of adversarial training, resulting in a lack of stylistic diversity in the generated results.
[0003] Although the proxy functions of existing methods reduce computational costs, they are not optimized for the unique topological structure of text; existing replacement losses focus more on global feature matching and ignore the redundancy problem caused by the local geometric constraints of glyphs; traditional proxy losses are difficult to adapt to the needs of multi-scale features in text rendering. The key difference between text rendering and general image generation tasks is that: glyph structures have strict topological constraints, and small geometric deformations may lead to significant changes in semantic understanding; text rendering quality assessment needs to consider not only pixel-level similarity, but also structural coherence and style consistency; rendering systems usually need to adapt to multi-scale scenarios from single words to paragraphs. These characteristics make it difficult for general image generation methods to be directly applied to high-quality text rendering. Therefore, there is an urgent need for a text rendering method that can improve text rendering efficiency while maintaining rendering quality. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a text rendering method, apparatus, device and storage medium based on proxy confrontation to improve the quality and efficiency of text rendering.
[0005] In order to solve the above technical problems, an embodiment of the present application provides a text rendering method based on proxy confrontation, including:
[0006] Get glyph encoding, style parameters and target text image;
[0007] Performing text rendering based on the glyph encoding and the style parameters by a rendering model to generate an initial text image;
[0008] Parallel calculation of adversarial loss, perceptual consistency loss, and reconstruction loss based on the initial text image and the target text image;
[0009] Calculating a total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retraining the rendering model based on the total model loss to generate a target rendering model;
[0010] Obtaining a to-be-processed glyph code and style parameters, and performing text rendering based on the to-be-processed glyph code and style parameters through a target rendering model to generate a target text rendering image.
[0011] In order to solve the above technical problems, an embodiment of the present application provides a text rendering device based on proxy confrontation, comprising:
[0012] A data acquisition module, used to obtain glyph encoding, style parameters and target text image;
[0013] A text rendering module, configured to perform text rendering based on the glyph encoding and the style parameters through a rendering model to generate an initial text image;
[0014] a loss calculation module, configured to concurrently calculate adversarial loss, perceptual consistency loss, and reconstruction loss based on the initial text image and the target text image;
[0015] a model training module, configured to calculate a total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retrain the rendering model based on the total model loss to generate a target rendering model;
[0016] The image generation module is used to obtain the to-be-processed glyph code and style parameters, and perform text rendering based on the to-be-processed glyph code and style parameters through a target rendering model to generate a target text rendering image.
[0017] To solve the above technical problems, a technical solution adopted by the present invention is: providing a computer device, including one or more processors; a memory for storing one or more programs, so that the one or more processors implement any of the above-mentioned proxy confrontation-based text rendering methods.
[0018] To solve the above technical problems, the present invention adopts a technical solution: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-mentioned proxy confrontation-based text rendering methods.
[0019] The embodiment of the present invention provides a text rendering method, device, equipment and storage medium based on proxy adversarial. The method includes: obtaining glyph encoding, style parameters and target text image; performing text rendering based on the glyph encoding and the style parameters by a rendering model to generate an initial text image; calculating adversarial loss, perceptual consistency loss and reconstruction loss in parallel based on the initial text image and the target text image; calculating the total model loss based on the adversarial loss, the perceptual consistency loss and the reconstruction loss, and retraining the rendering model based on the total model loss to generate a target rendering model; obtaining the glyph encoding and style parameters to be processed, and performing text rendering based on the glyph encoding and style parameters to be processed by a target rendering model to generate a target text rendering image. The embodiment of the present invention calculates the adversarial loss, perceptual consistency loss and reconstruction loss of the text image after rendering the text, and combines multiple losses to train the rendering model, which is conducive to improving the instructions of text rendering. In addition, the present application does not need to rely on the discriminator network to calculate the loss calculation, which reduces the complexity of the model and is conducive to improving the efficiency of text rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 1 is a schematic diagram of an application environment of a text rendering method based on proxy confrontation in one embodiment of the present invention;
[0022] Figure 2 This is a flowchart of the implementation process of the text rendering method based on proxy confrontation provided in an embodiment of the present application;
[0023] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S2;
[0024] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S3;
[0025] Figure 5 yes Figure 4A schematic flow chart of a specific implementation of step S31;
[0026] Figure 6 yes Figure 4 A schematic flow chart of a specific implementation of step S32;
[0027] Figure 7 yes Figure 2 A schematic flow chart of a specific implementation of step S4;
[0028] Figure 8 Schematic diagram of a text rendering device based on proxy confrontation provided in an embodiment of the present application;
[0029] Figure 9 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0032] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0033] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0034] It should be noted that the proxy confrontation-based text rendering method provided in the embodiments of the present application is generally executed by a server. Accordingly, the proxy confrontation-based text rendering device is generally configured in the server.
[0035] The text rendering method based on proxy confrontation provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can receive the glyph encoding and style parameters from the client; and generate a text rendering image according to the glyph encoding and style parameters. The server in the present invention sends the text rendering image to the client. The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0036] The proxy adversarial text rendering method provided in the embodiments of the present application can be applied to electronic health record systems, or can be applied to mobile banking and payment application scenarios.
[0037] Existing text rendering methods based on generative adversarial networks suffer from low training efficiency, large model sizes, and high training instability. Traditional methods rely on complex discriminator networks for adversarial training, making real-time deployment difficult on mobile devices. Knowledge distillation techniques, while enabling lightweight rendering, struggle to balance efficiency with the preservation of glyph details. Perceptual loss methods can easily lead to blurring or breakage at stroke junctions. Existing technologies lack targeted optimization for the unique topological constraints and multi-scale feature requirements of text, making it difficult to achieve efficient computation while maintaining glyph structural integrity.
[0038] In order to solve the above problems, the computational complexity of the discriminator in traditional adversarial training is an efficiency bottleneck, and the use of a lightweight proxy network to replace the complex discriminator is considered; in order to solve the problem that the glyph structure is easily distorted, it is proposed to combine the perceptual consistency loss with the adversarial loss; in order to solve the problem that the static loss weight is difficult to adapt to the training dynamics, a dynamic weight adjustment mechanism is introduced. Through a multi-dimensional loss collaborative optimization framework, the consistency of the stroke structure is guaranteed while reducing the computational overhead, achieving a balance between quality and efficiency. Therefore, the present application proposes to obtain glyph encoding, style parameters and target text images; perform text rendering based on the glyph encoding and style parameters through a rendering model to generate an initial text image; calculate the adversarial loss, perceptual consistency loss and reconstruction loss in parallel based on the initial text image and the target text image; calculate the total loss of the model based on the three types of losses and retrain the rendering model to generate a target rendering model; finally, the target rendering model is used to process the input parameters to generate a target text rendering image. The present application effectively reduces the training computational overhead of the text rendering model, enabling it to achieve real-time inference on mobile devices; maintains the structural integrity of the glyph stroke connections, avoiding the breakage or blurring problems caused by traditional methods; improves the stability of the training process through dynamic loss weight adjustment, and reduces the pattern collapse phenomenon.
[0039] See also Figure 2 , Figure 2A specific implementation of a text rendering method based on proxy confrontation is shown.
[0040] It should be noted that the method of the present invention is not limited to the method of Figure 2 The process sequence shown is limited to the following steps:
[0041] S1: Obtain glyph encoding, style parameters and target text image.
[0042] Among them, glyph encoding refers to the geometric feature representation generated by vector contour or stroke decomposition. Specifically, a convolutional autoencoder can be used to extract the glyph topology structure to control the stroke direction and connection relationship of the generated text. Style parameters refer to the vectorized representation of font style, color, and texture. Specifically, the high-level visual features of the target image can be extracted through the style transfer network to control the visual style of the generated text. The target text image is a real text image, which refers to a high-quality, standardized text image sample used to train and evaluate the generative model.
[0043] S2: Perform text rendering based on the glyph encoding and the style parameters through a rendering model to generate an initial text image.
[0044] Specifically, the rendering model adopts a U-Net architecture with modulated convolution, which consists of 8 downsampling layers and 8 upsampling layers. The rendering model receives glyph encoding and style parameters as input and outputs a rendered text image. In the U-Net architecture, the downsampling path captures global information, the upsampling path restores spatial details, and the skip connections preserve spatial information. Each convolutional layer is followed by batch normalization (BN) and ReLU activation function. Modulated convolution injects style information into the feature map through affine transformation, as follows:
[0045] y = Conv(x,W·s);
[0046] Here, x represents the input feature map, W represents the convolution kernel weight, s represents the style vector, and Conv represents the convolution operation. This modulation method enables the model to flexibly adjust the rendering effect according to different style parameters.
[0047] See also Figure 3 , Figure 3 A specific implementation of step S2 is shown, which is described in detail as follows:
[0048] S21: Concatenate the glyph code and the style parameters to generate concatenated parameters.
[0049] S22: Mapping the spliced parameters into affine transformation parameters through a fully connected layer, and adjusting the convolution kernel weights based on the affine transformation parameters.
[0050] S23: Performing downsampling processing based on the glyph encoding and the style parameters through an encoder to generate initial features.
[0051] S24: The decoder performs upsampling processing based on the initial features and the features of the encoder jump connection to generate the initial text image.
[0052] Specifically, during the parameter fusion stage, glyph encoding and style parameters are concatenated to form a joint representation, resolving the feature matching bias caused by the independent processing of the two types of parameters in traditional methods. The fully connected layer maps the concatenated parameters into affine transformation parameters and dynamically adjusts the convolution kernel weights, enabling the feature extraction process to adapt to the combined relationship between glyph structure and style features. The encoder employs a multi-level downsampling structure, preserving key geometric information such as stroke connection points while gradually extracting high-level semantic features. The decoder integrates features from each encoder layer through skip connections, simultaneously fusing shallow detail features with deep semantic features while restoring image resolution, effectively avoiding the stroke breakage phenomenon caused by information loss in traditional methods. This application addresses the problem of poor generated image quality caused by insufficient parameter fusion in traditional methods by effectively combining glyph encoding and style parameters to enhance the ability to preserve geometric structure during rendering. The dynamically adjusted convolution kernel accurately captures the associated features of glyphs and styles. The multi-level feature fusion mechanism ensures the structural integrity of stroke connections, and the skip connection feature transfer effectively avoids detail loss during upsampling, significantly improving the structural integrity and style consistency of the generated text while maintaining rendering efficiency.
[0053] Among them, splicing refers to connecting glyph encoding vectors and style parameter vectors of different dimensions along the channel dimension. Specifically, this can be achieved by using tensor splicing operations, allowing the subsequent network to simultaneously process glyph structure and style features. Mapping the fully connected layer to affine transformation parameters refers to converting the spliced high-dimensional vectors into affine transformation matrices through linear transformations. Specifically, this can be achieved using a multi-layer perceptron structure, with each affine transformation parameter corresponding to the spatial transformation coefficient of the convolution kernel. Adjusting the convolution kernel weights refers to dynamically modifying the spatial sampling position of the convolution kernel based on the affine transformation parameters. Specifically, this can be achieved through a deformable convolution mechanism, allowing the convolution operation to adaptively capture the associated features of glyphs and style. Encoder downsampling processing refers to extracting high-level semantic features by gradually reducing the resolution of feature maps. Specifically, this can be achieved by using strided convolution combined with residual connections, compressing feature dimensions while retaining key geometric structure information. Decoder upsampling processing refers to gradually restoring image resolution through transposed convolution or interpolation methods. Specifically, skip connections can be used to fuse features from different levels of the encoder with the corresponding layers of the decoder to avoid loss of details during the upsampling process.
[0054] S3: Calculate adversarial loss, perceptual consistency loss, and reconstruction loss in parallel based on the initial text image and the target text image.
[0055] Adversarial loss refers to the adversarial training loss calculated via a lightweight proxy network. Specifically, it employs depthwise separable convolution and attention mechanisms to construct a network structure, replacing the traditional discriminator to reduce computational complexity. Perceptual consistency loss refers to the loss term calculated by aligning the autoencoder's feature space. Specifically, it utilizes a pre-trained convolutional neural network to extract multi-scale features, which are used to constrain the structural integrity of the strokes in the generated image. Reconstruction loss refers to the loss calculated using the pixel-level L1 or L2 norm. Specifically, it employs a pixel-by-pixel comparison method to ensure that the generated result matches the low-level features of the target image.
[0056] See also Figure 4 , Figure 4 A specific implementation of step S3 is shown, which is described in detail as follows:
[0057] S31: Performing the adversarial loss based on the initial text image and the target text image through a proxy network.
[0058] S32: extracting image features from the initial text image and the target text image through an autoencoder, generating initial latent features and target latent features, and calculating the perceptual consistency loss based on the initial latent features and the latent features.
[0059] S33: Compare pixels of the initial text image with pixels of the target text image to calculate the reconstruction loss.
[0060] Specifically, the proxy network processes the initial and target text images separately through depthwise separable convolutional layers. The decomposed deep convolution extracts spatial features, and point-by-point convolution fuses channel information, reducing the number of parameters while maintaining feature extraction capabilities. The encoder portion of the autoencoder extracts latent features from the initial and target images, calculating the Manhattan distance to constrain high-level semantic consistency, while the gradient alignment loss enhances the sharpness of stroke edges. The reconstruction loss directly compares pixel-level differences to ensure accurate restoration of local details. The parallel computation mechanism of these three components preserves the adversarial training's ability to model complex style distributions while maintaining the integrity of the glyph topology through perceptual constraints.
[0061] See also Figure 5 , Figure 5 A specific implementation of step S31 is shown, which is described in detail as follows:
[0062] S311: Pre-training the lightweight network based on the discriminator through a bidirectional KL divergence distillation mechanism to generate the proxy network.
[0063] S312: Using the depthwise separable convolution layer of the proxy network to perform depthwise convolution and pointwise convolution on the initial text image and the target text image respectively, to generate an initial feature map and a target feature map.
[0064] S313: Performing global average pooling on the initial feature map and the target feature map through the SE attention module of the proxy network to generate an initial feature vector and a target feature vector.
[0065] S314: Outputting the adversarial loss based on the initial feature vector and the target feature vector through the fully connected layer of the proxy network.
[0066] Specifically, the present application approximates the behavior of complex discriminators through a lightweight network. The lightweight network consists of 4 layers of depth-separable convolution, each layer is followed by a SE (Squeeze-and-Excitation) attention module, and the total number of parameters is only 3.2M. The depth-separable convolution decomposes the standard convolution into depth-wise convolution and point-wise convolution, significantly reducing the amount of computation. The SE attention module captures inter-channel dependencies through global average pooling and enhances feature representation capabilities. The optimization goal of the proxy network is to minimize the KL divergence (Kullback-Leibler divergence, an asymmetric measure of the difference between two probability distributions) between the output distribution of the pre-trained discriminator. Its loss function is defined as:
[0067]
[0068] Where y represents the target text image; Represents rendering output; is the adversarial loss evaluation value of the proxy network; E denotes the mathematical expectation; KL denotes the Kullback-Leibler divergence, and || represents the KL divergence calculation, which measures the difference between two probability distributions P and Q. The first KL divergence term forces the proxy network to learn the discriminative features of real samples, while the second term enables it to capture the discriminative patterns of generated samples. This bidirectional KL divergence distillation mechanism ensures that the proxy network fully mimics the behavior of the discriminator.
[0069] The distillation training process of the proxy network is: the target text image y and the rendered output As input, features are extracted through depthwise separable convolution and SE modules; the extracted features are then globally averaged pooled to generate feature vectors. Finally, the fully connected layer outputs the adversarial loss evaluation value. Training objective: Minimize the bidirectional KL divergence with the discriminator output.
[0070] In the adversarial loss calculation process of the embodiment of the present application, the text structure perception ability of the original discriminator is first transferred to the lightweight proxy network using the bidirectional KL divergence distillation mechanism, so that the proxy network has the topological feature recognition ability comparable to that of the complex discriminator. Subsequently, the input image is subjected to feature extraction through a depth-wise separable convolution layer. The depth convolution stage uses a 3×3 convolution kernel to independently process the spatial features of each input channel. The point-by-point convolution stage uses a 1×1 convolution kernel to achieve cross-channel feature fusion. This combined operation significantly reduces the computational complexity while ensuring feature resolution. In the feature compression stage, the SE attention module performs global average pooling of the feature map in the channel dimension to generate a feature vector that represents the importance of each channel. The channel attention weights are learned through the fully connected layer and the feature response values are recalibrated to obtain higher activation strength for the features in the stroke intersection area. Finally, the fully connected layer maps the compressed feature vector to an adversarial loss value, and guides the rendering model through back propagation to generate an image that conforms to the text structure constraints.
[0071] The bidirectional KL divergence distillation mechanism uses bidirectional Kullback-Leibler divergence to simultaneously optimize the knowledge transfer between the teacher discriminator and the student proxy network. Specifically, a symmetric KL divergence loss function is used to align the probability distributions between the teacher and student networks. This mechanism effectively preserves the integrity of text structural features through bidirectional information flow. The SE attention module is a channel attention mechanism that includes a squeeze-excitation operation. Specifically, a global average pooling layer compresses spatial dimensions followed by two fully connected layers to generate channel attention weights. This module strengthens the weight distribution of stroke-connected regions by dynamically adjusting the response strength of channel features.
[0072] See also Figure 6, Figure 6 A specific implementation of step S32 is shown, which is described in detail as follows:
[0073] S321: extracting image features from the initial text image and the target text image respectively through the encoder to generate the initial latent features and the target latent features.
[0074] S322: Calculate the Manhattan distance between the initial latent feature and the target latent feature to generate a latent variable alignment loss.
[0075] S323: Calculate the gradient alignment constraints of the initial text image and the target text image to generate a gradient alignment loss.
[0076] S324: Calculate a weighted sum of the potential alignment loss and the gradient alignment loss to generate the perceptual consistency loss.
[0077] Specifically, the autoencoder focuses on extracting glyph structural features and includes three residual blocks. The autoencoder consists of an encoder and a decoder. The encoder outputs a 256-dimensional latent variable, and the decoder reconstructs the input image through channel-wise attention. To enhance structural awareness, this module uses only pure glyph data during training, ensuring that the feature space focuses on text structural characteristics.
[0078] Perceptual consistency loss consists of two parts: latent variable distance and gradient alignment constraint. Its mathematical definition is:
[0079]
[0080] Among them, E ψ represents the encoding function of the autoencoder, ||·||1 represents the L1 norm (i.e., the sum of absolute values), represents the square L2 norm (i.e. the square of the Euclidean distance), represents the image gradient calculated using the Sobel operator, with λ = 0.5 as the balancing factor. The first term constrains the distance between latent variables to ensure overall structural consistency; the second term enforces stroke edge alignment through spatial gradient differences.
[0081] In the text rendering process of the embodiment of the present application, the encoder performs dual-path feature extraction on the generated image and the real image to obtain high-dimensional latent features representing the glyph structure. The latent variable alignment loss measures the difference between the two latent features by the Manhattan distance, focusing on capturing the geometric deformation of the stroke joints and intersections. Due to the sensitivity of the L1 norm to sparse errors, it can effectively identify small structural deviations in local areas. The gradient alignment loss constrains the sharpness and coherence of the stroke edges by comparing the gradient map difference between the generated image and the target image, and adopts the directional gradient histogram statistical method to ensure that the gradient distribution of the generated image at the stroke turning point is consistent with the target image. The two loss functions are weighted fused to form a composite perceptual consistency loss, in which the latent variable alignment loss weight can be set to 0.7 and the gradient alignment loss weight can be set to 0.3, which enhances local details while maintaining overall structural consistency. Among them, Manhattan distance refers to the difference measure between two feature vectors calculated using the L1 norm. Compared with the Euclidean distance, it is more suitable for capturing sparse differences in the feature space. Specifically, it can be implemented by element-by-element absolute value summation. This calculation method is more sensitive to local deformations in the glyph structure. Gradient alignment constraint refers to the difference analysis of the image spatial gradient map. Specifically, the Sobel operator is used to calculate the image gradient and then the difference distribution in the edge area between the generated image and the target image is compared. This constraint can enhance the continuity of the stroke edges.
[0082] S4: Calculate the total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retrain the rendering model based on the total model loss to generate a target rendering model.
[0083] Specifically, the total model loss is calculated based on the adversarial loss, perceptual consistency loss and reconstruction loss, and the rendering model is retrained based on the total model loss to generate a target rendering model, including adjusting the weight coefficient based on the gradient norm and temperature parameter according to the dynamic temperature scaling mechanism; the adversarial loss, perceptual consistency loss and reconstruction loss are weightedly summed by the weight coefficient to generate the total model loss; the parameters of the rendering model are adjusted according to the total model loss, and the adjusted rendering model is iteratively trained to generate the target rendering model.
[0084] See also Figure 7 , Figure 7 A specific implementation of step S4 is shown, which is described in detail as follows:
[0085] S41: Adjust the weight coefficient based on the gradient norm and the temperature parameter according to the dynamic temperature scaling mechanism.
[0086] S42: Performing a weighted summation of the adversarial loss, the perceptual consistency loss, and the reconstruction loss using the weight coefficient to generate the total model loss.
[0087] Furthermore, the total loss of the model is calculated as:
[0088]
[0089] Among them, L surrogate is the total loss of the model, L perc is the perceptual consistency loss, L pixel is the reconstruction loss, is the adversarial loss, and α, β, and γ are weight coefficients.
[0090] The weight coefficient is dynamically adjusted by temperature scaling, and its formula is:
[0091]
[0092] Among them, T=0.1 is the temperature parameter, represents the gradient of the corresponding loss with respect to the rendering model parameter θ, ∑ represents the sum of all loss terms i, and exp represents the exponential function. This mechanism enables the system to focus on pixel-level reconstruction in the early stages of training, and gradually enhance the influence of adversarial and perceptual terms in the later stages, thus automating the multi-stage training strategy.
[0093] S43: Adjust the parameters of the rendering model according to the total loss of the model, and iteratively train the adjusted rendering model to generate the target rendering model.
[0094] In the embodiment of the present application, during the initial model training, the gradient norm analysis module monitors that the gradient magnitude of the adversarial loss is significantly higher than that of other loss terms. At this point, the dynamic temperature scaling mechanism automatically reduces its weight coefficient to prevent adversarial training from prematurely dominating the optimization direction. As training progresses, when the gradient norm of the perceptual consistency loss shows a continuous growth trend, the temperature parameter control module gradually increases its weight proportion, strengthening the constraint of the autoencoder feature space on the glyph structure. During the weight coefficient update phase, the weight of the reconstruction loss is dynamically adjusted based on the pixel error distribution of the current batch of samples, ensuring that local detail optimization and global style generation proceed simultaneously. Through multiple rounds of iterative training, the rendering model gradually establishes a nonlinear mapping relationship between glyph encoding and style parameters, achieving diverse text rendering effects while maintaining stroke connection accuracy. The dynamic temperature scaling mechanism refers to an algorithm that automatically adjusts the weights of different loss terms by monitoring the gradient change magnitude of each loss term during training. The gradient norm refers to the modulus of the derivative vector of the loss function with respect to the model parameters, which is used to quantify the degree of influence of different loss terms on parameter updates. The temperature parameter is a learnable variable that controls the magnitude of the loss term weight adjustment, which can be specifically optimized end-to-end using the backpropagation algorithm.
[0095] S5: Obtain the to-be-processed glyph code and style parameters, and perform text rendering based on the to-be-processed glyph code and style parameters through a target rendering model to generate a target text rendering image.
[0096] Specifically, in the inference stage or the actual application stage, the to-be-processed glyph encoding and style parameters are obtained, and text rendering is performed based on the to-be-processed glyph encoding and style parameters by a target rendering model to generate a target text rendering image.
[0097] This application is applied to different application scenarios in the field of financial technology, such as mobile banking and payment application scenarios, smart investment advisory and report generation scenarios, identity authentication and anti-fraud scenarios, etc. In mobile banking and payment application scenarios, high-fidelity electronic credentials (such as transaction receipts, electronic contracts, bills) can be generated in real time. This application can be applied to different application scenarios in the medical field, such as electronic health record (EHR) system scenarios, medical imaging report automatic annotation scenarios, and drug label and instruction manual generation scenarios. In the electronic health record (EHR) system, the scenario requirements are: rapid generation and editing of patient medical records, prescriptions and examination reports, which can accurately render medical-specific symbols (such as μg, ) and dosage units. Link text with images and laboratory data to improve diagnostic efficiency.
[0098] In an embodiment of the present application, a glyph encoding, style parameters and a target text image are obtained; text rendering is performed based on the glyph encoding and the style parameters by a rendering model to generate an initial text image; adversarial loss, perceptual consistency loss and reconstruction loss are calculated in parallel based on the initial text image and the target text image; the total model loss is calculated based on the adversarial loss, the perceptual consistency loss and the reconstruction loss, and the rendering model is retrained based on the total model loss to generate a target rendering model; the glyph encoding and style parameters to be processed are obtained, and text rendering is performed based on the glyph encoding and style parameters to be processed by a target rendering model to generate a target text rendering image. The embodiment of the present invention calculates the adversarial loss, perceptual consistency loss and reconstruction loss of the text image after rendering the text, and combines multiple losses to train the rendering model, which is conducive to improving the instructions of text rendering, and the present application does not need to rely on the discriminator network to calculate the loss calculation, which reduces the complexity of the model and is conducive to improving the efficiency of text rendering.
[0099] Please refer to Figure 8 , as a response to the above Figure 2 The present application provides an embodiment of a text rendering device based on proxy confrontation, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0100] like Figure 8 As shown, the text rendering device based on proxy confrontation in this embodiment includes: a data acquisition module 61, a text rendering module 62, a loss calculation module 63, a model training module 64 and an image generation module 65, wherein:
[0101] A data acquisition module 61 is used to acquire font encoding, style parameters and target text image;
[0102] A text rendering module 62 is configured to perform text rendering based on the glyph encoding and the style parameters using a rendering model to generate an initial text image;
[0103] A loss calculation module 63 is configured to concurrently calculate an adversarial loss, a perceptual consistency loss, and a reconstruction loss based on the initial text image and the target text image;
[0104] a model training module 64, configured to calculate a total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retrain the rendering model based on the total model loss to generate a target rendering model;
[0105] The image generation module 65 is configured to obtain the to-be-processed glyph code and style parameters, and perform text rendering based on the to-be-processed glyph code and style parameters using a target rendering model to generate a target text rendering image.
[0106] Furthermore, the text rendering module 62 includes:
[0107] A data splicing unit, configured to splice the glyph code and the style parameters to generate spliced parameters;
[0108] a data mapping unit, configured to map the concatenated parameters into affine transformation parameters through a fully connected layer, and adjust the convolution kernel weights based on the affine transformation parameters;
[0109] an initial feature generating unit, configured to generate initial features by performing downsampling processing based on the glyph code and the style parameters through an encoder;
[0110] The initial text image generating unit is configured to generate the initial text image by performing upsampling processing on the basis of the initial features and the features of the jump connection of the encoder through a decoder.
[0111] Furthermore, the loss calculation module 63 includes:
[0112] an adversarial loss calculation unit, configured to perform the adversarial loss calculation based on the initial text image and the target text image through a proxy network;
[0113] a perceptual consistency loss calculation unit, configured to extract image features from the initial text image and the target text image through an autoencoder, generate initial latent features and target latent features, and calculate the perceptual consistency loss based on the initial latent features and the latent features;
[0114] A reconstruction loss calculation unit is used to compare pixels of the initial text image with pixels of the target text image to calculate the reconstruction loss.
[0115] Furthermore, the adversarial loss calculation unit includes:
[0116] A pre-training unit, configured to pre-train the lightweight network based on the discriminator through a bidirectional KL divergence distillation mechanism to generate the proxy network;
[0117] A convolution unit, configured to perform depthwise convolution and pointwise convolution on the initial text image and the target text image respectively using the depthwise separable convolution layer of the proxy network to generate an initial feature map and a target feature map;
[0118] A pooling unit, configured to perform global average pooling on the initial feature map and the target feature map through the SE attention module of the proxy network to generate an initial feature vector and a target feature vector;
[0119] An adversarial loss generating unit is configured to output the adversarial loss based on the initial feature vector and the target feature vector through a fully connected layer of the proxy network.
[0120] Furthermore, the perceptual consistency loss calculation unit includes:
[0121] a latent feature generating subunit, configured to extract image features from the initial text image and the target text image respectively through the encoder to generate the initial latent features and the target latent features;
[0122] a latent variable alignment loss calculation subunit, configured to calculate the Manhattan distance between the initial latent feature and the target latent feature to generate a latent variable alignment loss;
[0123] A gradient alignment loss calculation subunit, configured to calculate the gradient alignment constraint of the initial text image and the target text image to generate a gradient alignment loss;
[0124] The perceptual consistency loss generating subunit is configured to calculate a weighted sum of the potential alignment loss and the gradient alignment loss to generate the perceptual consistency loss.
[0125] Furthermore, the model training module 64 includes:
[0126] a weight coefficient adjustment unit for adjusting the weight coefficient based on the gradient norm and the temperature parameter according to a dynamic temperature scaling mechanism;
[0127] a model total loss calculation unit, configured to perform a weighted summation of the adversarial loss, the perceptual consistency loss, and the reconstruction loss using the weight coefficient to generate the model total loss;
[0128] A target rendering model generating unit is used to adjust the parameters of the rendering model according to the total loss of the model, and iteratively train the adjusted rendering model to generate the target rendering model.
[0129] Furthermore, the calculation formula of the total loss of the model is:
[0130]
[0131] Among them, L surrogate is the total loss of the model, L perc is the perceptual consistency loss, L pixel is the reconstruction loss, is the adversarial loss, and α, β, and γ are weight coefficients.
[0132] To solve the above technical problems, the present application also provides a computer device. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.
[0133] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are interconnected through a system bus. It should be noted that Figure 9 Only a computer device 7 having three components, memory 71, processor 72, and network interface 73, is shown. However, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. It should be understood by those skilled in the art that a computer device herein is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0134] Computer devices can be desktop computers, laptops, PDAs, cloud servers, etc. Computer devices can interact with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0135] The memory 71 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 7. Of course, the memory 71 may also include both the internal storage unit of the computer device 7 and its external storage devices. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed on the computer device 7, such as the program code of the proxy-based text rendering method. In addition, the memory 71 may also be used to temporarily store various types of data that have been output or are about to be output.
[0136] In some embodiments, the processor 72 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to execute program code stored in the memory 71 or process data, such as executing the program code of the above-mentioned proxy-based text rendering method to implement various embodiments of the proxy-based text rendering method.
[0137] The network interface 73 may include a wireless network interface or a wired network interface. The network interface 73 is generally used to establish a communication connection between the computer device 7 and other electronic devices.
[0138] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program. The computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned proxy confrontation-based text rendering method.
[0139] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.
[0140] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A text rendering method based on proxy confrontation, characterized in that: include: Get glyph encoding, style parameters and target text image; Performing text rendering based on the glyph encoding and the style parameters by a rendering model to generate an initial text image; Parallel calculation of adversarial loss, perceptual consistency loss, and reconstruction loss based on the initial text image and the target text image; Calculating a total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retraining the rendering model based on the total model loss to generate a target rendering model; Obtaining a to-be-processed glyph code and style parameters, and performing text rendering based on the to-be-processed glyph code and style parameters through a target rendering model to generate a target text rendering image.
2. The text rendering method based on proxy confrontation according to claim 1, characterized in that: The step of performing text rendering based on the glyph encoding and the style parameters by using a rendering model to generate an initial text image includes: Splicing the glyph code and the style parameter to generate a spliced parameter; Mapping the concatenated parameters into affine transformation parameters through a fully connected layer, and adjusting the convolution kernel weights based on the affine transformation parameters; Performing downsampling processing based on the glyph code and the style parameters by an encoder to generate initial features; The decoder performs upsampling processing based on the initial features and the features of the encoder jump connection to generate the initial text image.
3. The text rendering method based on proxy confrontation according to claim 1, characterized in that: The parallel calculation of the adversarial loss, the perceptual consistency loss, and the reconstruction loss based on the initial text image and the target text image includes: Performing the adversarial loss based on the initial text image and the target text image through a proxy network; extracting image features from the initial text image and the target text image through an autoencoder to generate initial latent features and target latent features, and calculating the perceptual consistency loss based on the initial latent features and the latent features; Pixels of the initial text image are compared with pixels of the target text image to calculate the reconstruction loss.
4. The text rendering method based on proxy confrontation according to claim 3 is characterized in that: The performing the adversarial loss based on the initial text image and the target text image through the proxy network includes: Pre-training the lightweight network based on the discriminator through a bidirectional KL divergence distillation mechanism to generate the proxy network; Using the depthwise separable convolution layer of the proxy network to perform depthwise convolution and pointwise convolution on the initial text image and the target text image respectively, to generate an initial feature map and a target feature map; Performing global average pooling on the initial feature map and the target feature map through the SE attention module of the proxy network to generate an initial feature vector and a target feature vector; The adversarial loss is output based on the initial feature vector and the target feature vector through a fully connected layer of the proxy network.
5. The text rendering method based on proxy confrontation according to claim 3, characterized in that: The extracting image features from the initial text image and the target text image by the autoencoder, generating initial latent features and target latent features, and calculating the perceptual consistency loss based on the initial latent features and the latent features, includes: Extracting image features from the initial text image and the target text image respectively by the encoder to generate the initial latent features and the target latent features; Calculating the Manhattan distance between the initial latent feature and the target latent feature to generate a latent variable alignment loss; Calculating gradient alignment constraints of the initial text image and the target text image to generate a gradient alignment loss; A weighted sum of the potential alignment loss and the gradient alignment loss is calculated to generate the perceptual consistency loss.
6. The text rendering method based on proxy confrontation according to any one of claims 1 to 5, characterized in that: The calculating the total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retraining the rendering model based on the total model loss to generate a target rendering model includes: Adjust the weight coefficient based on the gradient norm and temperature parameter according to the dynamic temperature scaling mechanism; Performing a weighted summation of the adversarial loss, the perceptual consistency loss, and the reconstruction loss using the weight coefficient to generate the total model loss; The parameters of the rendering model are adjusted according to the total loss of the model, and the adjusted rendering model is iteratively trained to generate the target rendering model.
7. The text rendering method based on proxy confrontation according to claim 6, characterized in that: The total loss of the model is calculated as: Among them, L surrogate is the total loss of the model, L perc is the perceptual consistency loss, L pixel is the reconstruction loss, is the adversarial loss, and α, β, and γ are weight coefficients.
8. A text rendering device based on proxy confrontation, characterized in that: include: A data acquisition module, used to obtain glyph encoding, style parameters and target text image; A text rendering module, configured to perform text rendering based on the glyph encoding and the style parameters through a rendering model to generate an initial text image; a loss calculation module, configured to concurrently calculate an adversarial loss, a perceptual consistency loss, and a reconstruction loss based on the initial text image and the target text image; a model training module, configured to calculate a total model loss based on the adversarial loss, the perceptual consistency loss, and the reconstruction loss, and retrain the rendering model based on the total model loss to generate a target rendering model; The image generation module is used to obtain the to-be-processed glyph code and style parameters, and perform text rendering based on the to-be-processed glyph code and style parameters through a target rendering model to generate a target text rendering image.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the text rendering method based on proxy confrontation as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text rendering method based on proxy confrontation according to any one of claims 1 to 7 is implemented.