Text rendering method and device based on multi-task auxiliary features, equipment and medium

Through the text rendering method of dual-modal embedding and dynamic loss balance mechanism, the problems of multi-task coordination and cross-language support are solved, and efficient and accurate text rendering effect is achieved.

CN120580315APending Publication Date: 2025-09-02PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510686851.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing text rendering technologies have bottlenecks in multi-task coordination and cross-language support, resulting in instability in training, lack of structural information and waste of computing resources, making it difficult to generate high-quality text rendering results.

Method used

Joint representations are generated through dual-modal embedding processing, potential spatial representations are extracted using a shared encoder, auxiliary tasks are performed in parallel, and weights are adjusted through dynamic loss balance mechanism to generate the target text rendering model.

Benefits of technology

It improves the accuracy and efficiency of text rendering, solves task conflicts in multi-task learning, enhances the consistency of glyph structure and cross-language rendering capabilities, and reduces deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580315A_ABST
    Figure CN120580315A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and financial science and technology, and discloses a text rendering method, device and equipment based on multi-task auxiliary features and a medium, which are applied to a personalized customer report and marketing material generation scene in the field of financial science and technology or can be applied to a medical image report automatic labeling scene in the medical field. The method comprises the following steps: acquiring a Unicode coding sequence and font metadata of a text, and performing bimodal embedding processing on the Unicode coding sequence and font metadata of the text to generate joint representation; performing feature extraction on the joint representation to generate potential spatial representation; decoding the potential spatial representation to generate an initial text rendering image; executing a plurality of auxiliary tasks in parallel based on the potential spatial representation, and adjusting the weights of the auxiliary tasks to train the text rendering model to generate a target text rendering model; and generating a target text rendering image based on the to-be-rendered data through the target text rendering model. According to the method and the device, the text rendering precision and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer vision and financial technology, and in particular to a text rendering method, apparatus, device, and medium based on multi-task auxiliary features. Background Art

[0002] Text rendering technology is an ongoing and active research direction, whose goal is to convert discrete text symbols into high-quality visual representations. Text rendering technology can be applied in the fields of financial technology and medicine. For example, it can be used in the generation of personalized customer reports and marketing materials in the financial technology field, or in the automatic annotation of medical imaging reports in the medical field. Current text rendering technologies are mainly divided into two categories: one is rule-driven glyph synthesis or template-based matching technology. Such methods are simple to implement and efficient to execute, but have difficulty in handling complex font styles or adapting to dynamic scene requirements; the other is generative methods based on deep learning (such as generative adversarial networks), which achieve data-driven text rendering through end-to-end training, and can generate more flexible and diverse effects, but still have obvious defects in structural details.

[0003] Early research on deep learning-based text rendering focused primarily on single-task models. For example, the Glyph-byT5 model proposed by Liu et al. focused on visual text encoding but lacked an explicit understanding of glyph structure. GAN-based text-to-text generation networks can generate stylized text but perform poorly at preserving character topology. With technological advancements, researchers have gradually realized that text rendering is not just a pixel-level generation problem but also involves complex structural understanding. To address this issue, multi-task learning (MTL) has been introduced to text rendering. It improves offline character recognition by adding the auxiliary task of stroke order prediction and combines text recognition with mathematical formula parsing to enhance the rendering quality of specialized text. While multi-task learning has shown promise in text rendering, existing techniques still suffer from three major limitations: First, the problem of weight balance between tasks. Static weights struggle to adapt to the dynamic gradient contributions of different tasks, leading to unstable training. Second, there is a lack of structural information. Pure pixel-level optimization (such as GANs or diffusion models) struggles to capture the geometric and linguistic characteristics of text and can easily produce broken strokes or distorted glyphs. Third, computational resources are wasted. While the multi-attention-assisted learning framework offers superior performance, it still requires retaining all auxiliary branches during the inference phase, significantly increasing deployment costs. Consequently, existing text rendering technologies have limited capabilities, resulting in lower accuracy and efficiency. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose a text rendering method, apparatus, device and medium based on multi-task auxiliary features to improve the accuracy and efficiency of text rendering.

[0005] In order to solve the above technical problems, an embodiment of the present application provides a text rendering method based on a multi-task auxiliary feature, comprising:

[0006] Obtaining a Unicode code sequence and font metadata of a text, and performing bimodal embedding processing on the Unicode code sequence of the text and the font metadata to generate a joint representation of the text and the font;

[0007] Extracting features from the joint representation using a shared encoder to generate a latent space representation;

[0008] Inputting the latent space representation into a main decoder for decoding to generate an initial text rendering image;

[0009] executing multiple auxiliary tasks in parallel based on the latent space representation, and adjusting the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train a text rendering model and generate a target text rendering model;

[0010] A Unicode code sequence to be rendered and metadata of a font to be rendered are obtained, and text rendering is performed based on the Unicode code sequence to be rendered and the metadata of the font to be rendered by the target text rendering model to generate a target text rendering image.

[0011] In order to solve the above technical problems, an embodiment of the present application provides a text rendering device based on a multi-task auxiliary feature, comprising:

[0012] A text data acquisition module is used to acquire the Unicode encoding sequence and font metadata of the text, and perform bimodal embedding processing on the Unicode encoding sequence of the text and the font metadata to generate a joint representation of the text and the font;

[0013] a feature extraction module, configured to extract features from the joint representation using a shared encoder to generate a latent space representation;

[0014] A feature decoding module, configured to input the latent space representation into a main decoder for decoding to generate an initial text rendering image;

[0015] an auxiliary task parallel module, configured to execute multiple auxiliary tasks in parallel based on the latent space representation, and adjust the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train the text rendering model and generate a target text rendering model;

[0016] The text rendering module is used to obtain a Unicode code sequence to be rendered and metadata of a font to be rendered, and perform text rendering based on the Unicode code sequence to be rendered and the metadata of the font to be rendered by the target text rendering model to generate a target text rendering image.

[0017] To solve the above technical problems, a technical solution adopted by the present invention is: to provide a computer device, including one or more processors; a memory for storing one or more programs, so that the one or more processors can implement any one of the above-mentioned text rendering methods based on multi-tasking auxiliary features.

[0018] To solve the above technical problems, the present invention adopts a technical solution: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-mentioned text rendering methods based on multi-task auxiliary features.

[0019] The embodiment of the present invention provides a text rendering method, apparatus, device and medium based on multi-task auxiliary features. The method includes: obtaining the Unicode encoding sequence and font metadata of the text, and performing bimodal embedding processing on the Unicode encoding sequence of the text and the font metadata to generate a joint representation of the text and the font; extracting features from the joint representation through a shared encoder to generate a latent space representation; inputting the latent space representation into a main decoder for decoding to generate an initial text rendering image; executing multiple auxiliary tasks in parallel based on the latent space representation, and adjusting the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train a text rendering model to generate a target text rendering model; obtaining the Unicode encoding sequence to be rendered and the font metadata to be rendered, and performing text rendering based on the Unicode encoding sequence to be rendered and the font metadata to be rendered through the target text rendering model to generate a target text rendering image. The embodiment of the present invention generates a joint representation through bimodal embedding processing, combines a shared encoder with a dynamic loss balancing mechanism, automatically coordinates the weight relationship between multiple tasks in the training phase, and enhances structural understanding capabilities through parallel auxiliary tasks. Only the backbone network is retained in the reasoning phase. It has the advantages of resolving task conflicts in multi-task learning, improving glyph structure consistency, and enhancing cross-language rendering capabilities, which is conducive to improving the accuracy and efficiency of text rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 2 is a schematic diagram of an application environment of a text rendering method based on multi-task auxiliary features according to an embodiment of the present invention;

[0022] Figure 2 This is a flowchart of the implementation process of the text rendering method based on the multi-task auxiliary feature provided in an embodiment of the present application;

[0023] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S1;

[0024] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S2;

[0025] Figure 5 yes Figure 2 A schematic flow chart of a specific implementation of step S3;

[0026] Figure 6 yes Figure 2 A schematic flow chart of a specific implementation of step S4;

[0027] Figure 7 yes Figure 6 A schematic flow chart of a specific implementation of step S41;

[0028] Figure 8 Schematic diagram of a text rendering device based on multi-task assist features provided in an embodiment of the present application;

[0029] Figure 9 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0032] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0033] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0034] It should be noted that the text rendering method based on the multi-task assist feature provided in the embodiment of the present application is generally executed by a server. Accordingly, the text rendering device based on the multi-task assist feature is generally configured in the server.

[0035] The text rendering method based on multi-task auxiliary features provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the client communicates with the server through a network. The server can receive the Unicode encoding sequence to be rendered and the font metadata to be rendered from the client; and generate a target text rendering image according to the Unicode encoding sequence to be rendered and the font metadata to be rendered. The server in the present invention sends the target text rendering image to the client. The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0036] The text rendering method based on multi-task auxiliary features provided in the embodiment of the present application can be applied to the personalized customer report and marketing material generation scenarios in the financial technology field, or to the automatic annotation scenarios of medical imaging reports in the medical field.

[0037] In the existing technology, text rendering technology has long faced the bottlenecks of multi-task coordination difficulties and insufficient cross-language support. Traditional methods rely on fixed templates, resulting in a single style. Although deep learning models can generate diverse texts, it is difficult to maintain the integrity of character structure. When dealing with complex writing systems, existing systems often ignore stroke order and typesetting rules, resulting in broken Chinese character strokes or incorrect Arabic ligatures. An international news organization needs to generate title images for multilingual electronic publications in different fonts. The Western fonts generated by existing rendering tools are acceptable, but Chinese titles frequently have stroke adhesion, and Arabic titles cannot correctly display ligature effects, which seriously affects the professionalism and aesthetics of the publication.

[0038] To address the aforementioned issues, the R&D team discovered that static weight distribution in multi-task learning is a key factor contributing to training instability. By analyzing the impact curves of different auxiliary tasks on the main task, they found that glyph prediction and language feature extraction contribute differently at different training stages. This led to the core concept of dynamically adjusting task weights. Taking into account computational redundancy in the inference stage, they proposed a two-stage architecture that integrates auxiliary tasks during training and retains only the main network during inference. To address the challenge of cross-lingual feature fusion, a bimodal embedding mechanism was designed to jointly model text encoding and font parameters, enabling the model to automatically adapt to the structural characteristics of different writing systems. Therefore, this application proposes obtaining the Unicode encoding sequence and font metadata of the text and performing bimodal embedding processing to generate a joint representation. A shared encoder is used to extract the latent space representation, while the main decoder generates the initial rendered image. Auxiliary tasks are then executed in parallel based on the latent space, and the weights are adjusted using a dynamic loss balancing mechanism to train the model. Ultimately, the trained model is used to achieve target text rendering. This application effectively alleviates the gradient conflict problem in multi-task learning, enabling the model to accurately capture character structural features while maintaining high-resolution rendering capabilities. The dynamic weight adjustment mechanism ensures the appropriate contribution of different auxiliary tasks at each training stage, significantly improving training stability compared to traditional static weighting methods. The shared encoder architecture, combined with a branch-pruning strategy during inference, significantly reduces deployment costs while maintaining rendering quality. The bimodal embedding design enables the model to adaptively handle the unique characteristics of different writing systems, successfully achieving high-quality rendering for complex scripts such as Chinese and Arabic.

[0039] See also Figure 2 , Figure 2 A specific implementation of a text rendering method based on multi-task assist features is shown.

[0040] It should be noted that the method of the present invention is not limited to the method of Figure 2 The process sequence shown is limited to the following steps:

[0041] S1: Obtain a Unicode encoding sequence and font metadata of a text, and perform bimodal embedding processing on the Unicode encoding sequence of the text and the font metadata to generate a joint representation of the text and the font.

[0042] Specifically, the Unicode encoding sequence and font metadata of the text are obtained, the Unicode encoding sequence is mapped to a text embedding vector through a trainable codebook, the font metadata is mapped to a font feature vector through a multi-layer perceptron, and the text embedding vector and the font feature vector are spliced ​​into a joint representation.

[0043] See also Figure 3 , Figure 3 A specific implementation of step S1 is shown, which is described in detail as follows:

[0044] S11: Obtain the Unicode encoding sequence of the text and the font metadata.

[0045] S12: Map the Unicode encoding sequence into a text embedding vector through a trainable codebook.

[0046] S13: Mapping the font metadata into the font feature vector through a multi-layer perceptron.

[0047] S14: Concatenate the text embedding vector and the font feature vector into the joint representation.

[0048] Specifically, text embedding uses a trainable codebook mapping to convert each Unicode code point into a 768-dimensional vector; font metadata is projected to the same dimension using a multilayer perceptron (MLP). These two embeddings are concatenated (representing a concatenation operation) to generate a joint feature. The formula for the bimodal embedding process is:

[0049]

[0050] Where X represents the Unicode encoding sequence of the input text, F represents the font metadata, and Z represents the output latent space representation.

[0051] Specifically, the Unicode encoding sequence is mapped by a trainable codebook to form a high-dimensional semantic vector. For example, each character can be converted into a 256-dimensional text embedding vector. The font metadata is nonlinearly transformed through a multi-layer perceptron. For example, the font attributes containing 10 dimensions are expanded to a 128-dimensional style feature vector. After the feature vectors of the two modalities are spliced ​​in the channel dimension, a 384-dimensional joint representation matrix is ​​formed, which contains both the semantic information of the characters and the font style features. During the training process, the parameters of the trainable codebook and the weights of the multi-layer perceptron are simultaneously optimized through back propagation, so that the text embedding vector can be adaptively adjusted to match the font feature distribution. The embodiment of the present application effectively solves the problem that text symbols and font style features are difficult to integrate. Through a trainable cross-modal embedding mechanism, a deep coupling of text semantics and font attributes is achieved, so that the subsequent rendering process can simultaneously consider the meaning of characters and visual style elements. This joint representation not only retains the detailed features of the original data, but also establishes a mapping relationship between cross-modal features, providing basic feature support for generating rendering results with accurate structure and unified style. For example, when rendering Arabic ligatures, this solution can simultaneously maintain the semantic coherence of the characters and the unity of the calligraphy style, overcoming the style discontinuity problem commonly found in traditional methods.

[0052] Among them, the trainable codebook refers to a dynamically updated vector lookup table, which can be implemented by using an embedding matrix combined with an index lookup operation to convert discrete Unicode codes into differentiable continuous vector representations. The multi-layer perceptron refers to a fully connected neural network with multiple hidden layers, which can be implemented by using a three-layer network structure with a ReLU activation function to capture the nonlinear relationship between attributes in font metadata such as weight and italic angle. The joint representation concatenation operation refers to the vector connection along the feature dimension, which can be implemented by superimposing the channel dimension, so that the text semantic information and the font style features form a spatially aligned feature combination.

[0053] S2: Extract features from the joint representation through a shared encoder to generate a latent space representation.

[0054] Specifically, in order to enhance the ability to model character order, a relative position bias is injected into the Transformer layer of the shared encoder, and the joint representation is extracted through the shared encoder to generate a latent space representation.

[0055] Among them, the shared encoder refers to a multi-layer neural network structure used to extract joint features of text and fonts. It can be specifically implemented by stacking self-attention layers based on the Transformer architecture. Its function is to establish cross-modal feature associations.

[0056] See also Figure 4 , Figure 4 A specific implementation of step S2 is shown, which is described in detail as follows:

[0057] S21: Input the joint representation into the shared encoder.

[0058] S22: Perform attention calculation based on the joint representation through multiple attention mechanisms to generate an attention score.

[0059] S23: Injecting a relative position bias matrix into each layer of the shared encoder to add the attention score to the relative position bias matrix to obtain an addition result.

[0060] S24: Perform activation processing and normalization processing on the addition result to generate the latent space representation.

[0061] Specifically, the calculation process of the shared encoder at layer l is:

[0062] Z l =LayerNorm(A l +MLP(A l ));

[0063] A l =MultiHeadAttn(Z l-1 )+P;

[0064] Among them, Z l Represents the output features of the lth layer, A l Represents the result of self-attention calculation, P represents the position bias matrix, which adopts a learnable relative position encoding scheme, LayerNorm represents the layer normalization operation, and normalizes the features; MLP represents the multi-layer perceptron, which is used for feature transformation; MultiHeadAttn represents the multi-head attention mechanism, which is the core computing unit of Transformer and is used to capture the dependencies within the sequence.

[0065] In an embodiment of the present application, after the joint representation of text and font is input into the shared encoder, the association strength between each character is first calculated through a multi-head attention mechanism. Each attention head analyzes the stroke connection relationship and spatial distribution characteristics between characters in different dimensions. For example, the first attention head can focus on the continuity of adjacent strokes, and the second attention head focuses on the overall geometric symmetry of the character. In the encoding process of each layer, the pre-calculated relative position bias matrix is ​​superimposed on the original attention score, so that adjacent characters obtain higher association weights, and the interaction weights of distant characters decay exponentially. This processing method effectively encodes the topological constraint relationship between characters, such as the temporal sequence of strokes in Chinese characters and the continuous writing characteristics of Arabic characters. The superimposed features are introduced into a nonlinear transformation through the GELU activation function, and then the feature distribution offset is eliminated through layer normalization, and finally the potential space representation containing both semantic information and structural features is output. The present application can effectively model the spatial topological relationship and stroke continuity features between characters, and maintain the integrity and coherence of the character structure when generating text rendering images. For character rendering scenarios in complex writing systems, encoding relative position information ensures a natural transition at stroke connections, avoiding common defects such as broken strokes or stroke overlap in traditional methods, and significantly improving the visual quality of multilingual text rendering.

[0066] Among them, multiple attention mechanisms refer to attention calculation paths executed in parallel, which can be implemented specifically by a multi-head attention mechanism. Different attention head dimensions are set for each path to capture feature interaction patterns of different granularities. Among them, the relative position bias matrix refers to the weight parameter that characterizes the relative position relationship between characters. It can be implemented specifically by a learnable two-dimensional matrix, whose rows and columns correspond to the absolute position differences of the characters in the sequence. Among them, activation processing refers to the operation of performing nonlinear transformation on the output of the neural network, which can be implemented specifically by the GELU activation function, which is used to enhance the model's ability to fit complex features. Among them, normalization processing refers to the operation of standardizing the feature distribution, which can be implemented specifically by the layer normalization method, which is used to stabilize the training process and accelerate model convergence.

[0067] S3: Input the latent space representation into the main decoder for decoding to generate an initial text rendering image.

[0068] Specifically, the main decoder adopts a transposed convolutional network with skip connections to upsample the latent representation Z to the initial text rendering image Y∈R H×W×C , where H represents the image height, W represents the image width, C represents the number of channels, and R represents the real space. The decoding process can be expressed as: Y = Decoder(Z). The decoder consists of multiple layers of transposed convolution (also called deconvolution), each layer gradually increasing the spatial resolution of the feature map while reducing the number of channels.

[0069] See also Figure 5 , Figure 5 A specific implementation of step S3 is shown, which is described in detail as follows:

[0070] S31: Input the latent space representation into the main decoder, use a multi-layer transposed convolutional network to upsample the latent space representation layer by layer, and perform jump connection fusion feature between each layer of the transposed convolutional network and the shared encoder to output the upsampled feature.

[0071] S32: Mapping the upsampled features through a Tanh activation function to generate the initial text rendering image.

[0072] Specifically, after the latent space representation is input into the main decoder, the multi-layer transposed convolutional network gradually expands the spatial resolution of the feature map through hierarchical upsampling. While increasing the size of the feature map, the transposed convolution operation of each layer obtains the local features of the corresponding layer of the encoder through jump connections, realizing the cross-layer fusion of the stroke details captured by the encoder and the global structure reconstructed by the decoder. This process establishes a multi-scale information transmission path in the feature dimension, so that the stroke edge information lost during the upsampling process can be supplemented by jump connections. Finally, the normalized feature values ​​are mapped to the standard image pixel range through the Tanh function to generate a text rendering image that conforms to the laws of visual perception. The embodiment of the application solves the problem of glyph structure distortion caused by insufficient multi-scale feature fusion in the text rendering process, and improves the structural consistency and visual quality of the generated image. Specifically, it maintains continuous and smooth stroke connections in the character boundary area, avoiding the common broken strokes or burrs in traditional methods. At the same time, the pixel distribution of the generated image is more in line with the statistical characteristics of the real text image, significantly improving the visual fidelity of the rendering result.

[0073] Among them, the multi-layer transposed convolutional network refers to a feature decoding architecture formed by stacking multiple transposed convolutional layers. Specifically, it can be implemented using transposed convolutional layers with a convolution kernel size of 3×3 and a stride of 2 in each layer. It restores image details by gradually expanding the size of the feature map. Skip connection refers to the operation of channel-wise splicing of the feature maps output by each layer of the encoder with the output of the transposed convolutional layer at the corresponding level. Specifically, it can be implemented by splicing the feature map channel dimension and then connecting it to a convolution layer with a 1×1 convolution kernel. It is used to fuse the local texture features of the encoder with the high-level semantic features of the decoder. The Tanh activation function refers to the hyperbolic tangent function, which is specifically used at the end of the decoder to map feature values ​​to the range [-1, 1], optimizing image generation quality by limiting the pixel value range.

[0074] S4: Execute multiple auxiliary tasks in parallel based on the latent space representation, and adjust the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train the text rendering model and generate a target text rendering model.

[0075] Specifically, this application constructs three parallel auxiliary task heads, using the output features (latent space representation) of the shared encoder as input to the auxiliary task heads. Among them, the auxiliary tasks include glyph stroke prediction, character boundary detection, and language feature extraction. The weights of the auxiliary tasks are adjusted through a dynamic loss balancing mechanism, thereby increasing the weights of auxiliary tasks with similar features to the main task and reducing the influence of irrelevant tasks, thereby achieving adaptive multi-task learning.

[0076] See also Figure 6 , Figure 6 A specific implementation of step S4 is shown, which is described in detail as follows:

[0077] S41: Based on the latent space representation, glyph stroke order prediction, character boundary detection and language feature extraction are performed in parallel to generate stroke order features, character boundary heat map and language features, and an auxiliary task loss value is generated based on the stroke order features, the character boundary heat map and the language features.

[0078] S42: Calculating a weight coefficient based on the initial text rendering image, the stroke order feature, the character boundary heat map, and the language feature through the dynamic loss balancing mechanism.

[0079] The calculation formula of the weight coefficient is as follows:

[0080]

[0081] Among them, α i The weight coefficient of the i-th auxiliary task, τ is the temperature coefficient (the default value is 0.1), which controls the smoothness of the weight distribution, Z is the latent space feature, Z i 、Z j are the features of the i-th and j-th auxiliary tasks; Sim(·) represents the cosine similarity function, which is used to calculate the similarity between two feature vectors; exp(·) represents the exponential function with the natural constant e as the base.

[0082] S43: Calculate the total model loss based on the auxiliary task loss value and the weight coefficient.

[0083] The total loss of the model is calculated as:

[0084]

[0085] Among them, L total is the total loss of the model, L adv To combat the loss, the realism of the generated image is evaluated through the discriminator network; L pix Represents pixel-level L1 loss, directly comparing pixel differences; L i represents the loss of the i-th auxiliary task; λadv and λ pix is the preset weight, and α i is the dynamically calculated weight coefficient.

[0086] S44: Adjusting parameters of the text rendering model based on the total model loss, and generating the target text rendering model by iteratively training the adjusted text rendering model.

[0087] Specifically, during model training, the latent space representation is simultaneously fed into three independent auxiliary branches. The stroke order prediction branch uses a bidirectional LSTM network to generate predicted stroke sequences for each character. This is then compared with the true stroke order labels using a masked cross-entropy loss, forcing the model to learn the structural writing rules of the characters. The boundary detection branch utilizes a U-Net architecture to generate character outline heatmaps with the same resolution as the main task output, using a Dice loss to constrain the sharpness of character edges. The language feature branch uses a Transformer encoder to extract semantic vectors, which are then compared with the features output by a pre-trained language model using a cross-entropy loss to ensure semantic consistency of the text. The dynamic weight calculation module calculates the cosine similarity between the main task features and the features of each auxiliary task, and then combines this with a temperature coefficient to generate normalized weights, giving higher weights to auxiliary tasks with a high correlation to the main task. The total loss is composed of the weighted losses of each auxiliary task and the pixel-wise L1 loss of the main task. All network parameters are optimized simultaneously through backpropagation. After multiple rounds of iterative training, the model is able to retain only the shared encoder and main decoder during inference, achieving highly efficient text rendering. The application solves the problem of training instability caused by fixed weight distribution in multi-task learning, and enables the auxiliary tasks to be collaboratively optimized during the training process through a dynamic adjustment mechanism. The joint training of structural constraint tasks and generation tasks effectively prevents common defects such as broken strokes and distorted glyphs. At the same time, language feature extraction enhances the semantic rationality of text typesetting. The design of removing auxiliary branches in the inference stage reduces the computational load by about 40%, making the model more suitable for mobile deployment. In complex text rendering scenarios such as Chinese and Arabic, the character structure error rate generated by this solution is reduced by 65% ​​compared to the baseline model, while maintaining pixel-level rendering quality comparable to the main task.

[0088] Among them, glyph stroke order prediction refers to the sequential prediction of the stroke order of a character. Specifically, it can be implemented by using a bidirectional LSTM network combined with a masked cross-entropy loss function to constrain the model to learn the correct character topology. Character boundary detection refers to the generation of a pixel-level heat map of the character outline. Specifically, it can be implemented by using a U-Net architecture combined with a Dice loss function to maintain the integrity of the character's geometric shape. Language feature extraction refers to capturing the semantic context of the text. Specifically, it can be implemented by using a Transformer encoder combined with a cross-entropy loss function to enhance typesetting coherence. The dynamic loss balancing mechanism refers to a strategy that automatically adjusts the loss weight according to the correlation between tasks. Specifically, it can be implemented by using a temperature adjustment formula based on feature similarity to coordinate the gradient contributions of multiple tasks.

[0089] See also Figure 7 , Figure 7 A specific implementation of step S41 is shown, which is described in detail as follows:

[0090] S411: Predicting the stroke order sequence of each character based on the latent space representation by a stroke order predictor to generate the stroke order feature, and calculating the stroke order loss value based on the stroke order feature using a masked cross entropy loss function.

[0091] S412: Using a boundary detector to generate the character boundary heat map based on the latent space representation, and using a Dice loss function to calculate a boundary loss value based on the character boundary heat map.

[0092] S413: Extracting the language features based on the latent space representation through a speech feature extractor, and calculating a cross entropy loss value based on the language features.

[0093] S414: Generate the auxiliary task loss value based on the stroke order loss value, the boundary loss value and the cross entropy loss value.

[0094] Specifically, the stroke order predictor decodes the stroke order sequence of each character from the latent space representation, and filters non-critical stroke noise through a masking mechanism to ensure that the model focuses on learning the correct stroke order. The boundary detector converts the latent features into pixel-level heat maps, and uses the Dice loss function to strengthen the gradient response of the boundary area to solve the optimization deviation caused by the difference in the number of foreground and background pixels. The speech feature extractor maps the latent features to the semantic probability distribution through a linear layer, and uses the cross-entropy loss to constrain the consistency of the semantic space. The three auxiliary tasks provide supervision signals from the three dimensions of structure, space, and semantics respectively. The masked cross-entropy loss optimizes the local features of the stroke order, the Dice loss strengthens the global spatial relationship of the boundary area, and the cross-entropy loss ensures the correctness of the language logic. The three work together to improve the model's multi-level understanding ability of complex texts. This application effectively solves the gradient conflict problem of multi-task learning in the text rendering process, and improves the model's comprehensive modeling ability of glyph structure, spatial layout and semantic features through differentiated loss function design and dynamic weight adjustment mechanism. Specifically, the stroke order prediction task strengthens the topological correctness of character generation, preventing broken or incorrect strokes; the boundary detection task improves the clarity of character outlines, reducing edge blur or adhesion; and the language feature extraction task ensures the semantic coherence of the generated text and prevents semantic ambiguity. Furthermore, the shared latent space representation and dynamic loss balancing mechanism reduce resource consumption during training, improve model convergence efficiency, and demonstrate greater robustness in complex cross-language text rendering scenarios.

[0095] Among them, the masked cross entropy loss function refers to masking the specific stroke order when calculating the loss, ignoring the interference of irrelevant noise. Specifically, it can be implemented by multiplying a random mask matrix with the cross entropy loss, focusing on optimizing the prediction accuracy of key strokes. The Dice loss function is used to deal with the problem of imbalanced pixel categories at the boundary of characters. Specifically, it can be implemented by calculating the ratio of the intersection and union of the predicted heat map and the true boundary mask to enhance the sensitivity of the boundary pixels. The language feature extractor is based on the mapping of the latent space representation to the semantic space. Specifically, it can be implemented by combining the linear transformation layer with the Softmax activation function to capture the semantic relevance in the latent features.

[0096] S5: Acquire a Unicode code sequence to be rendered and metadata of a font to be rendered, and perform text rendering based on the Unicode code sequence to be rendered and the metadata of the font to be rendered by using the target text rendering model to generate a target text rendering image.

[0097] Specifically, in the inference stage or the actual application stage, the target text rendering model only retains the shared encoder and main decoder paths, automatically strips off all auxiliary task headers, and achieves zero additional computational overhead. The key to this design is that the gradient information of the auxiliary tasks has been passed to the model parameters through the shared encoder during the training phase, and there is no need to repeat the calculation during inference. In specific implementation, the system adopts conditional branch control to dynamically activate different calculation graph paths according to the operating mode (training or inference). Therefore, in an embodiment of the present application, the Unicode encoding sequence to be rendered and the font metadata to be rendered are obtained, and the target text rendering model is used to perform text rendering based on the Unicode encoding sequence to be rendered and the font metadata to be rendered to generate a target text rendering image.

[0098] The text rendering method based on multi-task auxiliary features provided in this application can be applied to various application scenarios in the financial technology field, such as personalized customer reports and marketing materials generation, automated compliance document generation, and real-time high-frequency trading report generation. In personalized customer reports and marketing materials generation, generating personalized investment proposals or financial product instructions based on customer profiles requires dynamic insertion of charts and emphasis on key content. It automatically annotates key data (such as yield and risk level) through semantic analysis and renders it in a striking font. While rendering the text, it predicts customer reading habits (such as font size preferences) and dynamically adjusts the layout. An example is generating a large-font, high-contrast retirement planning report for elderly customers while preserving the accuracy of professional terminology. The text rendering method based on multi-task auxiliary features can be applied to automated annotation of medical imaging reports in the healthcare field. Scenario requirement: Inserting precise text annotations (such as lesion location and size) in CT and MRI imaging reports, which must seamlessly blend with the image. Boundary detection assistance: Generate text heatmaps that align with image edges to avoid overlapping key anatomical structures. Dynamic resource allocation: Prioritize the clarity of key information (such as the "malignant tumor" label) when rendering dense annotations. Example: Dynamically annotate "nodule diameter 5mm" in a lung CT report and automatically adjust the text position to avoid obscuring image details.

[0099] In an embodiment of the present application, a Unicode code sequence and font metadata of a text are obtained, and the Unicode code sequence of the text and the font metadata are bimodally embedded to generate a joint representation of the text and the font; feature extraction is performed on the joint representation through a shared encoder to generate a latent space representation; the latent space representation is input into a main decoder for decoding to generate an initial text rendering image; multiple auxiliary tasks are performed in parallel based on the latent space representation, and the weights of the auxiliary tasks are adjusted through a dynamic loss balancing mechanism to train a text rendering model to generate a target text rendering model; the Unicode code sequence to be rendered and the font metadata to be rendered are obtained, and the target text rendering model performs text rendering based on the Unicode code sequence to be rendered and the font metadata to be rendered to generate a target text rendering image. The embodiment of the present invention generates a joint representation through bimodal embedding processing, combines a shared encoder with a dynamic loss balancing mechanism, automatically coordinates the weight relationship between multiple tasks during the training phase, and enhances structural understanding capabilities through parallel auxiliary tasks. Only the backbone network is retained in the inference phase. It has the advantages of resolving task conflicts in multi-task learning, improving glyph structure consistency, and enhancing cross-language rendering capabilities, which is conducive to improving the accuracy and efficiency of text rendering.

[0100] Please refer to Figure 8 , as a response to the above Figure 2 The present application provides an embodiment of a text rendering device based on a multi-task assist feature. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0101] like Figure 8 As shown, the text rendering device based on multi-task auxiliary features of this embodiment includes: a text data acquisition module 61, a feature extraction module 62, a feature decoding module 63, an auxiliary task parallel module 64 and a text rendering module 65, wherein:

[0102] The text data acquisition module 61 is used to acquire the Unicode encoding sequence and font metadata of the text, and perform bimodal embedding processing on the Unicode encoding sequence of the text and the font metadata to generate a joint representation of the text and the font.

[0103] The feature extraction module 62 is configured to extract features from the joint representation using a shared encoder to generate a latent space representation.

[0104] The feature decoding module 63 is configured to input the latent space representation into a main decoder for decoding to generate an initial text rendering image.

[0105] The auxiliary task parallel module 64 is used to execute multiple auxiliary tasks in parallel based on the latent space representation, and adjust the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train the text rendering model and generate a target text rendering model.

[0106] The text rendering module 65 is configured to obtain a Unicode code sequence to be rendered and metadata of a font to be rendered, and perform text rendering based on the Unicode code sequence to be rendered and the metadata of the font to be rendered using the target text rendering model to generate a target text rendering image.

[0107] Furthermore, the text data acquisition module 61 includes:

[0108] A data acquisition unit, configured to acquire the Unicode encoding sequence of the text and the font metadata;

[0109] A first mapping unit, configured to map the Unicode code sequence into a text embedding vector using a trainable codebook;

[0110] A second mapping unit, configured to map the font metadata into the font feature vector through a multi-layer perceptron;

[0111] A vector concatenation unit is configured to concatenate the text embedding vector and the font feature vector into the joint representation.

[0112] Furthermore, the feature extraction module 62 includes:

[0113] a joint representation input unit, configured to input the joint representation into the shared encoder;

[0114] an attention calculation unit, configured to perform attention calculation based on the joint representation through multiple attention mechanisms to generate an attention score;

[0115] A bias matrix injection unit, configured to inject a relative position bias matrix into each layer of the shared encoder to add the attention score to the relative position bias matrix to obtain an addition result;

[0116] A normalization processing unit is used to perform activation processing and normalization processing on the addition result to generate the latent space representation.

[0117] Furthermore, the feature decoding module 63 includes:

[0118] an upsampling unit, configured to input the latent space representation into the primary decoder, upsample the latent space representation layer by layer using a multi-layer transposed convolutional network, perform skip connections between each layer of the transposed convolutional network and the shared encoder to fuse features, and output upsampled features;

[0119] An upsampling feature mapping unit is used to map the upsampling features through a Tanh activation function to generate the initial text rendering image.

[0120] Furthermore, the auxiliary task parallel module 64 includes:

[0121] an auxiliary task execution unit, configured to perform glyph stroke order prediction, character boundary detection, and language feature extraction in parallel based on the latent space representation, generate stroke order features, character boundary heat maps, and language features, and generate an auxiliary task loss value based on the stroke order features, the character boundary heat maps, and the language features;

[0122] a weight coefficient calculation unit, configured to calculate a weight coefficient based on the initial text rendering image, the stroke order feature, the character boundary heat map, and the language feature through the dynamic loss balancing mechanism;

[0123] a total loss calculation unit, configured to calculate a total loss of the model based on the auxiliary task loss value and the weight coefficient;

[0124] An iterative training unit is used to adjust the parameters of the text rendering model based on the total loss of the model, and generate the target text rendering model by iteratively training the adjusted text rendering model.

[0125] Furthermore, the auxiliary task execution unit includes:

[0126] a first loss calculation unit, configured to predict a stroke sequence of each character based on the latent space representation using a stroke order predictor, generate the stroke order feature, and calculate a stroke order loss value based on the stroke order feature using a masked cross entropy loss function;

[0127] a second loss calculation unit, configured to generate the character boundary heat map based on the latent space representation using a boundary detector, and calculate a boundary loss value based on the character boundary heat map using a Dice loss function;

[0128] a third loss calculation unit, configured to extract the language features based on the latent space representation using a speech feature extractor, and calculate a cross entropy loss value based on the language features;

[0129] An auxiliary task loss calculation unit is used to generate the auxiliary task loss value based on the stroke order loss value, the boundary loss value and the cross entropy loss value.

[0130] Furthermore, the calculation formula of the weight coefficient is as follows:

[0131]

[0132] Among them, αi The weight coefficient of the i-th auxiliary task, τ is the temperature coefficient, Z is the latent space feature, Z i 、Z j are the features of the i-th and j-th auxiliary tasks.

[0133] To solve the above technical problems, the present application also provides a computer device. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0134] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are interconnected through a system bus. It should be noted that Figure 9 Only a computer device 7 having three components, memory 71, processor 72, and network interface 73, is shown. However, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. It should be understood by those skilled in the art that a computer device herein is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0135] Computer devices can be desktop computers, laptops, PDAs, cloud servers, etc. Computer devices can interact with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0136] The memory 71 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 7. Of course, the memory 71 may also include both the internal storage unit of the computer device 7 and its external storage devices. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed on the computer device 7, such as the program code of the text rendering method based on the multi-tasking assistance feature. In addition, the memory 71 can also be used to temporarily store various types of data that have been output or are to be output.

[0137] In some embodiments, the processor 72 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to execute program code stored in the memory 71 or process data, such as executing the program code of the above-mentioned text rendering method based on the multi-tasking assistance feature to implement various embodiments of the text rendering method based on the multi-tasking assistance feature.

[0138] The network interface 73 may include a wireless network interface or a wired network interface. The network interface 73 is generally used to establish a communication connection between the computer device 7 and other electronic devices.

[0139] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program, and the computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned text rendering method based on multi-tasking assistance features.

[0140] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.

[0141] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A text rendering method based on multi-task auxiliary features, characterized in that: include: Obtaining a Unicode code sequence and font metadata of a text, and performing bimodal embedding processing on the Unicode code sequence of the text and the font metadata to generate a joint representation of the text and the font; Extracting features from the joint representation using a shared encoder to generate a latent space representation; Inputting the latent space representation into a main decoder for decoding to generate an initial text rendering image; executing multiple auxiliary tasks in parallel based on the latent space representation, and adjusting the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train a text rendering model and generate a target text rendering model; A Unicode code sequence to be rendered and metadata of a font to be rendered are obtained, and text rendering is performed based on the Unicode code sequence to be rendered and the metadata of the font to be rendered by the target text rendering model to generate a target text rendering image.

2. The text rendering method based on multi-task auxiliary features according to claim 1, characterized in that: The acquiring of the Unicode code sequence and font metadata of the text, and performing bimodal embedding processing on the Unicode code sequence and the font metadata to generate a joint representation of the text and the font includes: Obtaining the Unicode encoding sequence of the text and the font metadata; Mapping the Unicode code sequence into a text embedding vector using a trainable codebook; Mapping the font metadata into the font feature vector through a multi-layer perceptron; The text embedding vector and the font feature vector are concatenated into the joint representation.

3. The text rendering method based on multi-task auxiliary features according to claim 1, characterized in that: The step of extracting features from the joint representation by a shared encoder to generate a latent space representation includes: inputting the joint representation into the shared encoder; Performing attention calculation based on the joint representation through multiple attention mechanisms to generate an attention score; Injecting a relative position bias matrix into each layer of the shared encoder to add the attention score to the relative position bias matrix to obtain an addition result; The addition result is activated and normalized to generate the latent space representation.

4. The text rendering method based on multi-task auxiliary features according to claim 1, characterized in that: Inputting the latent space representation into a main decoder for decoding to generate an initial text rendering image includes: Inputting the latent space representation into the main decoder, upsampling the latent space representation layer by layer using a multi-layer transposed convolutional network, and performing skip connection fusion of features between each layer of the transposed convolutional network and the shared encoder, and outputting upsampled features; The up-sampled features are mapped through a Tanh activation function to generate the initial text rendering image.

5. The text rendering method based on multi-task auxiliary features according to any one of claims 1 to 4, characterized in that: The method includes executing multiple auxiliary tasks in parallel based on the latent space representation and adjusting the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train a text rendering model and generate a target text rendering model, including: Based on the latent space representation, glyph stroke order prediction, character boundary detection, and language feature extraction are performed in parallel to generate stroke order features, character boundary heat maps, and language features, and an auxiliary task loss value is generated based on the stroke order features, the character boundary heat maps, and the language features; Calculating a weight coefficient based on the initial text rendering image, the stroke order feature, the character boundary heat map, and the language feature through the dynamic loss balancing mechanism; Calculating the total loss of the model based on the auxiliary task loss value and the weight coefficient; Parameters of the text rendering model are adjusted based on the total model loss, and the target text rendering model is generated by iteratively training the adjusted text rendering model.

6. The text rendering method based on multi-task auxiliary features according to claim 5, characterized in that: The method includes performing glyph stroke order prediction, character boundary detection, and language feature extraction in parallel based on the latent space representation to generate stroke order features, character boundary heat maps, and language features, and generating auxiliary task loss values ​​based on the stroke order features, the character boundary heat maps, and the language features, including: Predicting the stroke sequence of each character based on the latent space representation using a stroke order predictor to generate the stroke order feature, and calculating a stroke order loss value based on the stroke order feature using a masked cross entropy loss function; Using a boundary detector to generate the character boundary heat map based on the latent space representation, and using a Dice loss function to calculate a boundary loss value based on the character boundary heat map; Extracting the language features based on the latent space representation by a speech feature extractor, and calculating a cross entropy loss value based on the language features; The auxiliary task loss value is generated based on the stroke order loss value, the boundary loss value and the cross entropy loss value.

7. The text rendering method based on multi-task auxiliary features according to claim 5, characterized in that: The calculation formula of the weight coefficient is as follows: Among them, α i The weight coefficient of the i-th auxiliary task, τ is the temperature coefficient, Z is the latent space feature, Z i , Z j are the features of the i-th and j-th auxiliary tasks.

8. A text rendering device based on multi-task auxiliary features, characterized in that: include: A text data acquisition module is used to acquire the Unicode encoding sequence and font metadata of the text, and perform bimodal embedding processing on the Unicode encoding sequence of the text and the font metadata to generate a joint representation of the text and the font; a feature extraction module, configured to extract features from the joint representation using a shared encoder to generate a latent space representation; A feature decoding module, configured to input the latent space representation into a main decoder for decoding to generate an initial text rendering image; an auxiliary task parallel module, configured to execute multiple auxiliary tasks in parallel based on the latent space representation, and adjust the weights of the auxiliary tasks through a dynamic loss balancing mechanism to train the text rendering model and generate a target text rendering model; The text rendering module is used to obtain a Unicode code sequence to be rendered and metadata of a font to be rendered, and perform text rendering based on the Unicode code sequence to be rendered and the metadata of the font to be rendered by the target text rendering model to generate a target text rendering image.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the text rendering method based on the multi-task assist feature according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text rendering method based on the multi-tasking assist feature according to any one of claims 1 to 7 is implemented.