Style character generation method and device and computer readable storage medium

By constructing a style text generation model and utilizing multimodal feature extraction, cross-attention and diffusion network processing, we solved the problems of diversified requirements and insufficient imitation ability of traditional font generation methods and achieved high-quality style text generation.

CN120635252APending Publication Date: 2025-09-12HANHAI INFORMATION TECH SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510686825.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional font generation methods rely on a large font library, have limited ability to imitate specific font styles, cannot adapt to diverse user needs, and require users to provide multiple character samples to extract style features.

Method used

A style text generation model is constructed using multimodal feature extraction network, cross-attention network, gated fusion network, text semantic feature extraction network and diffusion network. Style text is generated through multimodal feature extraction, cross-attention processing, gated fusion and diffusion network processing, reducing dependence on font libraries and enhancing style imitation capabilities.

Benefits of technology

It achieves the generation of diverse styles of text without relying on a huge font library, improves the accuracy and generation quality of style imitation, and optimizes the user creation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635252A_ABST
    Figure CN120635252A_ABST
Patent Text Reader

Abstract

The invention provides a style character generation method and device and a computer readable storage medium, and relates to the technical field of content generation. The method comprises the following steps: respectively extracting a first character style feature of a character style text and a second character style feature of a character style image through a multi-modal feature extraction network; processing the first character style feature and the second character style feature through a cross attention network to obtain a third character style feature; processing the third character style feature and the second character style feature through a gating fusion network to obtain a fourth character style feature; extracting text semantic features of the content text through a text semantic feature extraction network; and processing the text semantic feature and the fourth character style feature through a diffusion network to obtain a content style text for the content text. By means of the technical means, the problems that in the related technology, style text generation needs to depend on a huge font library, and the capacity of simulating specific font styles is limited are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of content generation, and in particular to a method and device for generating style text, and a computer-readable storage medium. Background Art

[0002] In user-generated content scenarios, providing a rich and diverse range of font styles is crucial to enhancing the user experience. However, traditional font generation methods have several drawbacks: They rely on a vast font library to meet diverse needs, requiring the design of numerous font styles. They also have limited ability to mimic specific font styles or make subtle adjustments to existing fonts, making them incapable of adapting to diverse needs. Furthermore, to accurately extract the stylistic characteristics of a particular font, users are often required to provide multiple character samples. Summary of the Invention

[0003] The present disclosure provides a method, device, and computer-readable storage medium for generating stylized text, which at least to a certain extent meet the diverse needs of generating stylized text.

[0004] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0005] According to one aspect of the present disclosure, a method for generating style text is provided, comprising: constructing a style text generation model using a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network; inputting a content text, a text style text, and a text style image input by a target object into the style text generation model: extracting a first text style feature of the text style text and a second text style feature of the text style image respectively through a multimodal feature extraction network; processing the first text style feature and the second text style feature through a cross-attention network to obtain a third text style feature; processing the third text style feature and the second text style feature through a gated fusion network to obtain a fourth text style feature; extracting text semantic features of the content text through a text semantic feature extraction network; and processing the text semantic features and the fourth text style features through a diffusion network to obtain content style text for the content text.

[0006] In one embodiment of the present disclosure, the text style image is an image of at least one character written by the target object.

[0007] In one embodiment of the present disclosure, before processing the first text style feature and the second text style feature through a cross-attention network to obtain the third text style feature, the method also includes: when the target object only inputs content text, obtaining a preset text style image and a preset text style text; inputting the content text, the preset text style image and the preset text style text into a style text generation model; and extracting the first text style feature of the preset text style text and the second text style feature of the preset text style image through a multimodal feature extraction network.

[0008] In one embodiment of the present disclosure, the text semantic features and the fourth text style features are processed through a diffusion network to obtain content style text for the content text, including: adding noise to the text semantic features N times in a row through the diffusion network, where N is a positive integer greater than 1; based on the fourth text style features, predicting and removing the noise added to the text semantic features N times in a row through the diffusion network to obtain content style text.

[0009] In one embodiment of the present disclosure, before adding noise to the text semantic features N times in a row through the diffusion network, the method further includes: determining a complexity score of the text style text and the text style image; and determining N based on the complexity score, wherein the higher the complexity score, the larger N.

[0010] In one embodiment of the present disclosure, the complexity scores of text style text and text style images are determined, including: the text style text and text style image contain: font requirements, atmosphere requirements and special effects requirements; respectively determining the font score, atmosphere score and special effects score corresponding to the font requirements, atmosphere requirements and special effects requirements; and weightedly summing the font score, atmosphere score and special effects score according to preset weights to obtain a complexity score.

[0011] In one embodiment of the present disclosure, before inputting the content text, text style text and text style image of the target object into the style text generation model, the method further includes: obtaining training text, training style text and training style image, and inputting the training text, training style text and training style image into the style text generation model: extracting the fifth text style feature of the training style text and the sixth text style feature of the training style image respectively through a multimodal feature extraction network; processing the fifth text style feature and the sixth text style feature through a cross-attention network to obtain a seventh text style feature; processing the seventh text style feature and the sixth text style feature through a gated fusion network to obtain an eighth text style feature; extracting the training semantic feature of the training text through a text semantic feature extraction network; processing the training semantic feature and the eighth text style feature through a diffusion network to obtain a training style text for the training text; calculating the generation loss between the training style text and the corresponding label of the training text using a cross-entropy loss function; and optimizing the model parameters of the style text generation model based on the generation loss to complete the training of the style text generation model.

[0012] In one embodiment of the present disclosure, after extracting the training semantic features of the training text through the text semantic feature extraction network, the method further includes: the diffusion network includes a diffusion process and an inverse diffusion process; within the diffusion network: adding noise to the training semantic features through the diffusion process; based on the eighth text style feature, predicting and removing the noise added to the training semantic features in the diffusion process through the inverse diffusion process to obtain the training style text; using the mean square error loss function to calculate the diffusion loss between the noise added in the diffusion process and the noise predicted in the inverse diffusion process; and optimizing the model parameters of the diffusion network based on the diffusion loss to complete the training of the diffusion network.

[0013] According to another aspect of the present disclosure, a style text generation device is provided, characterized in that it includes: a construction module, configured to use a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network and a diffusion network to construct a style text generation model; an input module, configured to input the content text, text style text and text style image input by the target object into the style text generation model; a first extraction module, configured to extract the first text style feature of the text style text and the second text style feature of the text style image respectively through the multimodal feature extraction network; a first processing module, configured to process the first text style feature and the second text style feature through the cross-attention network to obtain a third text style feature; a second processing module, configured to process the third text style feature and the second text style feature through the gated fusion network to obtain a fourth text style feature; a second extraction module, configured to extract text semantic features of the content text through the text semantic feature extraction network; and a third processing module, configured to process the text semantic feature and the fourth text style feature through the diffusion network to obtain content style text for the content text.

[0014] According to yet another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above methods by executing the executable instructions.

[0015] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any of the above methods is implemented.

[0016] According to another aspect of the present disclosure, a computer program product is provided, including computer instructions stored in a computer-readable storage medium, and the computer instructions implement operating instructions of any of the above methods when executed by a processor.

[0017] In the embodiment of the present disclosure, a style text generation model is constructed using a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network; the content text, text style text, and text style image input by the target object are input into the style text generation model: the first text style feature of the text style text and the second text style feature of the text style image are extracted respectively by the multimodal feature extraction network; the first text style feature and the second text style feature are processed by the cross-attention network to obtain a third text style feature; the third text style feature and the second text style feature are processed by the gated fusion network to obtain a fourth text style feature; the text semantic feature of the content text is extracted by the text semantic feature extraction network; the text semantic feature and the fourth text style feature are processed by the diffusion network to obtain a content style text for the content text. Through the above technical means, the problems in the related art of relying on a huge font library to generate style text and the limited ability to imitate a specific font style are solved, thereby meeting diversified needs.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0020] Figure 1 A schematic diagram of a style text generation system in an embodiment of the present disclosure is shown.

[0021] Figure 2 A flow chart of a method for generating style text in an embodiment of the present disclosure is shown.

[0022] Figure 3 A flowchart of a content text processing method in an embodiment of the present disclosure is shown.

[0023] Figure 4 A flowchart of a complexity score calculation method in an embodiment of the present disclosure is shown.

[0024] Figure 5 A flowchart of a method for training a style text generation model in an embodiment of the present disclosure is shown.

[0025] Figure 6 A flowchart showing another method for generating stylized text in an embodiment of the present disclosure is shown.

[0026] Figure 7 A schematic diagram of a device for generating style text in an embodiment of the present disclosure is shown.

[0027] Figure 8 A schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0029] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0030] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0033] It should be pointed out that, in the absence of conflict, the embodiments of the present disclosure and the technical features therein may be combined with each other.

[0034] For ease of understanding, several terms involved in this disclosure are explained below:

[0035] A multimodal feature extraction network is a neural network structure used to extract feature information from data of different modalities (such as text and images). It can extract features of text and images separately and can use models such as CLIP, ViLBERT, and LXMERT.

[0036] Cross-Attention Network is a network that uses the attention mechanism to dynamically adjust the weights of input features in order to better capture the relationships between different types of data.

[0037] Gated Fusion Network is a method for integrating information from different sources or modalities, aiming to dynamically adjust the importance of each modality based on the input.

[0038] The text semantic feature extraction network is used to understand and extract the semantic features of the text, and can use BERT and RoBERTa, etc.

[0039] BERT (Bidirectional Encoder Representations from Transformers): A pre-trained natural language processing model that uses a bidirectional Transformer structure, which allows the model to take contextual information into account when processing text.

[0040] RoBERTa (Robustly Optimized BERT Approach): An improved version of the BERT model. The main goal of RoBERTa is to improve the performance of the model by optimizing the training process of BERT.

[0041] The Large Language Model (LLM) is a deep learning model designed to understand and generate natural language text. Based on a neural network architecture, particularly the Transformer architecture, the LLM is capable of handling a wide range of natural language processing tasks, such as text generation, machine translation, and question-answering, without requiring specialized tuning or training for each task.

[0042] Transformers are a type of deep learning model.

[0043] CLIP (Contrastive Language-Image Pre-training): A multimodal pre-training model. Its core idea is to embed images and text into the same high-dimensional space, so that the image and its related text description are close to each other in this space.

[0044] ViLBERT (Vision and Language BERT): A multimodal model for image-text matching tasks. The design philosophy of ViLBERT is to combine BERT's text processing capabilities with visual feature extraction. ViLBERT consists of two streams, one for processing text (language stream) and the other for processing images (visual stream). These two streams are connected through cross-modal interaction layers to learn the association between text and images.

[0045] LXMERT (Language and Cross-Modal BERT): A multimodal Transformer model designed for tasks such as visual question answering and image-text matching. LXMERT also contains three streams: a language stream, a visual stream, and a cross-modal stream. These streams are able to capture the complex interactions between different modalities and learn how to fuse them together. When processing language and visual information, LXMERT not only considers the internal relationships of a single modality, but also emphasizes the interactions between modalities. This makes LXMERT excel in tasks that require a deep understanding of visual content and textual descriptions.

[0046] Diffusion models are a class of deep learning models used for generative data. They excel at tasks such as image synthesis, audio generation, and other complex data types. Diffusion process: A diffusion network first defines a forward diffusion process that gradually converts real data into pure random noise. This process is typically a Markov chain, with each step adding a certain amount of noise to the data. Inverse diffusion process: In contrast to the forward process, the inverse diffusion process starts with random noise and gradually recovers the structure of the real data. This process is what the diffusion network needs to learn in order to reconstruct or generate data with a distribution similar to the real data.

[0047] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0048] Figure 1 A schematic diagram of a style text generation system according to an embodiment of the present disclosure is shown.

[0049] The user device 101 may be a mobile device such as a mobile phone, a tablet computer, an e-book reader, etc. The user device 101 may be installed with an application program to execute: uploading the text style image and / or text style text input by the target object to the server.

[0050] The server 102 may be composed of one server or several servers. An application program may be installed in the server 102 to execute the following steps: constructing a style text generation model using a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network; inputting the content text, text style text, and text style image of the target object into the style text generation model; extracting a first text style feature of the text style text and a second text style feature of the text style image respectively through the multimodal feature extraction network; processing the first text style feature and the second text style feature through the cross-attention network to obtain a third text style feature; processing the third text style feature and the second text style feature through the gated fusion network to obtain a fourth text style feature; extracting text semantic features of the content text through the text semantic feature extraction network; and processing the text semantic features and the fourth text style features through the diffusion network to obtain a content style text for the content text.

[0051] The user device 101 and the server 102 are connected via a communication network. Optionally, the communication network is a wired network or a wireless network.

[0052] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.

[0053] Figure 2 A flow chart of a method for generating style text in an embodiment of the present disclosure is shown. Figure 2 As shown, the following steps are included:

[0054] S201, uses multimodal feature extraction network, cross attention network, gated fusion network, text semantic feature extraction network and diffusion network to build a style text generation model.

[0055] The multimodal feature extraction network, cross attention network and gated fusion network are serially connected as one branch, and the text semantic feature extraction network is another branch. Both branches are connected to the diffusion network to obtain a style text generation model.

[0056] S202: Input the content text, text style text and text style image of the target object into a style text generation model.

[0057] The content text is the basic text content for which the user wishes to generate stylized text, for example, a piece of text that needs to be presented in a specific style.

[0058] The text style text is text content that describes the target font style, and can be a text description or sample characters related to a specific font style.

[0059] The text style image is image data containing a specific font style, such as an image sample of a certain calligraphy style or printing font.

[0060] S203 , extracting first text style features of the text style text and second text style features of the text style image respectively through a multimodal feature extraction network.

[0061] The first text style feature is vector information representing the font style of the text extracted from the text style text through a multimodal feature extraction network.

[0062] The second text style feature is vector information representing the font style of the image extracted from the text style image through a multimodal feature extraction network.

[0063] The multimodal feature extraction network extracts the corresponding font style features by processing text style text and text style images separately. Through this technical means, font styles can be learned from a small number of examples, thus avoiding dependence on font libraries.

[0064] S204 , processing the first text style feature and the second text style feature through a cross attention network to obtain a third text style feature.

[0065] The cross-attention network fuses the first and second text style features to generate a more comprehensive third text style feature. This technology captures the correlation between text style text and text style image, improving the accuracy of style transfer.

[0066] S205 , processing the third text style feature and the second text style feature through a gated fusion network to obtain a fourth text style feature.

[0067] The gated fusion network weights the third and second text style features to generate a fourth text style feature. The gated fusion network selectively retains useful information, generating a more accurate font style feature. This technology optimizes feature representation and reduces the impact of redundant information.

[0068] S206: extracting text semantic features of the content text through a text semantic feature extraction network.

[0069] Text semantic features represent the semantic meaning of the text and guide the diffusion network to generate text that aligns with the semantic style of the content. The text semantic feature extraction network ensures that the generated results are semantically consistent with the original content, preventing semantic deviations caused by style transfer.

[0070] S207 , processing the text semantic feature and the fourth text style feature through a diffusion network to obtain a content style text for the content text.

[0071] The diffusion network is responsible for combining the text semantic features with the fourth text style features to gradually generate text images with a specified style.

[0072] Content-style text is the final generated result, which is the conversion of content text into a text image with a specified text style.

[0073] The embodiments of the present disclosure use the above-mentioned technical means to solve the problems in the related art of relying on a huge font library to generate stylized text and having limited ability to imitate specific font styles, thereby meeting diverse needs.

[0074] For example, when designing a poster, a user might want to present a text in a certain calligraphy style. The user uploads a text description of the calligraphy style (text style text) and an image of the calligraphy work (text style image), along with the original text to be converted (content text). Using the aforementioned technical means, the font style information is extracted and integrated, combined with the semantics of the original text, to generate content style text in the calligraphy style, thus enabling personalized text generation without the need for font library support.

[0075] In one embodiment of the present disclosure, the text style image is an image of at least one character written by the target object.

[0076] In one exemplary embodiment, the characters in the text-themed image have specific fonts, ambiance, and special effects. The fonts include, but are not limited to, font shape, size, proportion, and stroke design. The ambiance includes, but is not limited to, cute and handsome. The special effects include, but are not limited to, shadows, glows, textures, and gradients.

[0077] In one exemplary embodiment, the text style text is used to specify that a word should have a specific font, atmosphere, and special effects.

[0078] In an exemplary embodiment, each text style feature includes a font style feature, an atmosphere style feature, and a special effect style feature.

[0079] Through these technical means, we enhance the ability to capture font style characteristics, improve the accuracy of font style imitation, and enhance the overall quality of generated content style text. At the same time, this method reduces the reliance on a large number of font libraries and optimizes the user's creative experience.

[0080] Figure 3 A flowchart of a content text processing method according to an embodiment of the present disclosure is shown. Figure 3 As shown, the following steps are included:

[0081] S301, when the target object only inputs content text, obtaining a preset text style image and a preset text style text;

[0082] S302: Input the content text, the preset text style image and the preset text style text into the style text generation model:

[0083] S303 , extracting first text style features of a text with a preset text style and second text style features of an image with a preset text style respectively through a multimodal feature extraction network.

[0084] Preset text style images are predefined or stored images that contain specific font styles, moods, and special effects. These images are typically prepared in advance to provide a reference for font styles. Preset text style text is predefined or stored descriptive content that describes the characteristics of a specific font style, including the font, mood, and special effects.

[0085] In this embodiment, if the target user does not input a text style image or text style text, the first text style features of the preset text style text and the second text style features of the preset text style image are extracted. Through the above technical means, support for diverse font styles is enhanced, the efficiency and accuracy of style extraction are improved, and user creative flexibility is increased. Furthermore, this approach reduces the need for real-time user input and the reliance on a large font library, optimizing the overall creative experience.

[0086] In an optional embodiment, when the target object only inputs content text and a text style image, a preset text style text is obtained; the content text, the text style image and the preset text style text are input into a style text generation model: and the first text style feature of the preset text style text and the second text style feature of the text style image are respectively extracted through a multimodal feature extraction network.

[0087] In an optional embodiment, when the target object only inputs content text and text style text, a preset text style image is obtained; the content text, the preset text style image and the text style text are input into a style text generation model: and the first text style feature of the text style text and the second text style feature of the preset text style image are respectively extracted through a multimodal feature extraction network.

[0088] In one embodiment of the present disclosure, the text semantic features and the fourth text style features are processed through a diffusion network to obtain content style text for the content text, including: adding noise to the text semantic features N times in a row through the diffusion network, where N is a positive integer greater than 1; based on the fourth text style features, predicting and removing the noise added to the text semantic features N times in a row through the diffusion network to obtain content style text.

[0089] Noise addition involves superimposing a certain amount of random interference on the semantic features of text, causing it to deviate from its original state. This operation simulates the data uncertainty that can occur in the real world. Predicting and removing noise involves the diffusion network inferring the noise components contained in the current features based on existing text style features and attempting to restore them to a feature representation that is closer to the target style.

[0090] In this embodiment, the diffusion network first adds noise to the text semantic features N times in a row, gradually disrupting the original semantic information. Then, based on the fourth text style feature, the diffusion network predicts and removes the added noise N times in a row, gradually restoring a text image with the specified style. Each denoising operation is guided by the fourth text style feature, ensuring that the generated result meets the requirements of the target style. Through this technical approach, text images that conform to a specific style can be gradually constructed without relying on a font library, improving generation quality, enhancing the model's ability to imitate complex styles, and enhancing the visual consistency of the generated results.

[0091] In one embodiment of the present disclosure, before adding noise to the text semantic features N times in a row through the diffusion network, the method further includes: determining a complexity score of the text style text and the text style image; and determining N based on the complexity score, wherein the higher the complexity score, the larger N.

[0092] The complexity score quantifies the structural details or semantic complexity of text and image styles. A higher complexity score indicates a more complex style and greater difficulty in generating it. For example, calligraphy typically has a higher complexity score than printed fonts because its strokes are more varied and its structure is more flexible.

[0093] In this embodiment, before the diffusion network begins processing, the complexity of the input text style text and text style image will be analyzed, and the number of iterations N of the diffusion process will be adjusted accordingly. Specifically, the complexity of the text style text and text style image can be evaluated separately by a preset evaluation model (the evaluation model can be a trained large language model) to obtain corresponding complexity scores. The number of times N that noise is added and removed in the diffusion network is determined based on the score. The above technical means can dynamically adjust the sophistication of the generation process according to the complexity of different styles, improve the quality adaptability of the generated results, and avoid the waste of resources caused by a fixed number of iterations.

[0094] Figure 4 A flowchart of a complexity score calculation method according to an embodiment of the present disclosure is shown. Figure 4 As shown, the following steps are included:

[0095] Text style text and text style images include: font requirements, atmosphere requirements and special effects requirements;

[0096] S401, determining the font score, atmosphere score, and special effect score corresponding to the font requirement, atmosphere requirement, and special effect requirement respectively;

[0097] S402: weighted summing the font score, atmosphere score, and special effect score according to preset weights to obtain a complexity score.

[0098] Font requirements are the specific description or presentation of font morphological features in text and text-style images, such as whether they include ligatures, decorative lines, or unique stroke structures. A font score is a numerical value derived from an evaluation of font requirements. A higher score on the font complexity indicator indicates a more complex font.

[0099] Atmosphere requirements are the overall visual atmosphere requirements for text and text-style images, such as abstract stylistic attributes like retro, technological, or dreamlike. The Atmosphere Score is a numerical value evaluated based on these requirements, used to measure the abstractness and diversity of the atmosphere. A higher score indicates a more difficult atmosphere to reproduce.

[0100] Special effects requirements refer to the need for additional visual effects in text-style text and text-style images, such as shadows, gradients, glows, and other graphical processing effects. The special effects score is a numerical value evaluated based on the special effects requirements, used to measure the number of special effects and the difficulty of implementation. A higher score indicates more complex special effects.

[0101] The preset weights are weighting coefficients set for the font score, atmosphere score, and special effects score respectively, which are used to control the proportion of the influence of the three in the final complexity score.

[0102] In this embodiment, the complexity score is decomposed into three dimensions: font, atmosphere, and special effects, and then evaluated separately and weighted fusion is performed. Specifically: the font requirements, atmosphere requirements, and special effects requirements in the input text style text and text style image are identified, and the font score, atmosphere score, and special effects score are calculated using the corresponding evaluation models. Then, these three scores are weighted and summed according to the preset weights to obtain the final complexity score. The above technical means can more finely quantify the style complexity, avoid the deviation caused by single-dimensional evaluation, and improve the adaptation accuracy of the generation process.

[0103] Figure 5 A flow chart of a method for training a style text generation model according to an embodiment of the present disclosure is shown. Figure 5 As shown, the following steps are included:

[0104] S501: Obtain training text, training style text, and training style image, and input the training text, training style text, and training style image into a style text generation model:

[0105] S502, extracting a fifth text style feature of the training style text and a sixth text style feature of the training style image respectively through a multimodal feature extraction network;

[0106] S503, processing the fifth character style feature and the sixth character style feature through a cross attention network to obtain a seventh character style feature;

[0107] S504, processing the seventh character style feature and the sixth character style feature through a gated fusion network to obtain an eighth character style feature;

[0108] S505, extracting training semantic features of the training text through a text semantic feature extraction network;

[0109] S506, processing the training semantic features and the eighth text style features through a diffusion network to obtain a training style text for the training text;

[0110] S507, using a cross entropy loss function to calculate the generation loss between the training style text and the corresponding label of the training text;

[0111] S508 , optimizing the model parameters of the style text generation model according to the generation loss to complete the training of the style text generation model.

[0112] Training text is the basic text data used in the model training phase, guiding the model to learn correct semantic representations. Training style text is the text style description information used for training, and together with training style images, it constitutes the style input source. Training style images are font style image samples used for training, and together with training style text, they provide multimodal style information.

[0113] The fifth text style feature is a font style vector extracted from the training style text using a multimodal feature extraction network. The sixth text style feature is a font style vector extracted from the training style image using a multimodal feature extraction network. The seventh text style feature is a fusion feature obtained by processing the fifth and sixth text style features using a cross-attention network. The eighth text style feature is the final integrated feature obtained by processing the seventh and sixth text style features using a gated fusion network. The training semantic feature is a semantic representation extracted from the training text using a text semantic feature extraction network and is used to ensure semantic consistency during the generation process.

[0114] The training style text is the simulated output generated by the diffusion network based on the training semantic features and the eighth text style features. It is used to compare with the true labels to calculate the loss. The generation loss is a measure of the difference between the training style text and the corresponding labels of the training text using the cross-entropy loss function, reflecting the current model generation quality.

[0115] In this embodiment, training text, training style text, and training style images are obtained and processed in sequence through a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network to obtain a training style text. Subsequently, the generation loss is calculated using a cross-entropy loss function, and the model parameters are back-propagated based on the loss to complete the training process. The above technical means can effectively improve the model's ability to learn font styles and semantic features, and enhance the accuracy and consistency of the generated results.

[0116] In one embodiment of the present disclosure, after extracting the training semantic features of the training text through the text semantic feature extraction network, the method further includes: the diffusion network includes a diffusion process and an inverse diffusion process; within the diffusion network: adding noise to the training semantic features through the diffusion process; based on the eighth text style feature, predicting and removing the noise added to the training semantic features in the diffusion process through the inverse diffusion process to obtain the training style text; using the mean square error loss function to calculate the diffusion loss between the noise added in the diffusion process and the noise predicted in the inverse diffusion process; and optimizing the model parameters of the diffusion network based on the diffusion loss to complete the training of the diffusion network.

[0117] Diffusion loss is a numerical value that measures the difference between the actual noise in the diffusion process and the noise predicted by the inverse diffusion process through the mean square error loss function, which is used to guide the optimization of diffusion network parameters.

[0118] In this embodiment, after extracting the training semantic features, noise is first added to these features through a diffusion process within the diffusion network to form a noisy representation. Then, based on the eighth text style feature, this noise is predicted and removed through an inverse diffusion process to obtain the training style text. A mean square error loss function is then used to calculate the diffusion loss between the actual noise added during the diffusion process and the model-predicted noise. This loss is then used to backpropagate and optimize the model parameters of the diffusion network, completing the training of the diffusion network. This technical approach allows for more accurate learning of noise distribution patterns, improving the ability to restore detail during style transfer.

[0119] Figure 6 A flow chart of another method for generating style text in an embodiment of the present disclosure is shown. Figure 6 As shown, the following steps are included:

[0120] The multimodal feature extraction network extracts text style features from text style images and / or text style text; the text semantic feature extraction network extracts text semantic features from the content text. The diffusion network generates content style text based on the text style features and text semantic features.

[0121] Based on the same inventive concept, the present disclosure also provides a style text generation device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.

[0122] Figure 7 A schematic diagram of a device for generating style text according to an embodiment of the present disclosure is shown. Figure 7 As shown, the style text generating device may include:

[0123] A construction module 701 is configured to construct a style text generation model using a multimodal feature extraction network, a cross attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network;

[0124] The input module 702 is configured to input the content text, text style text and text style image of the target object into the style text generation model:

[0125] A first extraction module 703 is configured to extract a first text style feature of the text style text and a second text style feature of the text style image respectively through a multimodal feature extraction network;

[0126] A first processing module 704 is configured to process the first text style feature and the second text style feature through a cross attention network to obtain a third text style feature;

[0127] The second processing module 705 is configured to process the third text style feature and the second text style feature through a gated fusion network to obtain a fourth text style feature;

[0128] The second extraction module 706 is configured to extract text semantic features of the content text through a text semantic feature extraction network;

[0129] The third processing module 707 is configured to process the text semantic feature and the fourth text style feature through a diffusion network to obtain a content style text for the content text.

[0130] In some embodiments, the first extraction module 703 is also configured to obtain a preset text style image and a preset text style text when the target object only inputs content text; input the content text, the preset text style image and the preset text style text into the style text generation model: and extract the first text style feature of the preset text style text and the second text style feature of the preset text style image respectively through the multimodal feature extraction network.

[0131] In some embodiments, the third processing module 707 is further configured to add noise to the text semantic features N times in a row through a diffusion network, where N is a positive integer greater than 1; based on the fourth text style feature, the noise added to the text semantic features is predicted and removed N times in a row through a diffusion network to obtain a content style text.

[0132] In some embodiments, the third processing module 707 is further configured to determine a complexity score of the text style text and the text style image; and determine N based on the complexity score, wherein the higher the complexity score, the greater N.

[0133] In some embodiments, the third processing module 707 is further configured so that the text style text and the text style image include: font requirements, atmosphere requirements and special effects requirements; respectively determine the font score, atmosphere score and special effects score corresponding to the font requirements, atmosphere requirements and special effects requirements; and weightedly sum the font score, atmosphere score and special effects score according to preset weights to obtain a complexity score.

[0134] In some embodiments, the input module 702 is further configured to obtain training text, training style text and training style image, and input the training text, training style text and training style image into the style text generation model: extract the fifth text style feature of the training style text and the sixth text style feature of the training style image respectively through the multimodal feature extraction network; process the fifth text style feature and the sixth text style feature through the cross-attention network to obtain the seventh text style feature; process the seventh text style feature and the sixth text style feature through the gated fusion network to obtain the eighth text style feature; extract the training semantic feature of the training text through the text semantic feature extraction network; process the training semantic feature and the eighth text style feature through the diffusion network to obtain the training style text for the training text; use the cross entropy loss function to calculate the generation loss between the training style text and the corresponding label of the training text; optimize the model parameters of the style text generation model based on the generation loss to complete the training of the style text generation model.

[0135] In some embodiments, the input module 702 is further configured as a diffusion network including a diffusion process and an inverse diffusion process; within the diffusion network: noise is added to the training semantic features through the diffusion process; based on the eighth text style feature, the noise added to the training semantic features during the diffusion process is predicted and removed through the inverse diffusion process to obtain the training style text; the mean square error loss function is used to calculate the diffusion loss between the noise added during the diffusion process and the noise predicted during the inverse diffusion process; and the model parameters of the diffusion network are optimized based on the diffusion loss to complete the training of the diffusion network.

[0136] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0137] Refer to the following Figure 8 800 according to this embodiment of the present disclosure will be described. Figure 8 The electronic device 800 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0138] like Figure 8 As shown, electronic device 800 is implemented as a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, the aforementioned at least one processing unit 810, the aforementioned at least one storage unit 820, and a bus 830 connecting various system components (including storage unit 820 and processing unit 810).

[0139] Among them, the storage unit stores a program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps described in the above "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 810 can execute the following steps of the above method embodiment: using a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network and a diffusion network to construct a style text generation model; inputting the content text, text style text and text style image of the target object into the style text generation model: extracting the first text style feature of the text style text and the second text style feature of the text style image respectively through the multimodal feature extraction network; processing the first text style feature and the second text style feature through the cross-attention network to obtain a third text style feature; processing the third text style feature and the second text style feature through the gated fusion network to obtain a fourth text style feature; extracting the text semantic feature of the content text through the text semantic feature extraction network; processing the text semantic feature and the fourth text style feature through the diffusion network to obtain a content style text for the content text.

[0140] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .

[0141] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0142] Bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0143] The electronic device 800 may also communicate with one or more external devices 840 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more styled text generators that enable a user to interact with the electronic device 800, and / or any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication may occur via an input / output (I / O) interface 850. Furthermore, the electronic device 800 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0144] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0145] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided. The computer-readable storage medium may be a readable signal medium or a readable storage medium.

[0146] In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present disclosure described in the above "Specific Implementation Methods" section of this specification.

[0147] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0149] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0150] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0151] The present disclosure provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for generating styled text provided in any of the optional embodiments of the present disclosure.

[0152] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0153] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0154] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0155] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope of the present disclosure being indicated by the appended claims.

Claims

1. A method for generating style text, characterized in that: include: A style text generation model is constructed using a multimodal feature extraction network, a cross-attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network. Input the target object's input content text, text style text, and text style image into the style text generation model: respectively extracting a first text style feature of the text style text and a second text style feature of the text style image through the multimodal feature extraction network; Processing the first text style feature and the second text style feature through the cross attention network to obtain a third text style feature; Processing the third text style feature and the second text style feature through the gated fusion network to obtain a fourth text style feature; Extracting text semantic features of the content text through the text semantic feature extraction network; The text semantic feature and the fourth text style feature are processed by the diffusion network to obtain a content style text for the content text.

2. The method according to claim 1, characterized in that The text style image is an image of at least one character written by the target object.

3. The method according to claim 1, characterized in that Before processing the first text style feature and the second text style feature through the cross attention network to obtain the third text style feature, the method further includes: When the target object only inputs the content text, obtaining a preset text style image and a preset text style text; Input the content text, the preset text style image and the preset text style text into the style text generation model: The first text style feature of the preset text style text and the second text style feature of the preset text style image are extracted respectively through the multimodal feature extraction network.

4. The method according to claim 1, wherein Processing the text semantic feature and the fourth text style feature through a diffusion network to obtain a content style text for the content text includes: Adding noise to the text semantic feature N times continuously through the diffusion network, where N is a positive integer greater than 1; Based on the fourth text style feature, the diffusion network is used to continuously predict N times and remove noise added to the text semantic feature to obtain the content style text.

5. The method according to claim 4, characterized in that Before adding noise to the text semantic feature N times continuously through the diffusion network, the method further includes: Determining complexity scores of the text style text and the text style image; N is determined based on the complexity score, wherein the higher the complexity score, the larger N is.

6. The method according to claim 5, characterized in that Determining complexity scores of the text style text and the text style image includes: The text style text and the text style image include: font requirements, atmosphere requirements and special effect requirements; respectively determining a font score, an atmosphere score, and a special effect score corresponding to the font requirement, the atmosphere requirement, and the special effect requirement; The font score, the atmosphere score, and the special effect score are weighted and summed according to preset weights to obtain the complexity score.

7. The method according to claim 1, characterized in that Before inputting the content text, text style text and text style image input by the target object into the style text generation model, the method further includes: Acquire training text, training style text, and training style image, and input the training text, the training style text, and the training style image into the style text generation model: respectively extracting a fifth text style feature of the training style text and a sixth text style feature of the training style image through the multimodal feature extraction network; Processing the fifth and sixth text style features through the cross attention network to obtain a seventh text style feature; Processing the seventh character style feature and the sixth character style feature through the gated fusion network to obtain an eighth character style feature; Extracting training semantic features of the training text through the text semantic feature extraction network; Processing the training semantic features and the eighth text style features through the diffusion network to obtain a training style text for the training text; Calculating the generation loss between the training style text and the label corresponding to the training text using a cross entropy loss function; The model parameters of the style text generation model are optimized according to the generation loss to complete the training of the style text generation model.

8. The method according to claim 7, characterized in that After extracting the training semantic features of the training text through the text semantic feature extraction network, the method further includes: The diffusion network includes a diffusion process and a reverse diffusion process; Inside the diffusion network: adding noise to the training semantic features through the diffusion process; Based on the eighth text style feature, predicting and removing noise added to the training semantic feature during the diffusion process through the inverse diffusion process to obtain the training style text; Calculating the diffusion loss between the noise added in the diffusion process and the noise predicted in the inverse diffusion process using a mean square error loss function; The model parameters of the diffusion network are optimized according to the diffusion loss to complete the training of the diffusion network.

9. A device for generating style text, characterized in that: include: A building module is configured to construct a style text generation model using a multimodal feature extraction network, a cross attention network, a gated fusion network, a text semantic feature extraction network, and a diffusion network; An input module is configured to input the content text, text style text, and text style image input by the target object into the style text generation model: a first extraction module configured to extract a first text style feature of the text style text and a second text style feature of the text style image respectively through the multimodal feature extraction network; a first processing module configured to process the first text style feature and the second text style feature through the cross attention network to obtain a third text style feature; a second processing module configured to process the third text style feature and the second text style feature through the gated fusion network to obtain a fourth text style feature; A second extraction module is configured to extract text semantic features of the content text through the text semantic feature extraction network; The third processing module is configured to process the text semantic feature and the fourth text style feature through the diffusion network to obtain a content style text for the content text.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.