Artificial intelligence-based text wordart method, apparatus and device, and medium

Through the Wensheng Art Character Method based on artificial intelligence, the large language model and diffusion model are used to generate art characters, which solves the problems of low efficiency and poor quality of art characters generation in the existing technology, and achieves an efficient and controllable art character generation effect.

CN120124588APending Publication Date: 2025-06-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510192051.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, there are problems of low efficiency and poor quality in the generation of artistic characters. Especially when pursuing unique artistic effects, designers need to spend a lot of time and energy to make fine adjustments, which cannot meet the modern society's demand for efficient creative artistic characters.

Method used

The Wensheng Art Word Method based on artificial intelligence is adopted, and the initial model and target loss function are constructed, and the large language model, discriminator and diffusion model are used to generate art words. The method includes rasterizing the target text, identifying semantic region coordinates, expanding style prompt words, and generating artistic characters that combine glyph recognition and artistic semantics through an adversarial training mechanism.

Benefits of technology

The efficiency and quality of artistic characters are improved, the higher controllability of text art generation and the high integration and balance between style and font are achieved, and the problems of low efficiency and poor quality in the existing technology are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124588A_ABST
    Figure CN120124588A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and finance, and provides an artificial intelligence-based text wordart method, device, equipment and medium, on one hand, a target text is rasterized based on a text wordart model to obtain a target raster image text, a vector text layer is converted into a pixel layer, and the pixel layer is converted into a pixel layer; the text is not limited by language types, the problem that the Chinese language structure is complex is avoided, and the text-to-graph problem is converted into the graph-to-graph problem; on one hand, a large language model is used for recognizing regional coordinates used for semantization in raster image characters and expanding style cue words, and the controllability degree of text art generation is improved; and on the other hand, the discriminator is utilized to guide the diffusion model to generate a font which can be obviously recognized, and meanwhile, through an adversarial training mechanism, the generated wordart has the advantages of clear font recognition and artistic semantic change, so that the style and the font are highly fused and balanced, and the problems of low efficiency and poor quality of the text wordart are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and financial technology, and particularly to a method, device, equipment and medium for generating artistic characters from text based on artificial intelligence. Background Art

[0002] As a unique font expression form, artistic characters are deformed fonts with unique visual effects formed on the basis of ordinary characters through elaborate artistic processing and creative transformation. They play a crucial role in today's design field, showing extremely high application value. Their application scenarios are very extensive, covering multiple creative design categories. For example, in the financial field, artistic characters can be used for publicity and promotion.

[0003] In the prior art, generating artistic characters faces two core challenges. One is typesetting, which requires reasonable arrangement of character spaces to achieve a beautiful and coordinated layout, including elements such as character spacing, proportion, alignment, etc., to form a unified and harmonious overall structure. The other is font design, which is the core of artistic characters, involving morphological innovation, stroke shaping, addition of decorative elements and color matching, aiming to show a unique artistic style, endow the font with personality and emotional expression, and enhance the visual impact.

[0004] In the traditional mode of artistic character creation, using professional design tools has always been the main design method. Designers rely on these tools to exert their professional skills to shape artistic characters, but this method has significant defects. Its operation is complex and requires extremely high professional capabilities, resulting in extremely low creation efficiency and being difficult to quickly generate highly artistic characters. Especially in the pursuit of unique artistic effects of characters, the limitations of these tools are prominent, and designers often need to spend a lot of time and energy on fine-tuning, unable to meet the efficient and creative creation needs of artistic characters in modern society.

[0005] In recent years, the rapid development of deep learning technology has injected new vitality into the field of artistic character generation. Many researchers have actively explored and used generative models to open up new paths for automatic generation of artistic characters. In this exploration process, the generation of artistic characters has gradually changed from simple basic forms to showing rich artistic styles, and significant progress has been made in improving the artistic expressiveness and accuracy of characters, providing important support for subsequent in-depth research and innovation.

[0006] Early generative models either could not generate accurate text images or lacked the ability to artfully deform text, and usually required designers to assist in repeated design, with low work efficiency as well.

[0007] In view of the above problems, it is necessary to provide a method for generating artistic characters from text based on artificial intelligence to improve the generation efficiency and quality of artistic characters. Summary of the Invention

[0008] In view of the above, it is necessary to provide a method, device, equipment and medium for generating artistic words from text based on artificial intelligence, aiming to solve the problems of low efficiency and poor quality in generating artistic words from text.

[0009] A method for generating artistic words from text based on artificial intelligence, the method for generating artistic words from text based on artificial intelligence includes:

[0010] Construct an initial model and construct a target loss function; wherein, the initial model includes a fine-tuned large language model, a discriminator, a first diffusion model and a second diffusion model;

[0011] Train the initial model based on the target loss function to obtain a model for generating artistic words from text;

[0012] In response to an artistic word generation instruction based on a target text and an initial style prompt, rasterize the target text based on the model for generating artistic words from text to obtain target raster image text;

[0013] Input the target raster image text and the initial style prompt into the large language model to obtain the coordinates of the target area for semanticization in the target raster image text, and an extended target style prompt;

[0014] Input the target style prompt into the first diffusion model to obtain a target style image;

[0015] Input the target raster image text, the target area coordinates, the target style prompt and the target style image into the second diffusion model to obtain target artistic words.

[0016] A device for generating artistic words from text based on artificial intelligence, the device for generating artistic words from text based on artificial intelligence includes:

[0017] A construction unit for constructing an initial model and constructing a target loss function; wherein, the initial model includes a fine-tuned large language model, a discriminator, a first diffusion model and a second diffusion model;

[0018] A training unit for training the initial model based on the target loss function to obtain a model for generating artistic words from text;

[0019] A rasterization unit for rasterizing the target text based on the model for generating artistic words from text in response to an artistic word generation instruction based on a target text and an initial style prompt to obtain target raster image text;

[0020] An input unit for inputting the target raster image text and the initial style prompt into the large language model to obtain the coordinates of the target area for semanticization in the target raster image text and the extended target style prompt;

[0021] The input unit is further configured to input the target style prompt into the first diffusion model to obtain a target style image;

[0022] The input unit is further configured to input the target raster image text, the target area coordinates, the target style prompt, and the target style image into the second diffusion model to obtain target artistic characters.

[0023] A computer device, the computer device includes:

[0024] A memory storing at least one instruction; and

[0025] A processor that executes the instructions stored in the memory to implement the method for generating artistic characters from text based on artificial intelligence.

[0026] A computer-readable storage medium storing at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the method for generating artistic characters from text based on artificial intelligence.

[0027] It can be seen from the above technical solutions that, on the one hand, rasterizing the target text based on the model for generating artistic characters from text to obtain the target raster image text, converting the vector text layer into a pixel layer, making the text not restricted by language types, avoiding the problem of complex Chinese language structures, and transforming the problem of generating images from text into the problem of generating images from images; on the other hand, using the large language model to identify the area coordinates for semanticization in the raster image text and expand the style prompt, improving the controllability of text art generation; on the other hand, using the discriminator to guide the diffusion model to generate clearly recognizable glyphs, and at the same time through the adversarial training mechanism, making the generated artistic characters have both clear glyph recognition and artistic semantic changes, achieving a high degree of integration and balance between style and glyph, thus solving the problems of low efficiency and poor quality in generating artistic characters from text. Description of the Drawings

[0028] Figure 1 is a flowchart of a preferred embodiment of the method for generating artistic characters from text based on artificial intelligence of the present invention.

[0029] Figure 2 is a functional module diagram of a preferred embodiment of the device for generating artistic characters from text based on artificial intelligence of the present invention.

[0030] Figure 3It is a schematic structural diagram of a computer device which is a preferred embodiment of the method for generating artistic characters from text based on artificial intelligence in the present invention. Detailed implementation manners

[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] As Figure 1 shown, it is a flowchart of a preferred embodiment of the method for generating artistic characters from text based on artificial intelligence in the present invention. According to different requirements, the order of steps in this flowchart can be changed and some steps can be omitted.

[0033] The method for generating artistic characters from text based on artificial intelligence is applied to one or more computer devices. The computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0034] The computer device can be any electronic product that can perform human-computer interaction with a user. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.

[0035] The computer device may further include a network device and / or a user device. Among them, the network device includes but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (Cloud Computing).

[0036] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0037] Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0038] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0039] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, Virtual Private Network (VPN), etc.

[0040] S10, construct an initial model and a target loss function; among them, the initial model includes a fine-tuned Large Language Model (LLM), a discriminator, a First Diffusion model, and a Second Diffusion model.

[0041] In this embodiment, the construction of the target loss function includes:

[0042] Construct the discriminator loss using the following formula:

[0043]

[0044] Among them, L dis represents the discriminator loss; D(I T ) represents the output data after inputting the raster image text I T into the discriminator; represents the output data after inputting the artistic word output by the second diffusion model into the discriminator;

[0045] Construct the diffusion model loss using the following formula:

[0046]

[0047] Among them, L dif represents the diffusion model loss; ε represents the original noise added to the second diffusion model; ε represents the predicted noise of the second diffusion model;

[0048] Construct the target loss function according to the discriminator loss and the diffusion model loss:

[0049] L = min Dif max Dis (L dif + βL dis );

[0050] where L represents the target loss function; β represents the weight.

[0051] Among them, β is used to control the influence of the discriminator.

[0052] Among them, D(x) represents the output after the data x enters the discriminator, and its value is between 0 and 1. 0 means the discriminator judges that x is a false sample, and 1 means the discriminator judges that x is a real sample.

[0053] Among them, the discriminator loss is used to measure the accuracy of the discriminator's judgment on real (i.e., raster image text) and false (i.e., the artistic words output by the second diffusion model) samples. For real samples, it is expected that the discriminator outputs close to 1; for false samples, it is expected that the discriminator outputs close to 0. Overall, when the discriminator can accurately distinguish true and false samples, the discriminator loss approaches 0, indicating that the discriminator has strong discrimination ability, and the discriminator can guide the second diffusion model to generate glyphs so that they can be clearly recognized.

[0054] S11. Train the initial model based on the target loss function to obtain a text-to-artistic-word model.

[0055] In this embodiment, training the initial model based on the target loss function to obtain a text-to-artistic-word model includes:

[0056] Perform adversarial learning training on the second diffusion model using the discriminator;

[0057] During the training process, maximize the discriminator loss and minimize the diffusion model loss until the value of the diffusion model loss no longer decreases, then stop training to obtain the text-to-artistic-word model;

[0058] Among them, during the training process, when the output data of the discriminator for raster image text is 1, and / or the output data of the discriminator for the artistic words output by the second diffusion model is 0, give negative feedback to the second diffusion model; when the output data of the discriminator for raster image text is 0, and / or the output data of the discriminator for the artistic words output by the second diffusion model is 1, give positive feedback to the second diffusion model.

[0059] During training, the second diffusion model cannot output a very excellent artistic word at the beginning, and this requires the discriminator (such as a Gan (Generative Adversarial Networks) model) to correct it. For example, the original second diffusion model is a person who makes fakes. The learning goal of the discriminator is to distinguish genuine and fake products, and the goal of the second diffusion model is to make fakes that are indistinguishable from the real ones. In the continuous learning and confrontation between the two models, the capabilities of both models will continuously become stronger until the second diffusion model generates a very high quality output that the discriminator cannot distinguish. Specifically in this case, the target grid image text and the target artistic word are used as the real sample and the fake sample respectively, and the discriminator has to learn to distinguish them. After training is completed and during inference, the discriminator is not required to participate, that is, the discriminator itself does not participate in the generation of the artistic word, but is just a feedback model that makes the generation of the artistic word better.

[0060] It can be seen that during the entire training process, the models continuously confront each other, and the discriminator is continuously optimized to try to distinguish real and fake samples to the greatest extent. This will bring optimization pressure to the second diffusion model, making the generated images more realistic and in line with expectations; while the second diffusion model is continuously optimized to try to minimize the overall loss in order to generate artistic word images that can better "fool" the discriminator and make it difficult for the discriminator to distinguish between true and false.

[0061] This adversarial training process promotes the improvement of the model performance. During the iterative process, the model continuously adjusts its parameters so that the finally generated artistic font can reach a highly integrated and balanced state in terms of style and glyph, while maintaining a certain degree of diversity and high quality. Through this dynamic balance of the adversarial mechanism, the effect and performance of the model in generating artistic fonts are improved.

[0062] Among them, the second diffusion model is different from the first diffusion model, and the parameters of the second diffusion model can be trained.

[0063] S12, in response to an artistic word generation instruction based on the target text and the initial style prompt word, rasterize the target text based on the text-to-artistic-word model to obtain the target grid image text.

[0064] In this embodiment, the target text is the text that needs to be converted into an artistic word. For example: the target text can be "Happy Birthday".

[0065] In this embodiment, the initial style prompt word is used to limit the style when converting the target text into an artistic word. For example: the initial style prompt word can be "joyful".

[0066] In this embodiment, the artistic text generation instruction can be triggered by relevant staff according to actual needs. For example, it can be triggered by promoters in the financial field.

[0067] In this embodiment, based on the text-to-artistic-text model, rasterizing the target text can convert the vector text layer into a pixel layer. After rasterization, the text layer loses its vector characteristics and becomes a pixel layer, which means that the text is no longer editable, but various editing operations (such as applying filters and effects) can be performed on it like a normal image. This enables the text to be unrestricted by language types, avoiding the problem of complex Chinese language structures, and transforming the problem from text-to-image into image-to-image.

[0068] S13. Input the target raster image text and the initial style prompt into the large language model to obtain the coordinates of the target area for semanticization in the target raster image text, and the extended target style prompt.

[0069] The style prompts input by the user are often relatively simple and prone to unclear expression. Therefore, in this embodiment, the large language model is fine-tuned to obtain better style prompts.

[0070] Specifically, the method further includes:

[0071] Obtain fine-tuning samples; wherein, the fine-tuning samples include text, as well as marked area coordinates and style prompts;

[0072] Fine-tune the large language model based on the fine-tuning samples so that the output data of the large language model approaches the marked area coordinates and style prompts.

[0073] The large language model is a language model constructed by a deep neural network containing more than tens of billions of parameters. It is usually trained through a large amount of unlabeled text using self-supervised learning methods to predict and generate text and other content. It has powerful generation, migration, and interaction capabilities, and has the ability comparable to that of humans in many aspects such as language expression, instant conversation, task planning, and logical deduction. In order to achieve the semanticization of characters, the present invention utilizes the powerful capabilities of the large model, controls the input and output formats through fine-tuning, so that the large language model can use its understanding of semantics and rich knowledge reserve, combined with the analysis of images and style prompts, to exert creativity and imagination, outline the image parts suitable for semanticization, and output more powerful prompts for the subsequent diffusion model to use.

[0074] The main purpose of fine-tuning the large model is to solidify the input and output formats, increase its sensitivity to area selection, and adjust the input style prompts, and appropriately use the large model for rewriting to obtain better image generation effects.

[0075] For example, after fine-tuning, the following instructions can be input to the fine-tuned large model: "Given a style prompt and an image, where the image is a piece of text to be artistically transformed. According to the style prompt, give the coordinates of the area on the image where the text is most suitable for transformation to reflect the style prompt, with the y-axis pointing downwards and the x-axis pointing to the right. Then, combining the style prompt and the image text, give a prompt suitable for input into a text-to-art model to generate an artistic image of the text."

[0076] S14. Input the target style prompt into the first diffusion model to obtain a target style image.

[0077] In this embodiment, the first diffusion model can be a pre-trained general diffusion model.

[0078] In this embodiment, the target style image is a sample for text semanticization.

[0079] S15. Input the target raster image text, the target area coordinates, the target style prompt, and the target style image into the second diffusion model to obtain a target artistic word.

[0080] In this embodiment, inputting the target raster image text, the target area coordinates, the target style prompt, and the target style image into the second diffusion model to obtain a target artistic word includes:

[0081] Fuse the target raster image text and the target style image based on the target area coordinates to obtain a mixed image;

[0082] Input the mixed image into the noise adder of the second diffusion model through the image encoder of the second diffusion model to obtain first intermediate data doped with noise;

[0083] Process the target style prompt using the text encoder of the second diffusion model to obtain second intermediate data;

[0084] Input the first intermediate data and the second intermediate data into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word.

[0085] Specifically, inputting the first intermediate data and the second intermediate data into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word includes:

[0086] Based on the cross-attention mechanism, multiple QKV (Query Key Value) modules included in the U-net structure are used to interact and process the first intermediate data and the second intermediate data to obtain third intermediate data;

[0087] The decoder of the second diffusion model is used to decode the third intermediate data to obtain the target artistic word.

[0088] In the above embodiment, under the guidance of region selection of the target raster image text and the target style image generated by the first diffusion model, the corresponding region in the target raster image text is replaced with the target style image, thereby forming a mixed image that combines the original characters and style. Further, the mixed image enters the noise adder after passing through the image encoder to form data x doped with noise ε t (When adding noise, t can be used as the time step, representing the degree of noise addition. t is an integer with a value range of 1 - T, and T is generally set to 1000. During the training process, t in each sample follows a uniform distribution selection strategy to randomly take a value within the interval. During the inference process, it decreases by 1 from t = T until t = 1). The target style prompt word polished by the large language model enters the text encoder and then enters the U-net structure together with x t and enters the U-net structure together.

[0089] Among them, the U-net structure is a successful Convolutional Neural Networks (CNN) architecture, which gradually restores the resolution of the image through the transposed convolutional layers in the expansion path. In this process, by combining the previously extracted features (including low-level detail features and high-level semantic features), the pixel information damaged by the noise can be accurately filled. In the present invention, the U-net is based on x t and continuously incorporates the influence of the prompt word through cross-attention during the encoding and decoding process, embeds the target style prompt word to generate a key vector, embeds the text sequence into x t to generate a query vector and a value vector, so that the prompt word can effectively affect the output of the diffusion model.

[0090] In addition, different from the traditional diffusion model that can only accept text to output pictures or pictures to output pictures, the structure of the traditional diffusion model is improved in this embodiment. In order to be able to perform text semanticization in a specific region, the input part of the traditional diffusion model is modified to enable it to accept region coordinates, style images, text, and style prompt words.

[0091] In this embodiment, after obtaining the target artistic word, the method further includes:

[0092] Obtain the trigger of the artistic text generation instruction;

[0093] Send the target artistic text to the trigger for confirmation;

[0094] When receiving the confirmation signal of the target artistic text from the trigger, fill the target artistic text into the specified material to obtain the target material;

[0095] Send the target material to the target customer.

[0096] For example: In the digital marketing scenario of insurance agents in the financial field, agents can quickly use marketing materials, such as posters and other promotional content, to send to customers, providing greetings or product information to get closer to customers. The artistic text in the materials can be generated by the method in this embodiment, thus avoiding repeated design by designers, and effectively improving work efficiency by using the model to generate.

[0097] The following uses a specific example to elaborate on the method of generating artistic text from text in this embodiment.

[0098] For example: The text to generate artistic text is "Happy Birthday", and the style prompt is "joyful".

[0099] The text is rasterized to form raster image text I T Together with the style prompt, it enters the large language model, and the large language model will output a coordinate, such as [10, 20, 20, 30], which respectively represent the upper left coordinate and the lower left coordinate of a box. The large language model will also output an extended prompt prompt': "Happy, birthday cake, birthday candles, only generate a single image, and the background is white."

[0100] The extended prompt prompt' enters the first diffusion model diffusionmodel-A to generate a style image, such as generating a cake with several candles inserted, and no other things are generated with the background.

[0101] This style image will be combined with the previously generated raster image text I T , the extended prompt prompt', and the coordinate are input into the second diffusion model diffusion model-B for text semanticization.

[0102] First, the raster image text I T Under the guidance of the coordinate with the style image, I TThe corresponding area is replaced with the style image to form a hybrid image that combines the original characters and the style image. After being encoded by the image encoder "encoder" and adding noise through the noise adder "Diffusion process", the hybrid image becomes the data x. t It enters the U-net.

[0103] After being encoded by the text encoder "encoder", "prompt" also enters the U-net. It continuously affects x through cross-attention. t Finally, x t is output from the U-net and enters the decoder "Decoder" to be decoded into a picture, that is, the finally generated artistic words are obtained.

[0104] It can be seen that in the above embodiment, by rasterizing the text into image-to-image generation, the problem of the complex Chinese language structure is avoided, which helps to improve the quality of generating Chinese-related artistic words; by utilizing the powerful capabilities of the large language model, the semantics are understood and combined with image and style prompts for analysis, and a more controllable prompt is output for the diffusion model to use, improving the controllability of text art generation; with the help of the rich knowledge reserve and creative imagination of the large language model, abstract style prompts can be better processed, making the generated artistic words more in line with the expected style; a discriminator is added to guide the diffusion model to generate clearly recognizable glyphs, and at the same time, through the adversarial training mechanism, the generated artistic words have both clear glyph recognition and artistic semantic changes, achieving a high degree of integration and balance between style and glyph.

[0105] It can be seen from the above technical solutions that on the one hand, based on the text-to-artistic-words model, the target text is rasterized to obtain the target raster image text, and the vector text layer is converted into a pixel layer, making the text not restricted by the language type, avoiding the problem of the complex Chinese language structure, and converting the text-to-image problem into an image-to-image problem; on the other hand, the large language model is used to identify the regional coordinates for semanticization in the raster image text and expand the style prompts, improving the controllability of text art generation; on the other hand, the discriminator is used to guide the diffusion model to generate clearly recognizable glyphs, and at the same time, through the adversarial training mechanism, the generated artistic words have both clear glyph recognition and artistic semantic changes, achieving a high degree of integration and balance between style and glyph, thus solving the problems of low efficiency and poor quality in text-to-artistic-words generation.

[0106] Such as Figure 2As shown, it is a functional block diagram of a preferred embodiment of the text-to-artificial-artwork device based on artificial intelligence according to the present invention. The text-to-artificial-artwork device 11 based on artificial intelligence includes a construction unit 110, a training unit 111, a rasterization unit 112, and an input unit 113. The modules / units referred to in the present invention refer to a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0107] The construction unit 110 is configured to construct an initial model and a target loss function; wherein, the initial model includes a fine-tuned large language model (LLM), a discriminator, a first diffusion model, and a second diffusion model.

[0108] In this embodiment, the construction unit 110 constructing the target loss function includes:

[0109] Construct the discriminator loss using the following formula:

[0110]

[0111] where L dis represents the discriminator loss; D(I T ) represents the output data after inputting the raster image text I T into the discriminator; represents the output data after inputting the artwork output by the second diffusion model into the discriminator;

[0112] Construct the diffusion model loss using the following formula:

[0113]

[0114] where L dif represents the diffusion model loss; ε represents the original noise added to the second diffusion model; ε represents the predicted noise of the second diffusion model;

[0115] Construct the target loss function according to the discriminator loss and the diffusion model loss:

[0116] L = min Dif max Dis (L dif + βL dis );

[0117] Wherein, L represents the target loss function; β represents the weight.

[0118] Wherein, β is used to control the influence of the discriminator.

[0119] Wherein, D(x) represents the output after the data x enters the discriminator, and its value is between 0 and 1. 0 represents that the discriminator determines that x is a false sample, and 1 represents that the discriminator determines that x is a real sample.

[0120] Wherein, the discriminator loss is used to measure the accuracy of the discriminator's judgment on real (i.e., raster image text) and false (i.e., the artistic words output by the second diffusion model) samples. For real samples, it is expected that the discriminator outputs close to 1; for false samples, it is expected that the discriminator outputs close to 0. Overall, when the discriminator can accurately distinguish true and false samples, the discriminator loss approaches 0, indicating that the discriminator has strong discrimination ability, and the discriminator can guide the second diffusion model to generate glyphs so that they can be clearly recognized.

[0121] The training unit 111 is configured to train the initial model based on the target loss function to obtain a text-to-artistic-word model.

[0122] In this embodiment, the training unit 111 training the initial model based on the target loss function to obtain a text-to-artistic-word model includes:

[0123] Performing adversarial learning training on the second diffusion model using the discriminator;

[0124] During the training process, maximizing the discriminator loss and minimizing the diffusion model loss until the value of the diffusion model loss no longer decreases, then stopping the training to obtain the text-to-artistic-word model;

[0125] Wherein, during the training process, when the output data of the discriminator for raster image text is 1, and / or the output data of the discriminator for the artistic words output by the second diffusion model is 0, giving negative feedback to the second diffusion model; when the output data of the discriminator for raster image text is 0, and / or the output data of the discriminator for the artistic words output by the second diffusion model is 1, giving positive feedback to the second diffusion model.

[0126] During training, the second diffusion model cannot output a very excellent artistic word at the beginning, and this requires the discriminator (such as a Gan (Generative Adversarial Networks) model) to correct it. For example, the original second diffusion model is a person who makes fakes. The learning goal of the discriminator is to distinguish genuine products from fakes, and the goal of the second diffusion model is to make fakes that are indistinguishable from the real ones. In the continuous learning and confrontation between the two models, the capabilities of both models will continuously become stronger until the second diffusion model generates a very high quality and makes the discriminator unable to distinguish. Specifically in this case, the target grid image text and the target artistic word are used as real samples and fake samples respectively, and the discriminator has to learn to distinguish them. After the training is completed and during inference, the discriminator is not required to participate, that is, the discriminator itself does not participate in the generation of the artistic word, but is only a feedback model that makes the generation of the artistic word better.

[0127] It can be seen that the models continuously confront each other throughout the training process, and the discriminator is continuously optimized, trying to distinguish real and fake samples to the greatest extent. This will bring optimization pressure to the second diffusion model, making the generated images more realistic and in line with expectations; while the second diffusion model is continuously optimized, trying to minimize the overall loss to generate artistic word images that can better "deceive" the discriminator and make it difficult for the discriminator to distinguish between true and false.

[0128] This adversarial training process promotes the improvement of the model performance. During the iteration process, the model continuously adjusts its parameters so that the finally generated artistic font can reach a highly integrated and balanced state in terms of style and glyph, while maintaining a certain degree of diversity and high quality. Through this dynamic balance of the adversarial mechanism, the effect and performance of the model in generating artistic fonts are improved.

[0129] Among them, the second diffusion model is different from the first diffusion model, and the parameters of the second diffusion model are trainable.

[0130] The rasterization unit 112 is configured to, in response to an artistic word generation instruction based on the target text and the initial style prompt word, rasterize the target text based on the text-to-artistic-word model to obtain target grid image text.

[0131] In this embodiment, the target text is the text that needs to be converted into an artistic word. For example: the target text can be "Happy Birthday".

[0132] In this embodiment, the initial style prompt word is used to limit the style when converting the target text into an artistic word. For example: the initial style prompt word can be "joyful".

[0133] In this embodiment, the artistic text generation instruction can be triggered by relevant staff according to actual needs. For example, it can be triggered by promoters in the financial field.

[0134] In this embodiment, based on the text-to-artistic-text model, rasterization processing is performed on the target text, which can convert the vector text layer into a pixel layer. After rasterization, the text layer loses its vector characteristics and becomes a pixel layer. This means that the text is no longer editable, but various editing operations (such as applying filters and effects) can be performed on it like ordinary images. This enables the text to be unrestricted by language types and can avoid the problem of complex Chinese language structures, converting the problem from text-to-image to image-to-image.

[0135] The input unit 113 is configured to input the target raster image text and the initial style prompt into the large language model to obtain the target region coordinates for semanticization in the target raster image text and the extended target style prompt.

[0136] The style prompts input by the user themselves are often relatively simple and prone to unclear expression. Therefore, in this embodiment, the large language model is fine-tuned to obtain better style prompts.

[0137] Specifically, fine-tuning samples are obtained; wherein, the fine-tuning samples include text, as well as marked region coordinates and style prompts.

[0138] Based on the fine-tuning samples, the large language model is fine-tuned to make the output data of the large language model approach the marked region coordinates and style prompts.

[0139] The large language model is a language model constructed by a deep neural network containing more than tens of billions of parameters. It is usually trained through a large amount of unlabeled text using self-supervised learning methods to predict and generate text and other content. It has powerful generation, transfer, and interaction capabilities and has the ability comparable to humans in many aspects such as language expression, instant conversation, task planning, and logical deduction. In order to achieve the semanticization of characters, the present invention utilizes the powerful capabilities of the large model and controls the input-output format through fine-tuning, so that the large language model can use its understanding of semantics and rich knowledge reserve, combined with the analysis of images and style prompts, to exert creativity and imagination, circle the image parts suitable for semanticization, and output more powerful prompts for the subsequent diffusion model to use.

[0140] The main purpose of fine-tuning the large model is to solidify the input-output format, increase its sensitivity to region selection, and adjust the input style prompts, and appropriately use the large model for rewriting to obtain better image generation effects.

[0141] For example: After fine-tuning, the following instructions can be input to the fine-tuned large model: "Given a style prompt and an image, where the image is a piece of text to be artistically transformed. Please, based on the style prompt, give the coordinates of the area on the image where the text is most suitable for transformation to reflect the style prompt, with the y-axis pointing downwards and the x-axis pointing to the right. Then, combining the style prompt and the image text, give a prompt suitable for input into a text-to-art model to generate an artistic image of the text."

[0142] The input unit 113 is further configured to input the target style prompt into the first diffusion model to obtain a target style image.

[0143] In this embodiment, the first diffusion model can be a pre-trained general diffusion model.

[0144] In this embodiment, the target style image is a sample for text semanticization.

[0145] The input unit 113 is further configured to input the target raster image text, the target area coordinates, the target style prompt, and the target style image into the second diffusion model to obtain a target artistic word.

[0146] In this embodiment, the input unit 113 inputs the target raster image text, the target area coordinates, the target style prompt, and the target style image into the second diffusion model to obtain a target artistic word, including:

[0147] Fusing the target raster image text and the target style image based on the target area coordinates to obtain a mixed image;

[0148] Feeding the mixed image through the image encoder of the second diffusion model into the noise adder of the second diffusion model to obtain first intermediate data doped with noise;

[0149] Processing the target style prompt using the text encoder of the second diffusion model to obtain second intermediate data;

[0150] Inputting the first intermediate data and the second intermediate data into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word.

[0151] Specifically, inputting the first intermediate data and the second intermediate data into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word includes:

[0152] Based on the cross-attention mechanism, multiple QKV (Query Key Value) modules included in the U-net structure are used to interact and process the first intermediate data and the second intermediate data to obtain third intermediate data;

[0153] The decoder of the second diffusion model is used to decode the third intermediate data to obtain the target artistic word.

[0154] In the above embodiment, under the guidance of region selection, the corresponding region in the target grid image text is replaced with the target style image generated by the first diffusion model, so as to form a hybrid image that combines the original characters and style. Further, the hybrid image enters the noise adder through the image encoder to form the data x doped with noise ε t (When adding noise, t can be used as the time step, representing the degree of noise addition. t is an integer with a value range of 1 - T, and T is generally set to 1000. During the training process, t in each sample follows a uniform distribution selection strategy and randomly takes a value within the interval. During the inference process, t decreases by 1 each time from t = T until t = 1). The target style prompt word polished by the large language model enters the text encoder and then enters the U-net structure together with x t and enters the U-net structure together.

[0155] Among them, the U-net structure is a successful Convolutional Neural Networks (CNN) architecture, which gradually restores the resolution of the image through the transposed convolutional layers in the expansion path. In this process, by combining the previously extracted features (including low-level detail features and high-level semantic features), the pixel information damaged by noise can be accurately filled. In the present invention, the U-net is based on x t and continuously incorporates the influence of the prompt word through cross-attention during the encoding and decoding process, embeds the target style prompt word to generate the key vector, embeds the text sequence into x t to generate the query vector and the value vector, so that the prompt word can effectively affect the output of the diffusion model.

[0156] In addition, different from the traditional diffusion model that can only accept text to output pictures or pictures to output pictures, the structure of the traditional diffusion model is improved in this embodiment. In order to be able to perform text semanticization in a specific region, the input part of the traditional diffusion model is modified to enable it to accept region coordinates, style images, text, and style prompt words.

[0157] In this embodiment, after obtaining the target artistic word, the trigger of the artistic word generation instruction is obtained;

[0158] Send the target artistic text to the trigger for confirmation;

[0159] When receiving the confirmation signal of the target artistic text from the trigger, fill the target artistic text into the specified material to obtain the target material;

[0160] Send the target material to the target customer.

[0161] For example: In the digital marketing scenario of insurance agents in the financial field, agents can quickly use marketing materials, such as posters and other promotional content, to send to customers, provide greetings or product information to get closer to customers. The artistic text in the materials can be generated by the method in this embodiment, thus avoiding designers from repeating the design, and effectively improving work efficiency by using the model to generate.

[0162] The following uses a specific example to elaborate on the method of generating artistic text from text in this embodiment.

[0163] For example: The text to generate artistic text is "Happy Birthday", and the style prompt is "joyful".

[0164] The text is rasterized to form raster image text I T Enter into the large language model together with the style prompt. The large language model will output a coordinate, such as [10, 20, 20, 30], which respectively represent the upper left coordinate and the lower left coordinate of a box. The large language model will also output an extended prompt prompt': "Happy, birthday cake, birthday candles, only generate a single image, and the background is white."

[0165] The extended prompt prompt' enters the first diffusion model diffusionmodel-A to generate a style image, such as generating a cake with several candles inserted, and no other things are generated with the background.

[0166] This style image will be combined with the previously generated raster image text I T , the extended prompt prompt', and the coordinate are input into the second diffusion model diffusion model-B for text semanticization.

[0167] First, the raster image text I T Under the guidance of the coordinate with the style image, replace the corresponding area of I T with the style image to form a mixed image that combines the original characters and the style image. After the mixed image is encoded by the image encoder encoder and noise is added by the noise adder Diffusion process, it becomes data x t and enters the U-net.

[0168] The 'prompt' also enters the U-net after being encoded by the text encoder. It continuously affects x through cross-attention t , and finally x t is output from the U-net and enters the decoder Decoder to be decoded into an image, that is, the finally generated artistic text is obtained.

[0169] It can be seen that in the above embodiments, by rasterizing the text into image-to-image generation, the problem of complex Chinese language structure is avoided, which helps to improve the quality of generating Chinese-related artistic text; by utilizing the powerful capabilities of the large language model to understand semantics and combining image and style prompts for analysis, more controllable prompts are output for the diffusion model to use, improving the controllability of text art generation; with the help of the rich knowledge reserve and creative imagination of the large language model, abstract style prompts can be better processed, making the generated artistic text more in line with the expected style; the discriminator is added to guide the diffusion model to generate clearly recognizable glyphs, and at the same time, through the adversarial training mechanism, the generated artistic text has both clear glyph recognition and artistic semantic changes, achieving a high degree of integration and balance between style and glyph.

[0170] It can be seen from the above technical solutions that on the one hand, based on the text-to-artistic-text model, rasterization processing is performed on the target text to obtain the target raster image text, and the vector text layer is converted into a pixel layer, making the text not restricted by language types, avoiding the problem of complex Chinese language structure, and converting the text-to-image problem into an image-to-image problem; on the other hand, the large language model is used to identify the regional coordinates for semanticization in the raster image text and expand the style prompts, improving the controllability of text art generation; on the other hand, the discriminator is used to guide the diffusion model to generate clearly recognizable glyphs, and at the same time, through the adversarial training mechanism, the generated artistic text has both clear glyph recognition and artistic semantic changes, achieving a high degree of integration and balance between style and glyph, thus solving the problems of low efficiency and poor quality in text-to-artistic-text.

[0171] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment for implementing the text-to-artistic-text method based on artificial intelligence of the present invention.

[0172] The computer device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a text-to-artistic-text program based on artificial intelligence.

[0173] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the computer device 1 can also include input and output devices, network access devices, etc.

[0174] It should be noted that the computer device 1 is only an example. Other existing or future possible electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.

[0175] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 12 can be an internal storage unit of the computer device 1 in some embodiments. For example, the mobile hard disk of the computer device 1. The memory 12 can also be an external storage device of the computer device 1 in other embodiments. For example, the plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can be used not only to store the application software installed on the computer device 1 and various types of data, such as the code of the text-to-art program based on artificial intelligence, etc., but also to temporarily store the data that has been output or will be output.

[0176] The processor 13 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting all the components of the entire computer device 1 through various interfaces and lines. By running or executing the programs or modules stored in the memory 12 (such as executing the text-to-art program based on artificial intelligence), and calling the data stored in the memory 12, it can execute various functions of the computer device 1 and process data.

[0177] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of the method for generating artistic words based on artificial intelligence, such as Figure 1 the steps shown.

[0178] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program can be divided into a construction unit 110, a training unit 111, a rasterization unit 112, and an input unit 113.

[0179] The above-mentioned integrated units implemented in the form of software function modules can be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the method for generating artistic words based on artificial intelligence described in the various embodiments of the present invention.

[0180] If the integrated module / unit of the computer device 1 is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented.

[0181] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, etc.

[0182] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store the operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.

[0183] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0184] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to achieve the connection and communication between the memory 12 and at least one processor 13, etc.

[0185] Although not shown, the computer device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may also include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The computer device 1 may further include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0186] Furthermore, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.

[0187] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.

[0188] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0189] Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0190] In combination with Figure 1 , the memory 12 in the computer device 1 stores multiple instructions to implement a method for generating artistic words from text based on artificial intelligence, and the processor 13 can execute the multiple instructions to implement:

[0191] Construct an initial model and a target loss function; wherein, the initial model includes a fine-tuned large language model, a discriminator, a first diffusion model, and a second diffusion model;

[0192] Train the initial model based on the target loss function to obtain an artistic word generation model from text;

[0193] In response to an artistic word generation instruction based on a target text and an initial style prompt, rasterize the target text based on the artistic word generation model from text to obtain target raster image text;

[0194] Input the target raster image text and the initial style prompt into the large language model to obtain the target region coordinates for semanticization in the target raster image text, and an extended target style prompt;

[0195] Input the target style prompt into the first diffusion model to obtain a target style image;

[0196] Input the target raster image text, the target region coordinates, the target style prompt, and the target style image into the second diffusion model to obtain target artistic words.

[0197] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.

[0198] It should be noted that all the data involved in this case are legally obtained. The non-company software tools or components that appear in the embodiments of this application are only for illustrative introduction and do not represent actual use.

[0199] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0200] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0201] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0202] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0203] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0204] Therefore, in all aspects, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0205] In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices described in the present invention can also be implemented by one element or device through software or hardware. The terms such as "first" and "second" are used to denote names and do not denote any particular order.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for writing artistic characters based on artificial intelligence, characterized in that: The artificial intelligence-based Wensheng artistic word method comprises: Constructing an initial model and a target loss function; wherein the initial model includes a fine-tuned large language model, a discriminator, a first diffusion model, and a second diffusion model; Training the initial model based on the target loss function to obtain a Wensheng art character model; In response to an artistic word generation instruction based on a target text and an initial style prompt word, rasterizing the target text based on the Wensheng artistic word model to obtain a target raster image text; Inputting the target grid image text and the initial style prompt word into the large language model to obtain the target area coordinates for semanticization in the target grid image text and the expanded target style prompt word; Inputting the target style prompt word into the first diffusion model to obtain a target style image; The target grid image text, the target area coordinates, the target style prompt word and the target style image are input into the second diffusion model to obtain the target artistic word.

2. The method for producing artistic characters based on artificial intelligence as claimed in claim 1, characterized in that: The constructing target loss function comprises: The discriminator loss is constructed using the following formula: Among them, L dis Denotes the discriminator loss; D(I T ) means to convert the raster image text I T Output data after input into the discriminator; The artistic word representing the output of the second diffusion model Output data after input into the discriminator; The diffusion model loss is constructed using the following formula: Among them, L dif represents the diffusion model loss; ε represents the original noise added to the second diffusion model; ε represents the predicted noise of the second diffusion model; The objective loss function is constructed according to the discriminator loss and the diffusion model loss: L=min Dif max Dis (L dif +βL dis ); Wherein, L represents the target loss function; β represents the weight.

3. The method for producing artistic characters based on artificial intelligence as claimed in claim 1, characterized in that: The method further comprises: Obtaining a fine-tuning sample; wherein the fine-tuning sample includes text, marked area coordinates, and style prompt words; The large language model is fine-tuned based on the fine-tuning samples so that output data of the large language model approaches the marked region coordinates and style prompt words.

4. The method for producing artistic characters based on artificial intelligence as claimed in claim 2, characterized in that: The training of the initial model based on the target loss function to obtain the Wensheng art character model comprises: Using the discriminator to perform adversarial learning training on the second diffusion model; During the training process, the discriminator loss is maximized and the diffusion model loss is minimized until the value of the diffusion model loss no longer decreases, and the training is stopped to obtain the Wensheng art character model; Wherein, during the training process, when the output data of the discriminator for the raster image text is 1, and / or the output data of the artistic characters output by the second diffusion model is 0, negative feedback is given to the second diffusion model; when the output data of the discriminator for the raster image text is 0, and / or the output data of the artistic characters output by the second diffusion model is 1, positive feedback is given to the second diffusion model.

5. The method for producing artistic characters based on artificial intelligence as claimed in claim 1, characterized in that: The step of inputting the target grid image text, the target area coordinates, the target style prompt word and the target style image into the second diffusion model to obtain the target artistic word comprises: Based on the target area coordinates, the target grid image text is merged with the target style image to obtain a mixed image; Passing the mixed image through the image encoder of the second diffusion model and entering the noise adder of the second diffusion model to obtain first intermediate data doped with noise; Processing the target style prompt word using the text encoder of the second diffusion model to obtain second intermediate data; The first intermediate data and the second intermediate data are input into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word.

6. The method for producing artistic characters based on artificial intelligence as claimed in claim 5, characterized in that: The step of inputting the first intermediate data and the second intermediate data into the U-net structure of the second diffusion model for diffusion processing to obtain the target artistic word comprises: Based on the cross attention mechanism, the first intermediate data and the second intermediate data are interactively processed by using the multiple QKV modules included in the U-net structure to obtain third intermediate data; The third intermediate data is decoded by using the decoder of the second diffusion model to obtain the target artistic word.

7. The method for producing artistic characters based on artificial intelligence as claimed in claim 1, characterized in that: After obtaining the target artistic word, the method further includes: Obtaining the trigger of the artistic word generation instruction; Sending the target artistic word to the triggerer for confirmation; When receiving a confirmation signal of the triggerer for the target artistic word, filling the target artistic word into the designated material to obtain the target material; The target material is sent to the target customer.

8. An artificial intelligence-based Wensheng art character device, characterized in that: The Wensheng art character device based on artificial intelligence includes: A construction unit, used to construct an initial model and a target loss function; wherein the initial model includes a fine-tuned large language model, a discriminator, a first diffusion model, and a second diffusion model; A training unit, used for training the initial model based on the target loss function to obtain a Wensheng art character model; A rasterization unit, for responding to an artistic character generation instruction based on a target text and an initial style prompt word, and performing rasterization processing on the target text based on the Wensheng artistic character model to obtain a target raster image text; An input unit, used for inputting the target grid image text and the initial style prompt word into the large language model, and obtaining the target area coordinates for semanticization in the target grid image text, and the expanded target style prompt word; The input unit is further used to input the target style prompt word into the first diffusion model to obtain a target style image; The input unit is further used to input the target grid image text, the target area coordinates, the target style prompt word and the target style image into the second diffusion model to obtain a target artistic word.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes the instructions stored in the memory to implement the artificial intelligence-based Wensheng artistic character method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the artificial intelligence-based Wensheng artistic character method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • 3D wordart generation method based on structure-view angle double-stage diffusion model

    CN121414972A

  • Anti-counterfeiting tracing method and system based on font steganography coding and enterprise production center data center

    CN121544276A