Image-text generation method and device, electronic equipment and storage medium

By encoding and decoding image and text features with convolutional attention and segmentation, the method addresses the challenge of capturing key regions in complex backgrounds, improving the accuracy of text layout generation.

CN120318375APending Publication Date: 2025-07-15SHENZHEN PINGAN COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495788.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing graphic and text generation methods are difficult to accurately capture key areas in the image when dealing with complex backgrounds, resulting in the generated text layout being unreasonable and low accuracy.

Method used

By acquiring the target image and copy features, feature fusion, convolutional attention prediction and segmentation encoding are performed, copy layout description text is generated, and target graphics and text are obtained through segmentation and decoding, improving the capture accuracy of key areas.

Benefits of technology

It improves the accuracy of graphic and text generation, ensures the coordination and rationality of text layout and image content, and reduces the adverse effects of complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318375A_ABST
    Figure CN120318375A_ABST
Patent Text Reader

Abstract

The invention provides an image-text generation method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for the field of financial science and technology. The method comprises the following steps: acquiring a target image, and carrying out image coding on the target image to obtain target image features; obtaining a target copywriting, and carrying out text coding on the target copywriting to obtain a target copywriting feature; performing feature fusion on the target image features and the target copywriting features to obtain image-text fusion features; performing convolution attention prediction on the image-text fusion features to obtain a copywriting layout description text; encoding the copywriting layout description text and the target copywriting to obtain copywriting layout features; performing segmentation coding on the copywriting layout features and the target image features to obtain image-text layout coding features; and segmenting and decoding the copywriting layout features and the image-text layout coding features to obtain a target image-text. The method can reduce the adverse effect caused by the complex background of the image, and improves the image-text generation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology and is applicable to the field of fintech, and particularly relates to a method and device for generating text and images, an electronic device, and a storage medium. Background Art

[0002] The method for generating text and images is a technology for generating a reasonable text layout in an image. It is an important task in the fields of computer vision and image processing, and can be applied to scenarios such as advertising design, poster generation, and image annotation in the field of fintech. For example, in the poster generation scenario, a user expects to generate a promotional poster introducing new energy vehicles. At this time, the text input by the user includes the performance information of the new energy vehicle (such as: a large battery life of 900 kilometers, a full vehicle intelligent cockpit), and the input image includes the external view, interior view, and vehicle model view of the new energy vehicle, etc. Then, through a text and image generation model, the input text and image are used for poster generation to obtain a promotional poster introducing the new energy vehicle.

[0003] Existing methods for generating text and images have made certain progress in the task of generating text and images. However, when dealing with the complex background of an image (such as an image that may have objects with various colors, textures, patterns, or shapes, and these objects may be confused with the key areas of the foreground), it is often difficult to accurately capture the key areas in the image, resulting in an unreasonable text layout generated in the image, that is, the accuracy of generating text and images is relatively low.

[0004] Therefore, there is a problem of relatively low accuracy in generating text and images in the related art. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a method and device for generating text and images, an electronic device, and a storage medium, which can reduce the adverse effects brought by the complex background of the image and improve the accuracy of generating text and images.

[0006] To achieve the above object, a first aspect of the embodiments of the present application proposes a method for generating text and images, the method comprising:

[0007] Obtain a target image, and perform image encoding on the target image to obtain target image features;

[0008] Obtain a target text, and perform text encoding on the target text to obtain target text features;

[0009] Perform feature fusion on the target image features and the target text features to obtain text and image fusion features;

[0010] Perform convolutional attention prediction on the text and image fusion features to obtain a text description of the layout of the text;

[0011] Encode the copy layout description text and the target copy to obtain copy layout features;

[0012] Segment and encode the copy layout features and the target image features to obtain text-image layout encoding features;

[0013] Segment and decode the copy layout features and the text-image layout encoding features to obtain the target text-image.

[0014] Optionally, the convolutional attention prediction of the text-image fusion features to obtain the copy layout description text includes:

[0015] Perform self-attention calculation on the text-image fusion features to obtain text-image attention features;

[0016] Perform convolutional attention calculation on the text-image attention features to obtain text-image layout perception features;

[0017] Perform layout prediction on the text-image layout perception features to obtain the copy layout description text.

[0018] Optionally, the convolutional attention calculation on the text-image attention features to obtain text-image layout perception features includes:

[0019] Perform convolutional processing on the text-image attention features to obtain text-image convolutional features;

[0020] Perform weight calculation on the text-image convolutional features to obtain the attention weights at each spatial position in the text-image attention features;

[0021] Perform weighted summation on each spatial position in the text-image attention features according to the attention weights to obtain the text-image layout perception features.

[0022] Optionally, the segmenting and encoding of the copy layout features and the target image features to obtain text-image layout encoding features includes:

[0023] Multiply a preset query weight matrix by the target image features to obtain an image layout query vector;

[0024] Multiply a preset key weight matrix by the copy layout features to obtain a text layout key vector;

[0025] Multiply a preset value weight matrix by the copy layout features to obtain a text layout value vector;

[0026] Perform weight calculation on the image layout query vector and the text layout key vector to obtain text-image layout attention scores;

[0027] Perform non-linear activation processing on the graphic layout attention score to obtain a graphic layout weight vector;

[0028] Perform weighted summation on the graphic layout weight vector and the text layout value vector to obtain the graphic layout encoded feature.

[0029] Optionally, the feature fusion of the target image feature and the target copywriting feature to obtain a graphic-text fusion feature includes:

[0030] Perform dimensionality expansion on the target copywriting feature to make the spatial dimension of the target copywriting feature the same as the spatial dimension of the image feature;

[0031] Perform element-wise addition on the target image feature and the target copywriting feature to obtain the graphic-text fusion feature.

[0032] Optionally, the feature fusion of the target image feature and the target copywriting feature to obtain a graphic-text fusion feature includes: performing feature fusion on the target image feature and the target copywriting feature through a fusion layer of a preset layout extractor to obtain a graphic-text fusion feature; the performing convolutional attention prediction on the graphic-text fusion feature to obtain a copywriting layout description text includes: performing convolutional attention prediction on the graphic-text fusion feature through a layout prediction layer of the layout extractor to obtain a copywriting layout description text; wherein, before performing feature fusion on the target image feature and the target copywriting feature through the fusion layer of the preset layout extractor to obtain a graphic-text fusion feature, the method further includes:

[0033] Pre-train the layout extractor, specifically including:

[0034] Obtain a sample masked image, and perform image restoration on the sample masked image through a preset image restoration model to obtain a sample restored image;

[0035] Perform image encoding on the sample restored image to obtain a sample image feature;

[0036] Obtain a sample copywriting, and perform text encoding on the sample copywriting to obtain a sample copywriting feature;

[0037] Perform layout information extraction on the sample image feature and the sample copywriting feature through a preset initial layout extractor to obtain a sample copywriting layout description text; wherein, the initial layout extractor includes a fusion layer and a layout prediction layer;

[0038] Calculate a loss according to the sample copywriting layout description text and a preset labeled copywriting layout description text to obtain layout matching loss data;

[0039] Train the initial layout extractor according to the layout matching loss data to obtain the layout extractor.

[0040] Optionally, the sample text layout description text includes sample text layout prediction positions; the segmenting and encoding the copy layout features and the target image features to obtain the graphic-text layout encoded features includes: segmenting and encoding the copy layout features and the target image features through a segmentation encoder of a preset segmentation network to obtain the graphic-text layout encoded features; the segmenting and decoding the copy layout features and the graphic-text layout encoded features to obtain the target graphic-text includes: segmenting and decoding the copy layout features and the graphic-text layout encoded features through a segmentation decoder of the segmentation network to obtain the target graphic-text; wherein, before the segmenting and encoding the copy layout features and the target image features through the segmentation encoder of the preset segmentation network to obtain the graphic-text layout encoded features, the method further includes:

[0041] Pre-training the segmentation network specifically includes:

[0042] Performing text encoding on the sample text layout description text and the sample text to obtain sample text layout features;

[0043] Applying preset sample noise to the sample image features to obtain sample noise-added image features;

[0044] Segmenting and encoding the sample text layout features and the sample noise-added image features through a segmentation encoder of a preset initial segmentation network to obtain sample noise-added graphic-text layout encoded features;

[0045] Segmenting and decoding the sample text layout features and the sample noise-added graphic-text layout encoded features through a segmentation decoder of the initial segmentation network to obtain sample noise-added graphic-text and predicted noise;

[0046] Calculating a loss according to the sample noise and the predicted noise to obtain noise matching loss data;

[0047] Removing the noise from the sample noise-added graphic-text according to the predicted noise to obtain sample noise-removed graphic-text;

[0048] Performing image recognition on the sample noise-removed graphic-text to obtain sample text layout recognition positions;

[0049] Calculating a loss according to the sample text layout prediction positions and the sample text layout recognition positions to obtain text position control loss data;

[0050] Training the initial segmentation network according to the noise matching loss data and the text position control loss data to obtain the segmentation network.

[0051] To achieve the above object, a second aspect of the embodiments of the present application provides a graphic and text generation device, the device comprising:

[0052] An image encoding module, configured to obtain a target image and perform image encoding on the target image to obtain target image features;

[0053] A text encoding module, configured to obtain a target text and perform text encoding on the target text to obtain target text features;

[0054] A feature fusion module, configured to perform feature fusion on the target image features and the target text features to obtain graphic and text fusion features;

[0055] A layout prediction module, configured to perform convolutional attention prediction on the graphic and text fusion features to obtain a text layout description text;

[0056] A fusion encoding module, configured to encode the text layout description text and the target text to obtain text layout features;

[0057] A segmentation encoding module, configured to perform segmentation encoding on the text layout features and the target image features to obtain graphic and text layout encoding features;

[0058] A segmentation decoding module, configured to perform segmentation decoding on the text layout features and the graphic and text layout encoding features to obtain a target graphic and text.

[0059] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the graphic and text generation method described in the first aspect when executing the computer program.

[0060] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, the storage medium being a computer-readable storage medium, the storage medium storing a computer program, and the computer program implementing the graphic and text generation method described in the first aspect when executed by a processor.

[0061] The text and image generation method, device, electronic device, and storage medium proposed in this application first obtain a target image and target text. The target image is respectively subjected to image encoding to obtain target image features, and the target text is subjected to text encoding to obtain target text features. Further, the image features and text features are subjected to feature fusion, and then layout prediction is performed to obtain a text layout description text. This text layout description text can indicate the approximate layout position of the target text in the target image. Further, the text layout description text and the target text are subjected to text encoding to obtain text layout features. These text layout features can not only more accurately indicate the layout position of the target text in the target image but also indicate the text content at this layout position. Further, the text layout features and the target image features are subjected to segmentation encoding to obtain text and image layout encoding features. In addition to indicating the layout position of the text in the target image and the text content at this layout position, these text and image layout encoding features also indicate the key area information of the target image. Further, through a segmentation decoder, the text layout features and the text and image layout encoding features are subjected to segmentation decoding to obtain the target text and image. In summary, in this application, first, a text layout description text is extracted based on the target image features and the target text features. Secondly, joint text encoding with the target text can obtain text layout features with relatively rich text layout information. Further, the text layout features and the target image features are subjected to segmentation encoding, and the text layout features and the text and image layout encoding features are subjected to segmentation decoding, that is, there is a same input feature (i.e., the text layout features) for both the segmentation encoder and the segmentation decoder. This can achieve effective control of the text layout and improve the accuracy of capturing key areas in the image, thereby improving the rationality of the text layout generated in the image. In conclusion, this application can reduce the adverse effects brought by the complex background of the image and improve the accuracy of text and image generation.

[0062] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. Brief Description of the Drawings

[0063] Figure 1 is a flowchart of the text and image generation method provided by an embodiment of this application;

[0064] Figure 2 is Figure 1 the flowchart of step 104 in

[0065] Figure 3 is Figure 2 the flowchart of step 202 in

[0066] Figure 4 is Figure 1 the flowchart of step 106 in

[0067] Figure 5 is a flowchart of a graphic and text generation method provided by another embodiment of the present application;

[0068] Figure 6 is a flowchart of a graphic and text generation method provided by another embodiment of the present application;

[0069] Figure 7 is a block diagram of the module structure of a graphic and text generation device provided by an embodiment of the present application;

[0070] Figure 8 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0071] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.

[0072] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0074] First, several nouns involved in the present application are analyzed:

[0075] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0076] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information image processing, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing, etc.

[0077] Traditional image-text generation methods usually rely on manually designed rules or heuristic algorithms. These methods often perform poorly when dealing with complex backgrounds or diverse images and are difficult to generate both beautiful and readable image-text. With the development of deep learning technology, neural network-based image-text generation methods have gradually become mainstream. These methods can automatically learn layout rules from data and to a certain extent improve the quality of generating text layouts in images.

[0078] Currently, some studies have attempted to introduce the attention mechanism into the image-text generation task. For example, the spatial relationships in the image can be captured through the attention mechanism, thereby improving the generation quality of the text layout. Another example is to use the self-attention mechanism to capture long-range dependency relationships and generate coherent and beautiful text layouts. However, although the existing methods have made some progress in the image-text generation task, they still face the following challenges: First, when dealing with complex backgrounds, the existing methods often have difficulty accurately capturing the key regions in the image, resulting in an unreasonable text layout. Second, when capturing spatial relationships, the existing methods usually rely on the global attention mechanism, which may cause the model to perform poorly when dealing with local details. Finally, when generating text layouts, the existing methods often lack sufficient utilization of context information, resulting in a lack of coordination between the generated layout and the image content.

[0079] Based on this, the embodiments of this application propose an image-text generation method, an image-text generation device, an electronic device, and a computer-readable storage medium, which can reduce the adverse effects brought by the complex background of the image and improve the accuracy of image-text generation.

[0080] The image text layout generation method provided by the embodiments of the present application can be applied to terminals and servers, or can be software running on the server side. The server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the image text layout generation method, etc., but is not limited to the above forms.

[0081] The present application can be used in many general or special computer system environments or configurations. For example: server computers, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0082] The embodiments of the present application provide an image text layout generation method, an image text layout generation device, an electronic device, and a computer-readable storage medium, which will be specifically described through the following embodiments. First, the image text layout generation method in the embodiments of the present application will be described.

[0083] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing that needs to be based on data related to the user's identity or characteristics, such as the user's image data, the user's permission or consent will be obtained first, and moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards.

[0084] Refer to Figure 1 , Figure 1 is an optional flowchart of the graphic text generation method provided by the embodiments of the present application, which may include but is not limited to steps 101 to 107.

[0085] Step 101, obtain a target image, and perform image encoding on the target image to obtain target image features;

[0086] Step 102, obtain a target copywriting, and perform text encoding on the target copywriting to obtain target copywriting features;

[0087] Step 103: Perform feature fusion on the target image features and the target copywriting features to obtain text-image fusion features;

[0088] Step 104: Perform convolutional attention prediction on the text-image fusion features to obtain a copywriting layout description text;

[0089] Step 105: Encode the copywriting layout description text and the target copywriting to obtain copywriting layout features;

[0090] Step 106: Perform segmentation encoding on the copywriting layout features and the target image features to obtain text-image layout encoding features;

[0091] Step 107: Perform segmentation decoding on the copywriting layout features and the text-image layout encoding features to obtain the target text-image.

[0092] Steps 101 to 107 illustrated in the embodiments of the present application first extract a copywriting layout description text based on the target image features and the target copywriting features. Secondly, encoding in combination with the target copywriting can obtain copywriting layout features with relatively rich text layout information. Further, perform segmentation encoding on the copywriting layout features and the target image features, and then perform segmentation decoding on the copywriting layout features and the text-image layout encoding features. That is, there is a same input feature (i.e., the copywriting layout feature) for both the segmentation encoder and the segmentation decoder. This can effectively control the copywriting layout and improve the accuracy of capturing key regions in the image, thereby improving the rationality of the text layout generated in the image. In summary, the present application can reduce the adverse effects brought by the complex background of the image and improve the accuracy of text-image generation.

[0093] In one example, in the advertising design scenario in the fintech field, a user may expect to generate a printed advertisement (such as a magazine, newspaper). At this time, the target copywriting input by the user includes news text content related to insurance products (for example: an insurance company launches an insurance product with no age limit for insurance) and the target image input includes insurance product illustrations, insurance product element pictures, insurance icons, etc.

[0094] In another example, in the poster generation scenario in the fintech field, a user expects to generate a promotional poster introducing new energy vehicles. At this time, the target copywriting input by the user includes performance information of new energy vehicles (for example: a large cruising range of 900 kilometers, a full vehicle intelligent cockpit) and the target image input includes exterior views, interior views, and car model pictures of new energy vehicles, etc.

[0095] In step 101 of some embodiments, a target image is obtained, and the target image is encoded to obtain target image features. The target image refers to the image material used to generate the picture and text. The target image can be encoded by an image encoder to obtain target image features. The image encoder can adopt the ResNet-50 architecture. The main purpose is to extract high-level visual features from the input target image. These visual features capture the global and local information of the image, providing a basis for subsequent text layout generation. ResNet-50 solves the problem of gradient disappearance in deep networks through residual connections and can effectively capture the global and local information of the image. For example, after the target image is processed by ResNet-50, the target image features are obtained, which can be denoted as feature map F. The size of feature map F is H×W×C, where H and W are the height and width of feature map F respectively, and C is the number of channels.

[0096] In step 102 of some embodiments, a target copy is obtained, and the target copy is encoded to obtain target copy features. The target copy refers to the text material used to generate the picture and text. The target copy can be encoded by a text encoder to obtain target copy features. The text encoder can adopt the BERT model. BERT captures the context information of the target copy through a bidirectional Transformer structure and generates semantically rich target text features. After being processed by BERT, the target copy features are obtained, which can be denoted as text feature vector E, and its dimension is D, where D is the hidden layer dimension of the BERT model.

[0097] In step 103 of some embodiments, the target image features and the target copy features are fused to obtain picture and text fusion features. The target image features and the target copy features can be fused by a fusion layer to obtain picture and text fusion features.

[0098] In one embodiment, step 103 may include: expanding the dimension of the target copy features to make the spatial dimension of the target copy features the same as that of the target image features; adding the target image features and the target copy features element by element to obtain picture and text fusion features.

[0099] Specifically, the picture and text fusion features can be expressed as: U = F + B roa d cast (E). Among them, Broadcast(E) represents expanding the text feature vector E to the dimension of H×W×D, and U refers to the picture and text fusion features.

[0100] The benefits of the above embodiments are as follows: first, ensure the semantic consistency between the text layout and the image content through dimensionality expansion, and then improve the accuracy of semantic fusion by element-wise addition, thereby enhancing the information richness and semantic accuracy of the graphic-text fusion features.

[0101] In step 104 of some embodiments, convolutional attention prediction is performed on the graphic-text fusion features to obtain the text layout description text. Convolutional attention prediction mainly dynamically calculates the weights of each spatial position in the graphic-text fusion features based on the convolutional attention mechanism, and then predicts the text layout with context-aware information, that is, the text layout description text.

[0102] In one embodiment, referring to Figure 2 , step 104 may include:

[0103] Step 201, perform self-attention calculation on the graphic-text fusion features to obtain graphic-text attention features;

[0104] Step 202, perform convolutional attention calculation on the graphic-text attention features to obtain graphic-text layout perception features;

[0105] Step 203, perform layout prediction on the graphic-text layout perception features to obtain the text layout description text.

[0106] In step 201, the process of performing self-attention calculation on the graphic-text fusion features is as follows:

[0107] Where L refers to the graphic-text attention features, softmax is the activation function, p uq refers to the query vector corresponding to the graphic-text fusion feature U, p uk refers to the key vector corresponding to the graphic-text fusion feature, p uv refers to the value vector corresponding to the graphic-text fusion feature U, and d refers to the dimension of the graphic-text fusion feature U. By pre-setting the query weight matrix W uq , the key weight matrix W uk , and the value weight matrix W uv perform vector operations on the graphic-text fusion feature U respectively to obtain p uq , p uk , and p uv .

[0108] In one embodiment, referring to Figure 3 , step 202 may include:

[0109] Step 301, perform convolutional processing on the graphic-text attention features to obtain graphic-text convolutional features;

[0110] Step 302, calculate the attention weights of each spatial position in the graphic-text attention features for the graphic-text convolutional features;

[0111] Step 303: Weighted sum is performed on each spatial position in the graphic-text attention feature according to the attention weight to obtain the graphic-text layout perception feature.

[0112] In step 301, the process of performing convolution processing on the graphic-text attention feature is as follows:

[0113] S = K * L, where S refers to the graphic-text convolution feature and K refers to the weight matrix for convolution processing.

[0114] In step 302, the process of calculating the weight for the graphic-text convolution feature is as follows:

[0115] where α i,j is the attention weight of each spatial position (i, j) in the graphic-text convolution feature S, H is the height, W is the width, k refers to the k-th row, l refers to the l-th column, and S k,l refers to the feature at the k-th row and k-th column in the graphic-text convolution feature S.

[0116] In step 303, the graphic-text layout perception feature is a context-aware representation (CAR). By using the attention weight to perform weighted sum on the feature map, a context-aware feature representation is generated. Specifically, the process of performing weighted sum is as follows:

[0117] where L′ refers to the graphic-text layout perception feature, and L i,j refers to all the channel values of the graphic-text attention feature L at the i-th row and j-th column, which is a vector of length C.

[0118] The advantage of the embodiments of the above steps 301 to 303 is that, through convolution processing and the method of weighted sum based on the attention weight, a context-aware feature representation indicating the graphic-text layout can be generated more accurately, that is, the feature representation ability of the graphic-text layout perception feature is improved, which further helps to improve the accuracy of graphic-text generation.

[0119] In step 203, the graphic and text layout perception feature L′ is input into the fully connected layer to predict the position, size, and style of the text, obtaining the text layout description text. For example, for the poster generation scenario in the above text, at this time, the text layout description text can be: Part of the text is located in the upper left area of the background area of the appearance picture, with a size of occupying 80% of the upper left area, and the style includes Song typeface, No. 4 font size, red color, single line spacing, and 0 paragraph spacing; Part of the text is located in the lower right area of the background area of the interior picture, with a size of occupying 50% of the lower right area, and the style includes Song typeface, No. 5 font size, black color, single line spacing, and 0 paragraph spacing; There is no text in the background area of the vehicle model picture.

[0120] The benefits of the embodiments of the above steps 201 to 203 are that, based on the method of first self-attention processing, then convolutional attention calculation, and finally layout prediction, more accurate information describing the layout of the text in the image can be obtained, that is, the richness of the layout information of the text layout description text is improved.

[0121] In step 105 of some embodiments, the text layout description text and the target text are encoded to obtain the text layout feature. Step 105 may include: concatenating the text layout description text and the target text to obtain the target fusion text; performing text encoding on the target fusion text to obtain the text layout feature. When concatenating, the text layout description text and the target text can be directly concatenated. For example, text1 = "Hello", text2 = "World", result = text1 + text2 = "HelloWorld". When concatenating, a delimiter can also be set between the text layout description text and the target text. text1 = "Hello", text2 = "World", result = text1 + " / " + text2 = "Hello / World". The target fusion text can be text-encoded through a text encoder to obtain the text layout feature. The text encoder can adopt the BERT model. The introduction of the BERT model can refer to the description of step 102 in the above text and will not be elaborated here.

[0122] In step 106 of some embodiments, the text layout feature and the target image feature are split-encoded to obtain the graphic and text layout encoded feature. The text layout feature and the target image feature can be split-encoded through the split encoder of a preset split network to obtain the graphic and text layout encoded feature.

[0123] The segmentation network includes a segmentation encoder and a segmentation decoder. The segmentation encoder includes multiple convolutional layers, pooling layers, and attention modules. Deep features of the data are gradually extracted through convolutional and pooling operations, reducing the spatial resolution and increasing the number of channels. The attention module is used to capture the long-range dependencies in the feature map and weight the features at different positions. The segmentation decoder is symmetric to the segmentation encoder and consists of multiple upsampling layers, convolutional layers, and attention modules. The spatial resolution of the data is gradually restored through upsampling and convolutional operations, reducing the number of channels, and at the same time, the attention module is used to better fuse the feature information at different levels.

[0124] In one embodiment, referring to Figure 4 , step 106 may include:

[0125] Step 401, multiplying a preset query weight matrix by the target image features to obtain an image layout query vector;

[0126] Step 402, multiplying a preset key weight matrix by the copywriting layout features to obtain a text layout key vector;

[0127] Step 403, multiplying a preset value weight matrix by the copywriting layout features to obtain a text layout value vector;

[0128] Step 404, calculating the weights of the image layout query vector and the text layout key vector to obtain a graphic-text layout attention score;

[0129] Step 405, performing non-linear activation processing on the graphic-text layout attention score to obtain a graphic-text layout weight vector;

[0130] Step 406, performing weighted summation on the graphic-text layout weight vector and the text layout value vector to obtain a graphic-text layout encoded feature.

[0131] In step 401, the image layout query vector is as follows:

[0132] Q = xW Q , where Q refers to the image layout query vector, x refers to the target image features; W Q refers to the trainable query weight matrix.

[0133] In step 402, the text layout key vector is as follows:

[0134] K = EW K , where K refers to the text layout key vector, E refers to the copywriting layout features; W K refers to the trainable key weight matrix.

[0135] In step 403, the text layout key vector is as follows:

[0136] V = EW V , where V refers to the text layout value vector, E refers to the copywriting layout feature; W V refers to the trainable value weight matrix.

[0137] In step 404, the graphic-text layout attention score is as follows:

[0138] where S att refers to the graphic-text layout attention score, QK T represents the dot product operation of the image layout query vector Q and the text layout key vector K, and d K refers to the dimension of the text layout key vector K.

[0139] In steps 405 to 406, the graphic-text attention score S att after being processed by non-linear activation (such as Softmax operation), the graphic-text layout weight vector Attn is obtained. By weighted summing the graphic-text layout weight vector Attn with the text layout value vector V, the final output matrix, that is, the graphic-text layout encoding feature Z, can be obtained. Thus, the text information is integrated into the image information, enabling the text to control the generation of the image.

[0140] The benefits of the embodiments of the above steps 401 to 406 are that the integration degree of text information and image information in the graphic-text layout encoding feature is further improved, and the accuracy of graphic-text generation is improved.

[0141] In step 107 of some embodiments, the copywriting layout feature and the graphic-text layout encoding feature are segmented and decoded to obtain the target graphic-text. The segmentation and decoding can be implemented through the segmentation decoder of a preset segmentation network. The specific structure of the segmentation decoder has been described in the description of step 106 above and will not be elaborated here.

[0142] In one embodiment, step 103 includes: performing feature fusion on the target image feature and the target copywriting feature through the fusion layer of a preset layout extractor to obtain the graphic-text fusion feature; and step 104 includes: performing convolutional attention prediction on the graphic-text fusion feature through the layout prediction layer of the layout extractor to obtain the copywriting layout description text. Refer to Figure 5 , before step 103, the graphic-text generation method may further include: pre-training the layout extractor, specifically including:

[0143] Step 501, obtaining a sample masked image, and performing image restoration on the sample masked image through a preset image restoration model to obtain a sample restored image;

[0144] Step 502, performing image encoding on the sample restored image to obtain a sample image feature;

[0145] Step 503: Obtain a sample text, perform text encoding on the sample text to obtain sample text features;

[0146] Step 504: Use a preset initial layout extractor to extract layout information from the sample image features and the sample text features to obtain a sample text layout description text; wherein, the initial layout extractor includes a fusion layer and a layout prediction layer;

[0147] Step 505: Calculate a loss based on the sample text layout description text and a preset labeled text layout description text to obtain layout matching loss data;

[0148] Step 506: Train the initial layout extractor according to the layout matching loss data to obtain a layout extractor.

[0149] In step 501, the sample masked image refers to an image after masking a partial area of the image. For example, by setting a threshold, the part of the image with pixel values higher than the threshold is retained, and the part lower than the threshold is set to black (0) to obtain the masked image, that is, the sample masked image. The image repair model is a neural network model for repairing the masked area in the image, such as InpNet.

[0150] In step 502, the sample repaired image can be encoded by an image encoder to obtain sample image features. The image encoder can adopt the ResNet-50 architecture. The main purpose is to extract high-level visual features from the input target image. These visual features capture the global and local information of the image, providing a basis for subsequent text layout generation.

[0151] The specific process of step 503 is basically the same as that of step 102 in the above text, and will not be elaborated here.

[0152] In step 504, the initial layout extractor includes a fusion layer and a layout prediction layer. Step 604 includes: performing feature fusion on the sample image features and the sample text features through the fusion layer of the initial layout extractor to obtain sample text-image fusion features; performing convolutional attention prediction on the sample text-image fusion features to obtain a sample text layout description text.

[0153] The layout prediction layer can include a self-attention layer, a convolutional attention layer, and a fully connected layer. The above step of performing convolutional attention prediction on the sample text-image fusion features to obtain a sample text layout description text can include:

[0154] Performing self-attention calculation on the sample text-image fusion features through the self-attention layer to obtain sample text-image attention features;

[0155] Performing convolutional attention calculation on the sample text-image attention features to obtain sample text-image layout perception features;

[0156] Perform layout prediction on the layout perception features of the sample graphic text to obtain the sample text layout description text.

[0157] Regarding the specific processes of self-attention calculation, convolutional attention calculation, and layout prediction, reference can be made to the detailed descriptions of steps 201 to 203 in the above text, which will not be elaborated here.

[0158] In step 505, the process of loss calculation is as follows:

[0159] Among them, L layout refers to the layout matching loss data, B refers to the label text layout description text, refers to the sample text layout description text. Other loss error functions can also be selected, such as mean squared error loss, etc., which are not specifically limited in this embodiment.

[0160] In step 506, adjust the parameters of the initial layout extractor according to the layout matching loss data until the layout matching loss data is less than the predetermined loss threshold to obtain the layout extractor.

[0161] The benefits of the embodiments of the above steps 501 to 506 are that by introducing the sample mask image to increase the uncertainty of layout extraction for the initial layout extractor, the trained layout extractor can flexibly handle images with different complex backgrounds, improving the accuracy of layout extraction.

[0162] In one embodiment, the sample text layout description text includes the predicted position of the sample text layout. For example, for the poster generation scenario in the above text, the predicted position of the sample text layout can be "part of the text is located in the upper left area of the background area of the appearance diagram".

[0163] In one embodiment, step 106 may include: performing segmentation encoding on the text layout feature and the target image feature through the segmentation encoder of the preset segmentation network to obtain the graphic text layout encoding feature; and step 107 may include: performing segmentation decoding on the text layout feature and the graphic text layout encoding feature through the segmentation decoder of the segmentation network to obtain the target graphic text. Refer to Figure 6 Before step 106, the graphic text generation method may further include: pre-training the segmentation network, specifically including:

[0164] Step 601, perform text encoding on the sample text layout description text and the sample text to obtain the sample text layout feature;

[0165] Step 602, apply a preset sample noise to the sample image feature to obtain the sample noise-added image feature;

[0166] Step 603: Segment and encode the sample text layout feature and the sample noisy image feature through the segmentation encoder of the preset initial segmentation network to obtain the sample noisy text and image layout encoded feature;

[0167] Step 604: Segment and decode the sample text layout feature and the sample noisy text and image layout encoded feature through the segmentation decoder of the initial segmentation network to obtain the sample noisy text and image and the predicted noise;

[0168] Step 605: Calculate the loss based on the sample noise and the predicted noise to obtain the noise matching loss data;

[0169] Step 606: Remove the noise from the sample noisy text and image according to the predicted noise to obtain the sample denoised text and image;

[0170] Step 607: Perform image recognition on the sample denoised text and image to obtain the recognized position of the sample text layout;

[0171] Step 608: Calculate the loss based on the predicted position of the sample text layout and the recognized position of the sample text layout to obtain the text position control loss data;

[0172] Step 609: Train the initial segmentation network according to the noise matching loss data and the text position control loss data to obtain the segmentation network.

[0173] In step 601, text encoding can be implemented through a text encoder.

[0174] In step 602, the sample noisy image feature is as follows: After adding the sample noise at the t-th step to the sample image feature G, the sample noisy image feature G t can be obtained. Here, t refers to the t-th round of training.

[0175] The specific process of step 603 is the same as that of step 106 above and will not be elaborated here.

[0176] The specific process of step 604 is basically the same as that of step 107 above and will not be elaborated here. In addition, since the sample noisy text and image layout encoded feature contains noise, in addition to decoding the sample noisy text and image, the segmentation encoder can also decode the predicted noise.

[0177] In step 605, the noise matching loss data is as follows: Among them, L noise refers to the noise matching loss data, ε refers to the sample noise, and

[0178] refers to the predicted noise.

[0179] In step 607, the denoised sample text image can be recognized through optical character recognition technology to obtain the recognized position of the sample text layout. Optical Character Recognition (OCR): is a technology that converts text in paper documents, pictures, etc. into a computer-editable text format. In this embodiment, the OCR model is used to recognize the recognized position of the sample text layout from the denoised sample text image.

[0180] In step 608, the text position control loss data is as follows:

[0181] where L control refers to the text position control loss data, refers to the predicted position of the sample text layout, and C refers to the recognized position of the sample text layout. By using to participate in the loss calculation, it can promote that the text layout position decoded by the segmentation network (U-net) is controlled by the layout conditions.

[0182] In step 609, first, the noise matching loss data and the text position control loss data can be fused to obtain the target loss data; then, the parameters of the initial segmentation network can be adjusted according to the target loss data to obtain the segmentation network.

[0183] Specifically, the target loss data is as follows:

[0184] L U =αL noise +βL control where L U refers to the target loss data, α refers to the noise loss weight parameter, and β refers to the position loss weight parameter.

[0185] The benefit of the embodiments of the above steps 601 to 609 is that the segmentation network is jointly trained through the position control loss and the noise matching loss, so that the segmentation network has high accuracy in generating the text layout, especially extremely high precision in the generation accuracy of the text position.

[0186] Combining the above embodiments, the present application can at least achieve the following beneficial effects: introducing a convolutional attention mechanism to dynamically calculate the attention weights of each spatial position in the image feature map, and preferentially selecting areas suitable for text placement. Generating a context-aware feature representation through weighted summation to ensure the coordination between the text layout and the image content. Ensuring the accuracy of the generated text position through the training of the segmentation network.

[0187] Please refer to Figure 7 , the embodiments of the present application also provide a text and image generation device, which can implement the above text and image generation method.Figure 7 The block diagram of the module structure of the graphic text generation device provided by the embodiment of the present application. The device includes:

[0188] An image encoding module 701, configured to obtain a target image and perform image encoding on the target image to obtain target image features;

[0189] A text encoding module 702, configured to obtain a target text and perform text encoding on the target text to obtain target text features;

[0190] A feature fusion module 703, configured to perform feature fusion on the target image features and the target text features to obtain graphic text fusion features;

[0191] A layout prediction module 704, configured to perform convolutional attention prediction on the graphic text fusion features to obtain a text layout description text;

[0192] A fusion encoding module 705, configured to encode the text layout description text and the target text to obtain text layout features;

[0193] A segmentation encoding module 706, configured to perform segmentation encoding on the text layout features and the target image features to obtain graphic text layout encoding features;

[0194] A segmentation decoding module 707, configured to perform segmentation decoding on the text layout features and the graphic text layout encoding features to obtain a target graphic text.

[0195] In one embodiment, the graphic text generation device further includes: a first training module, configured to pre-train a layout extractor.

[0196] In one embodiment, the graphic text generation device further includes: a second training model, configured to pre-train a segmentation network.

[0197] It should be noted that the specific implementation manner of this graphic text generation device is basically the same as the specific embodiment of the above graphic text generation method, and will not be elaborated here.

[0198] The embodiment of the present application further provides an electronic device. The electronic device includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, the above graphic text generation method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0199] Please refer to Figure 8 , Figure 8 , which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0200] The processor 801 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0201] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the graphic generation method of the embodiments of the present application;

[0202] The input / output interface 803 is used to implement information input and output;

[0203] The communication interface 804 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0204] The bus 805 transmits information between the various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);

[0205] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.

[0206] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned graphic generation method.

[0207] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0208] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0209] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0212] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above figures are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0213] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0214] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0215] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0217] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0218] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.

Claims

1. A method for generating graphics and texts, characterized in that The method includes: Obtain a target image, and perform image encoding on the target image to obtain target image features; Obtain a target text, and perform text encoding on the target text to obtain target text features; Perform feature fusion on the target image features and the target text features to obtain text-image fusion features; Perform convolutional attention prediction on the text-image fusion features to obtain a text layout description text; Encode the text layout description text and the target text to obtain text layout features; Perform segmentation encoding on the text layout features and the target image features to obtain text-image layout encoding features; Perform segmentation decoding on the text layout features and the text-image layout encoding features to obtain a target text-image.

2. The method according to claim 1, characterized in that, The performing convolutional attention prediction on the text-image fusion features to obtain a text layout description text includes: Perform self-attention calculation on the text-image fusion features to obtain text-image attention features; Perform convolutional attention calculation on the text-image attention features to obtain text-image layout perception features; Perform layout prediction on the text-image layout perception features to obtain the text layout description text.

3. The method according to claim 2, wherein The performing convolutional attention calculation on the text-image attention features to obtain text-image layout perception features includes: Perform convolutional processing on the text-image attention features to obtain text-image convolutional features; Perform weight calculation on the text-image convolutional features to obtain the attention weight of each spatial position in the text-image attention features; Perform weighted summation on each spatial position in the text-image attention features according to the attention weight to obtain the text-image layout perception features.

4. The method according to any one of claims 1 to 3, characterized in that, The performing segmentation encoding on the text layout features and the target image features to obtain text-image layout encoding features includes: Multiply a preset query weight matrix by the target image features to obtain an image layout query vector; Multiply a preset key weight matrix by the text layout features to obtain a text layout key vector; Multiply a preset value weight matrix by the text layout features to obtain a text layout value vector; Perform weight calculation on the image layout query vector and the text layout key vector to obtain text-image layout attention scores; Perform non-linear activation processing on the text-image layout attention scores to obtain a text-image layout weight vector; Perform weighted summation on the text-image layout weight vector and the text layout value vector to obtain the text-image layout encoding features.

5. The method according to any one of claims 1 to 3, characterized in that, The performing feature fusion on the target image features and the target text features to obtain text-image fusion features includes: Perform dimensionality expansion on the target text features so that the spatial dimension of the target text features is the same as the spatial dimension of the image features; Perform element-wise addition on the target image features and the target text features to obtain the text-image fusion features.

6. The method according to any one of claims 1 to 3, wherein Performing feature fusion on the target image feature and the target copywriting feature to obtain a graphic-text fusion feature includes: performing feature fusion on the target image feature and the target copywriting feature through a fusion layer of a preset layout extractor to obtain a graphic-text fusion feature; Performing convolutional attention prediction on the graphic-text fusion feature to obtain a copywriting layout description text includes: performing convolutional attention prediction on the graphic-text fusion feature through a layout prediction layer of the layout extractor to obtain a copywriting layout description text; Wherein, before performing feature fusion on the target image feature and the target copywriting feature through a fusion layer of a preset layout extractor to obtain a graphic-text fusion feature, the method further includes: Pre-training the layout extractor, specifically including: Obtaining a sample masked image, and performing image restoration on the sample masked image through a preset image restoration model to obtain a sample restored image; Performing image encoding on the sample restored image to obtain a sample image feature; Obtaining a sample copywriting, and performing text encoding on the sample copywriting to obtain a sample copywriting feature; Performing layout information extraction on the sample image feature and the sample copywriting feature through a preset initial layout extractor to obtain a sample copywriting layout description text; wherein, the initial layout extractor includes a fusion layer and a layout prediction layer; Calculating a loss according to the sample copywriting layout description text and a preset labeled copywriting layout description text to obtain layout matching loss data; Training the initial layout extractor according to the layout matching loss data to obtain the layout extractor.

7. The method according to claim 6, characterized in that The sample copywriting layout description text includes sample text layout prediction positions; Performing segmentation encoding on the copywriting layout feature and the target image feature to obtain a graphic-text layout encoding feature includes: performing segmentation encoding on the copywriting layout feature and the target image feature through a segmentation encoder of a preset segmentation network to obtain a graphic-text layout encoding feature; Performing segmentation decoding on the copywriting layout feature and the graphic-text layout encoding feature to obtain a target graphic-text includes: performing segmentation decoding on the copywriting layout feature and the graphic-text layout encoding feature through a segmentation decoder of the segmentation network to obtain a target graphic-text; Wherein, before performing segmentation encoding on the copywriting layout feature and the target image feature through a segmentation encoder of a preset segmentation network to obtain a graphic-text layout encoding feature, the method further includes: Pre-training the segmentation network, specifically including: Performing text encoding on the sample copywriting layout description text and the sample copywriting to obtain a sample copywriting layout feature; Applying a preset sample noise to the sample image feature to obtain a sample noise-added image feature; Performing segmentation encoding on the sample copywriting layout feature and the sample noise-added image feature through a segmentation encoder of a preset initial segmentation network to obtain a sample noise-added graphic-text layout encoding feature; Performing segmentation decoding on the sample copywriting layout feature and the sample noise-added graphic-text layout encoding feature through a segmentation decoder of the initial segmentation network to obtain a sample noise-added graphic-text and a predicted noise; Perform loss calculation based on the sample noise and the predicted noise to obtain noise matching loss data; Perform noise removal on the sample noisy text image according to the predicted noise to obtain a sample denoised text image; Perform image recognition on the sample denoised text image to obtain the recognized position of the sample text layout; Perform loss calculation based on the predicted position of the sample text layout and the recognized position of the sample text layout to obtain text position control loss data; Train the initial segmentation network according to the noise matching loss data and the text position control loss data to obtain the segmentation network.

8. A graphic and text generation device, characterized in that, The device includes: An image encoding module, configured to obtain a target image and perform image encoding on the target image to obtain target image features; A text encoding module, configured to obtain a target copywriting and perform text encoding on the target copywriting to obtain target copywriting features; A feature fusion module, configured to perform feature fusion on the target image features and the target copywriting features to obtain text-image fusion features; A layout prediction module, configured to perform convolutional attention prediction on the text-image fusion features to obtain a text layout description text; A fusion encoding module, configured to encode the text layout description text and the target copywriting to obtain text layout features; A segmentation encoding module, configured to perform segmentation encoding on the text layout features and the target image features to obtain text-image layout encoding features; A segmentation decoding module, configured to perform segmentation decoding on the text layout features and the text-image layout encoding features to obtain a target text image.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the text-image generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by the processor, it implements the text-image generation method according to any one of claims 1 to 7.