Figure generation method and device, storage medium and electronic equipment

Through the target pooled text features and the autoregressive model of literary and biographics, the problem of low generation efficiency of literary and biographics technology is solved, and high-quality images are efficiently generated.

CN120472022APending Publication Date: 2025-08-12AIBASHI JAPAN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510383398.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing literary and biographical technology has low efficiency in generating images, making it difficult to easily generate high-quality images based on text.

Method used

The target pooled text features and target text image autoregression model are used to generate the target generated image through multi-scale image feature extraction and decoding processing.

Benefits of technology

Improves the efficiency and quality of image generation, ensuring consistency and creativity between images and text descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472022A_ABST
    Figure CN120472022A_ABST
Patent Text Reader

Abstract

The invention provides a graph generation method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining target text data, calling a target text coding model, carrying out the feature extraction of the target text data, and obtaining a feature extraction result of the target text data, the feature extraction result of the target text data comprises target pooling text features of the target text data; a target text graph autoregression model is called, based on the target pooling text features, image features of the target text data under each scale in K scales are determined in sequence, and K is a positive integer; wherein the image feature of the target text data under the Kth scale is determined based on the image feature of the target text data under the (K-1) th scale; and decoding the image features of the target text data under the Kth scale to obtain a target generated image of the target text data. According to the embodiment of the invention, the image can be conveniently generated according to the text to improve the image generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a text-generating diagram method, device, storage medium, and electronic equipment. Background Art

[0002] Currently, text-to-image generation technology (i.e., text-to-image generation technology) allows users to generate realistic and creative images based on text instructions. This technology has applications in digital painting, virtual try-on, prototyping, and other fields. However, related technologies typically use generative adversarial networks and diffusion models to introduce latent representations to generate images, which often takes a long time and results in low image generation efficiency. Therefore, there is currently no effective solution for quickly generating images from text to improve image generation efficiency. Summary of the Invention

[0003] In view of this, the embodiments of the present invention provide a text-generated image method, device, storage medium and electronic device to solve the problem of low efficiency of generating images from text in related technologies; that is, the embodiments of the present invention can conveniently generate target generated images through target pooling text features and target text-generated image autoregressive model, that is, it can conveniently generate images based on text, so as to effectively improve the efficiency of image generation.

[0004] According to one aspect of an embodiment of the present invention, a method for creating a cultural image is provided, the method comprising:

[0005] Acquire target text data, call a target text encoding model, perform feature extraction on the target text data, and obtain a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes a target pooled text feature of the target text data;

[0006] Calling a target text-image autoregressive model, based on the target pooled text features, sequentially determining the image features of the target text data at each of K scales, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale;

[0007] The image features of the target text data at the K-th scale are decoded to obtain a target generated image of the target text data.

[0008] According to one aspect of an embodiment of the present invention, another method for a Vincent diagram is provided, the method comprising:

[0009] Acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and label text of each training image in the at least one training image;

[0010] For any training image of the at least one training image, select M training scales from the first P scales of the K scales, call the multi-scale image coding model, perform image coding on the any training image, and obtain image features of the any training image at each of the M training scales, where K, P, and M are all positive integers;

[0011] Performing feature extraction on the label text of any one of the training images to obtain a feature extraction result of the label text of any one of the training images, wherein the feature extraction result of the label text of any one of the training images includes a label pooling text feature of the label text of any one of the training images;

[0012] Calling the initial text-image autoregressive model, determining the image features of the label text of any training image at each training scale based on the label pooled text features of the label text of any training image; and determining the model loss value for any training image based on the image features of the any training image at each training scale and the image features of the label text of any training image at each training scale;

[0013] After obtaining the model loss value under each training image, the loss value of the text graph autoregression model is calculated based on the model loss value under each training image; and the model parameters in the initial text graph autoregression model are optimized in the direction of reducing the loss value of the text graph autoregression model to obtain the model-optimized initial text graph autoregression model, and the target text graph autoregression model is determined based on the model-optimized initial text graph autoregression model; wherein the target text graph autoregression model supports a target generated image for generating target text data.

[0014] According to another aspect of an embodiment of the present invention, a cultural image device is provided, the device comprising:

[0015] A first acquiring unit, configured to acquire target text data;

[0016] a first processing unit, configured to call a target text encoding model, perform feature extraction on the target text data, and obtain a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes a target pooled text feature of the target text data;

[0017] The first processing unit is further configured to call a target text-image autoregressive model, and determine, based on the target pooled text features, the image features of the target text data at each of K scales, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale;

[0018] The first processing unit is further configured to decode the image features of the target text data at the K-th scale to obtain a target generated image of the target text data.

[0019] According to another aspect of an embodiment of the present invention, another Wensheng diagram device is provided, the device comprising:

[0020] A second acquisition unit is configured to acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and a label text of each training image in the at least one training image;

[0021] a second processing unit, configured to, for any training image among the at least one training image, select M training scales from first P scales of the K scales, and call a multi-scale image coding model to perform image coding on the any training image to obtain image features of the any training image at each of the M training scales, where K, P, and M are all positive integers;

[0022] The second processing unit is further configured to perform feature extraction on the label text of any training image to obtain a feature extraction result of the label text of any training image, wherein the feature extraction result of the label text of any training image includes a label pooling text feature of the label text of any training image;

[0023] The second processing unit is further configured to call the initial text-image autoregressive model to determine, based on the label pooled text features of the label text of the any training image, the image features of the label text of the any training image at each training scale; and determine, based on the image features of the any training image at each training scale and the image features of the label text of the any training image at each training scale, the model loss value for the any training image;

[0024] The second processing unit is further configured to, after obtaining the model loss value under each training image, calculate a loss value of a text graph autoregressive model based on the model loss value under each training image; and optimize the model parameters in the initial text graph autoregressive model in a direction of reducing the loss value of the text graph autoregressive model to obtain an optimized initial text graph autoregressive model, and determine a target text graph autoregressive model based on the optimized initial text graph autoregressive model; wherein the target text graph autoregressive model supports a target generated image for generating target text data.

[0025] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the above-mentioned method.

[0026] According to another aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned method.

[0027] The embodiment of the present invention can obtain target text data, call the target text encoding model, perform feature extraction on the target text data, and obtain the feature extraction results of the target text data, which include the target pooled text features of the target text data. Then, the target text-image autoregressive model can be called to determine the image features of the target text data at each of the K scales based on the target pooled text features, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale. Furthermore, the image features of the target text data at the Kth scale can be decoded to obtain the target generated image of the target text data. The target text-image autoregressive model is obtained by training the model through a graph-text matching dataset, which can effectively improve the model performance of the target text-image autoregressive model. It can be seen that the embodiment of the present invention can conveniently generate a target generated image through the target pooling text features and the target text-image autoregressive model. That is to say, the embodiment of the present invention can conveniently generate an image based on the text (such as conveniently generating a target generated image through the target text data) to effectively improve the image generation efficiency, that is, effectively improve the efficiency of generating images based on text, and effectively improve the quality of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Further details, features and advantages of the present invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0029] Figure 1 A schematic flow chart of a method for producing a cultural graph according to an exemplary embodiment of the present invention is shown;

[0030] Figure 2 A schematic diagram of a Wensheng graph autoregressive model according to an exemplary embodiment of the present invention is shown;

[0031] Figure 3 A schematic diagram of an attention model according to an exemplary embodiment of the present invention is shown;

[0032] Figure 4 A schematic flow chart of another method for generating a cultural graph according to an exemplary embodiment of the present invention is shown;

[0033] Figure 5 A schematic diagram of model training according to an exemplary embodiment of the present invention is shown;

[0034] Figure 6 A schematic block diagram of a cultural graph device according to an exemplary embodiment of the present invention is shown;

[0035] Figure 7 A schematic block diagram of another cultural image device according to an exemplary embodiment of the present invention is shown;

[0036] Figure 8 A block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention is shown. DETAILED DESCRIPTION

[0037] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0038] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0039] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0040] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0041] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0042] It should be noted that the execution subject of the text graph method provided in the embodiment of the present invention may be one or more electronic devices, which is not limited in the embodiment of the present invention; wherein, the electronic device may be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and the multiple electronic devices include at least one terminal and at least one server, the text graph method provided in the embodiment of the present invention may be jointly executed by the terminal and the server. Accordingly, the terminals mentioned here may include but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, and the like. The server mentioned here may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing (cloud computing), cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, and the like.

[0043] Based on the above description, an embodiment of the present invention proposes a text-generated graph method, which can be executed by the electronic device (terminal or server) mentioned above; or, the text-generated graph method can be executed by the terminal and the server together. For the sake of convenience, the following description will be based on an example of an electronic device executing the text-generated graph method; Figure 1 As shown, the text-generated graph method may include the following steps S101-S103:

[0044] S101, acquiring target text data, calling a target text encoding model, performing feature extraction on the target text data, and obtaining a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes target pooled text features of the target text data.

[0045] Optionally, the target text data may be any text (also referred to as a text prompt or text data), which is not limited in the embodiment of the present invention.

[0046] In the embodiment of the present invention, the target text data may be obtained in the following ways, but is not limited to:

[0047] The first acquisition method: a plurality of text data may be stored in the storage space of the electronic device itself, and the electronic device may select target text data from the plurality of text data (such as randomly selecting or selecting in sequence, etc.) to obtain the target text data.

[0048] The second acquisition method: the electronic device can obtain a text download link, and use the text data downloaded based on the text download link as the target text data to achieve acquisition of the target text data.

[0049] The third acquisition method: the electronic device may have a text input interface, that is, a displayable text input interface. In this case, the user may perform a text input operation on the text input interface, and the electronic device may respond to the text input operation detected on the text input interface to use the text data input by the text input operation as the target text data, and so on.

[0050] Optionally, a text encoding model (also referred to as a text encoder) can be a clip (Contrastive Language-Image Pre-Training) model (a deep learning model, a pre-trained model that can process text and images simultaneously, and can be a picture-text matching model). Optionally, the target text encoding model can be pre-trained, such as an electronic device can pre-train the target text encoding model to store the target text encoding model, thereby facilitating the call of the target text encoding model; or, the target text encoding model can be set according to experience or actual needs, or can be obtained from other devices, etc.; the embodiments of the present invention are not limited to this.

[0051] Optionally, the electronic device can extract the target refined text features of the target text data through the target text encoding model (such as the concatenation result between the representation vectors of each text recognition result in at least one text recognition result corresponding to the target text data, which can be a feature of length L). That is, the target text encoding model can be called to perform feature extraction on the target text data to obtain the target refined text features of the target text data; then, the target pooled text features can be calculated based on the target refined text features, that is, the target refined text features can be pooled to obtain the target pooled text features; optionally, the pooling process can be set according to experience or according to actual needs, and the embodiment of the present invention is not limited to this. Optionally, L can be a positive integer, such as L can be the sum of the dimensions of the representation vectors of each text recognition result. Optionally, a pooled text feature can be a 1×1 feature map (i.e., a feature map with a scale of 1×1), and the depth of a pooled text feature can be 2D, that is, a pooled text feature can be a 1×1×2D feature map (i.e., an image feature), that is, the dimension of the feature of a grid block can be 2D, and D is a positive integer. It should be understood that the feature of a grid block may include the value of the corresponding grid block in each channel dimension in 2D channel dimensions (ie, dimensions, which may also be referred to as feature dimensions).

[0052] Optionally, the electronic device may include a text feature guidance module. The electronic device may then call the text feature guidance module to extract generalizable text features (such as pooled text features, etc.) through the target text encoding model, thereby using the pooled text features (such as the target pooled text features) as the starting token, that is, as the image features at the first scale. Based on this, the embodiments of the present invention provide universal text features that help adapt to new text scenarios and ensure consistency between text descriptions and visual outputs.

[0053] S102, calling the target text-image autoregressive model, based on the target pooled text features, sequentially determining the image features of the target text data at each of K scales, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale.

[0054] Optionally, a text graph autoregressive model may include but is not limited to at least one of the following: at least one attention model, at least one interpolation module, and at least one position encoding module, etc.; this embodiment of the present invention is not limited to this. Optionally, each attention model in a text graph autoregressive model may be the same, that is, the model parameters of each attention model in a text graph autoregressive model may be the same. Exemplarily, the target text graph autoregressive model may include K-1 target attention models, K-1 interpolation modules, and K-1 position encoding modules; wherein the k-1th target attention model, the k-1th interpolation module, and the k-1th position encoding module in the target text graph autoregressive model can be used to determine the image features of the target text data at the kth scale through the image features of the target text data at the k-1th scale, and so on.

[0055] Optionally, the K scales can be set based on experience or actual needs, and this is not limited in this embodiment of the present invention. In this embodiment of the present invention, a scale can be an image size, such as 32×32; accordingly, a size can be used to indicate the horizontal and vertical number of grid blocks; optionally, a grid block can be used to represent a pixel point, that is, in this embodiment of the present invention, a grid block can also be referred to as a pixel point, and so on.

[0056] Optionally, an attention model can be a Transformer model (a self-attention model), which can also be called a self-attention model. Based on this, an attention model can include at least one network layer, and a network layer can include a self-attention layer and a feedforward layer. For example, Figure 2 As shown, a text-graph autoregressive model may include an attention model to generate image features of the target text data at the kth scale; it should be understood that Figure 2The Wensheng graph autoregressive model is merely exemplified and is not limited in this regard in the embodiments of the present invention. For example, a Wensheng graph autoregressive model may further include a position encoding module. For another example, an attention model (such as a Transformer) and an interpolation module may be included between any two scales, and so on.

[0057] In an embodiment of the present invention, the electronic device can call the target text-graph autoregressive model, use the target pooled text features as the image features of the target text data at the first scale, and determine the image features of the target text data at each scale in k-1 scales in sequence based on the image features of the target text data at the first scale, k∈[2,K]; that is, the image features of the target text data at the second scale can be determined based on the image features of the target text data at the first scale, and the image features of the target text data at the third scale can be determined based on the image features of the target text data at the second scale, and so on, until the image features of the target text data at the k-1th scale are determined.

[0058] Furthermore, the electronic device may perform interpolation processing on the image features of the target text data at the k-1th scale to obtain interpolated image features, where the scale of the interpolated image features is equal to the k-th scale, and the k-th scale is greater than the k-1th scale; optionally, for any grid block at the k-th scale, the features of each of the first T grid blocks closest to any grid block at the k-th scale may be determined from the image features of the target text data at the k-1th scale, and a weighted sum (such as a mean operation, etc.) may be performed on the features of each grid block in the first T grid blocks to obtain the features of any grid block in the interpolated image features (i.e., the features of any grid block at the k-th scale in the interpolated image features), where T is a positive integer.

[0059] Based on this, the electronic device can determine the first image feature based on the interpolated image feature. In one embodiment, the electronic device can use the interpolated image feature as the first image feature. In another embodiment, the electronic device can determine the target normalized grid scale, and calculate the position coding features of each grid block at the kth scale according to the target normalized grid scale and the interpolated image feature; then, the first image feature can be determined based on the position coding features of each grid block, and so on. Optionally, the target normalized grid scale can be set according to experience or according to actual needs, and the embodiment of the present invention is not limited to this; illustratively, the target normalized grid scale can be the maximum scale in the corresponding Wensheng graph autoregressive model, such as the Kth scale.

[0060] Optionally, the electronic device can calculate the position coding features of each grid block at the kth scale according to the target normalized grid scale through two-dimensional normalized rotation position encoding (Normalized RoPE), so that the relative positions between any tokens (i.e., any grid blocks) are normalized to a uniform scale; based on this, the embodiment of the present invention can ensure a unified understanding of the relative positions in feature maps of different scales, avoid confusion caused by simultaneous encoding of positions of different scales, and better adapt to scale-based feature map prediction tasks; and the normalized rotation position encoding proposed in the embodiment of the present invention does not require additional parameters, is easier to train, and provides potential possibilities for the generation of higher resolution images.

[0061] Optionally, given a sequence position i (e.g., a grid block in row i at scale k, where i is a positive integer) and a feature dimension D, the electronic device may calculate a rotation factor θ using formula 1.1: i :

[0062]

[0063] Accordingly, the rotation factor can be used to calculate the rotation matrix, which is then applied to feature x. For feature x with dimension d (e.g., from 1 to D / 2), formula 1.2 can be used to perform a rotation operation on feature x with dimension d:

[0064]

[0065] Among them, RoPE can represent rotational position encoding, RoPE x (i) can represent the result of rotating each dimension of feature x (i.e., rotating position encoding), and the row number of feature x is the i-th row under the k-th scale. In the embodiment of the present invention, for a given scale k (i.e., the k-th scale), a size of h can be constructed. k ×w k Two-dimensional grid (i.e. the kth scale can be h k ×w k ), h k and w k can represent the height (i.e., the number of rows) and width (i.e., the number of columns) of the feature map at the kth scale respectively. Assuming that the target normalized grid scale is H×W, the electronic device can use formula 1.3 to calculate the normalized rotation position code PE(i, j) of any position (i, j) in the two-dimensional grid (i.e., any grid block at the kth scale), which can also be called the position code feature of any grid block at the kth scale:

[0066]

[0067] Among them, i and j can represent the row and column numbers of any grid block at the kth scale respectively. The concatenation operation can be expressed in the channel dimension based on the sequence position i / h k .H rotation position encoding (which may include the normalized number of rows of any grid block at the kth scale in the first D dimensions of the features of any grid block (i.e. x h ) and the rotation position encoding in each dimension of the sequence j / w k .W rotation position encoding (which may include the normalized number of columns of any grid block at the kth scale in the last D dimensions of the features of any grid block (i.e. x w ) in each dimension), determine the position encoding feature of any grid block at the kth scale; where x h , x w The first D dimensions of the features of any grid block and the last D dimensions of the features of any grid block can be represented respectively. Accordingly, the electronic device can use formula 1.2 to calculate the rotational position coding of the normalized number of columns of any grid block at the k-th scale in each dimension of the first D dimensions of the features of any grid block, and calculate the rotational position coding of the normalized number of columns of any grid block at the k-th scale in each dimension of the last D dimensions of the features of any grid block, thereby obtaining the position coding features of any grid block at the k-th scale; wherein the features of any grid block here can be the features of any grid block in the interpolated image features, so as to realize the calculation of the position coding features of each grid block at the k-th scale according to the target normalized grid scale and the interpolated image features.

[0068] Based on this, the embodiments of the present invention can ensure that the position information in the horizontal and vertical dimensions can be effectively integrated, thereby generating a position code with rich spatial relationship expression; it can be seen that the embodiments of the present invention can more accurately process position information at different scales, and can normalize the relative positions between any discrete features to a unified scale, avoiding the differences and confusion in position interpretation between different scales, which not only improves the efficiency of position encoding and reduces model parameters, but also supports more precise position control, so that the model has better performance and adaptability when generating high-resolution images.

[0069] Optionally, when determining the first image feature based on the position coding features of each grid block, for any grid block at the kth scale, the electronic device may use the position coding features of any grid block as the features of that grid block in the first image feature to determine the first image feature, and so on. Optionally, the features of a grid block in an image feature may also be referred to as the features of the corresponding grid block in the corresponding image feature. In other embodiments, the electronic device may also use the interpolated image features as the first image feature, which is not limited by the present invention.

[0070] Furthermore, the electronic device may perform feature extraction on the first image feature to obtain the second image feature, that is, the target attention model (such as the k-1th target attention model) in the target text graph autoregressive model may be used to perform feature extraction on the first image feature to obtain the second image feature. Optionally, for the self-attention layer in the target attention model, the first image feature may be used as the key and query (that is, the interpolated image feature may be normalized RoPE to obtain the key and query under the self-attention layer), and the interpolated image feature may be used as the value; or, the first image feature may be used as the key, query, and value, etc.; this is not limited in the embodiment of the present invention. Optionally, the feature extraction result of the target text data may also include the target refined text feature of the target text data, the target text graph autoregressive model may include the target attention model, and the target attention model includes at least one cross-attention layer; optionally, there may be a cross-attention layer (such as ) between the self-attention layer (also called multi-head attention) and the feedforward layer of each network layer of an attention model. Figure 3As shown), the keys and values under any cross-attention layer can be the refined text features of the corresponding text data, and a cross-attention layer can also be called multi-head cross-attention; based on this, the electronic device can call the target attention model, and based on the target refined text features, perform feature extraction on the first image features to obtain the second image features. At this time, the target refined text features can be used as the keys (i.e., represented as K) and values (which can be represented as V) under any cross-attention layer in the target attention model, and the query under any cross-attention layer (which can be represented as Q) can be determined based on the output of the previous self-attention layer. Optionally, when extracting features from the first image features, feature extraction can be performed on the features of each grid block in the first image features respectively (i.e., the features of each grid block in the first image features can be used as a query respectively), and the feature extraction results of each grid block at the kth scale can be used as the features of the corresponding grid blocks in the second image features to obtain the second image features. It should be understood that during the model training process, the feature extraction result of a label text may also include the label-refined text features of the corresponding label text, the attention model to be trained may also include at least one cross-attention layer, and the electronic device may also determine the image features of the label text of any training image at the kth scale based on the label-refined text features of the label text of any training image to perform model training, etc.; it should be noted that the specific implementation method of determining the image features of the label text of any training image at the kth scale during the model training process may be the same as the specific implementation method of determining the image features of the target text data at the kth scale, and the embodiments of the present invention will not be repeated here. It can be seen that the embodiments of the present invention can provide detailed text guidance for feature maps at each scale by refining text features. That is to say, attention can further obtain fine-grained text prompts. These cross-attention layers can use the features extracted by the pre-trained text encoder (such as the target refined text features) to interact with the currently generated image features, so that each generation step can be accurately adjusted according to the text instructions, so that the details of the text can more accurately guide the generation of the image at different scales, thereby enhancing the robust consistency between the image and the text; it should be understood that during the generation process, the image details of different scales require different degrees of text guidance. The embodiments of the present invention can adjust the generation of the image at each scale according to the instructions provided by the text (i.e., refined text features) through the cross-attention mechanism to ensure that each detail matches the text description. This inter-layer cross-attention mechanism allows the model to better process complex information in the text while maintaining image quality. In other embodiments, the target attention model may not include a cross-attention layer, and the present invention is not limited to this.

[0071] Based on this, the electronic device can determine the image features of the target text data at the kth scale based on the second image features. Optionally, for the features of any grid block in the second image features, the electronic device can classify and predict the features of any grid block, obtain the probability vector of any grid block, and determine the predicted category of any grid block based on the probability vector of any grid block; then, based on the predicted category of any grid block, the reference features indicated by the predicted category of any grid block can be determined from the reference feature data, and based on the reference features indicated by the predicted category of any grid block, the features of any grid block in the image features of the target text data at the kth scale can be determined, so as to achieve the determination of the image features of the target text data at the kth scale. Optionally, when determining the predicted category of any grid block based on the probability vector of any grid block, the electronic device may adopt a specified sampling strategy to determine the categories with the largest top N probability values from the probability vector of any grid block, and use the categories with the largest top N probability values as the predicted category of any grid block, N is a positive integer, and the specified sampling strategy can be used to indicate the value of N; then correspondingly, when determining the features of any grid block in the image features of the target text data at the kth scale based on the benchmark features indicated by the predicted category of any grid block, when N is 1, the number of predicted categories of any grid block is 1, and at this time the benchmark features indicated by the predicted category of any grid block can be used as the image features of the target text data at the kth scale. The features of any grid block in the features, when N is greater than 1, the number of prediction categories of any grid block can be multiple. At this time, the benchmark features indicated by the prediction category of any grid block can include the benchmark features indicated by each prediction category of any grid block. Then, the benchmark features indicated by each prediction category of any grid block can be weighted summed (such as mean operation, etc.) to obtain the features of any grid block in the image features of the target text data at the kth scale, or a benchmark feature can be randomly selected from the benchmark features indicated by each prediction category of any grid block to use the randomly selected benchmark feature as the feature of any grid block in the image features of the target text data at the kth scale, and so on; the embodiment of the present invention is not limited to this. Optionally, the implementation method for determining the prediction category in the training process may be different from the implementation method for determining the prediction category in the application process, that is, the value of N may be different; or, the value of N may be the same, and the embodiment of the present invention is not limited to this. Optionally, the specified sampling strategy can be set according to experience or according to actual needs, and the embodiment of the present invention is not limited to this. It should be understood that existing generative models typically rely on category embeddings as starting tokens. These category embeddings are usually associated with a predefined set of categories and are susceptible to mode collapse, which limits the diversity of generated images. However, the embodiments of the present invention can introduce more flexible latent representations and learning mechanisms by specifying sampling strategies, thereby enhancing the model's ability to generate diverse images and ensuring that the generated images are both realistic and creative.

[0072] Furthermore, the electronic device may use the image features of the target text data at the kth scale as the image features of the target text data at the k-1th scale, and iteratively perform the above-mentioned interpolation processing on the image features of the target text data at the k-1th scale until the image features of the target text data at the Kth scale are obtained; that is, the above-mentioned interpolation processing on the image features of the target text data at the k-1th scale may be iteratively performed to obtain the image features of the target text data at the kth scale, until the image features of the target text data at the Kth scale are obtained.

[0073] Based on this, an embodiment of the present invention can realize a scale-based discrete feature map (i.e., image feature) prediction paradigm to apply the scale-based discrete feature map autoregressive generation paradigm to open set text guided generation. It can be based on a series of discrete feature maps of different scales defined by discrete image encoders (such as modules between any two scales, which may include but are not limited to at least one of the following: attention model, interpolation module, and position encoding module, etc.), and gradually predict higher resolution discrete feature maps based on lower resolution discrete feature maps (i.e., predict higher scale discrete feature maps, one scale can represent one resolution), avoiding the problem that the original autoregressive scheme is difficult to model image modality (i.e., avoiding the problem that the original discrete token-based autoregressive scheme is difficult to model image modality), so that the generated image has higher authenticity. It can be seen that the embodiments of the present invention take into account the shortcomings of existing autoregressive models in processing the bidirectional and two-dimensional structural dependencies of images, and develop a new autoregressive model architecture, which can better understand and generate images with complex two-dimensional structures; that is, existing autoregressive models usually directly model image distributions and gradually construct images by generating tokens (such as usually using structures such as PixelRNN (Pixel Recurrent Neural Networks, a generative model) or PixelCNN (Pixel Convolutional Neural Networks, another generative model) to construct images, so as to use the previous pixel in an image to determine the next pixel in the image, etc.), and usually cannot effectively process the bidirectional and two-dimensional structural dependencies in the image, resulting in a decrease in image quality. The embodiments of the present invention can optimize the network structure of the existing model or develop a new network module, which can effectively improve the structured processing capability of the autoregressive model, so as to more accurately capture and utilize the spatial relationship in the image, thereby determining the next image feature by the previous image feature, and effectively improving the accuracy of the generated image.

[0074] S103 , decoding the image features of the target text data at the Kth scale to obtain a target generated image of the target text data.

[0075] Optionally, the electronic device may call a discrete image decoder (also referred to as a target image decoder) to decode the image features of the target text data at the Kth scale to obtain a target generated image of the target text data. Optionally, the discrete image decoder may be a multi-scale VQ-VAE model (vector quatization-variational autoencoder, a generative model), that is, it may be a decoder in a VQ-VAE model. Optionally, the target image decoder may also be a module in a target text-graph autoregressive model, that is, the target text-graph autoregressive model may also include a target image decoder, and so on.

[0076] For example, the scale-based autoregressive paradigm can be shown as formula 1.4:

[0077]

[0078] Among them, r1, r2, …, r K It can be a feature map (i.e., image feature) of different scales constrained by a discrete image encoder, that is, it can be an image feature of different scales constrained by the Wensheng graph autoregressive model proposed in an embodiment of the present invention, such as r k It can represent the image features of the target text data at the kth scale. Based on this, we can start from the initial image feature r1 of unit size and gradually generate new image features r2,…,r with larger resolution. K , the image features of the largest scale (i.e., the Kth scale) are generated and the target generated image (also called the generated image corresponding to the specified text prompt) of the target text data (also called the specified text prompt) is obtained through the discrete image decoder D(.).

[0079] The embodiment of the present invention can obtain target text data and call a target text encoding model to perform feature extraction on the target text data to obtain a feature extraction result of the target text data. The feature extraction result of the target text data includes the target pooled text features of the target text data. Then, the target text-image autoregressive model can be called to sequentially determine the image features of the target text data at each of K scales based on the target pooled text features, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale. Furthermore, the image features of the target text data at the Kth scale can be decoded to obtain a target generated image of the target text data. It can be seen that the embodiment of the present invention can conveniently generate a target generated image through the target pooled text features and the target text-image autoregressive model. In other words, the embodiment of the present invention can conveniently generate an image based on text (e.g., conveniently generate a target generated image based on target text data), thereby effectively improving the efficiency of image generation, that is, effectively improving the efficiency of generating images based on text.

[0080] Based on the above description, another embodiment of the present invention further proposes a method for generating text graphs. Accordingly, the method for generating text graphs can be executed by the electronic device (terminal or server) mentioned above; or, the method for generating text graphs can be executed by the terminal and the server together. For the sake of convenience, the following description will be based on an example of an electronic device executing the method for generating text graphs; please refer to Figure 4 The text-generated image method may include the following steps S401-S405:

[0081] S401 : Acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and label text of each training image in the at least one training image.

[0082] Optionally, the electronic device may store a graphic-text matching dataset in its own storage space, and the graphic-text matching dataset may be obtained from its own storage space; or, a graphic-text matching data download link may be obtained, and the graphic-text matching dataset may be downloaded through the graphic-text matching data download link, etc.; this is not limited in this embodiment of the present invention.

[0083] S402, for any training image in the at least one training image, select M training scales from the first P scales of K scales, and call the multi-scale image coding model to perform image coding on the any training image to obtain image features of the any training image at each training scale in the M training scales, where K, P, and M are all positive integers.

[0084] Optionally, P may be less than or equal to K; it should be understood that when P is equal to K, the first P scales may be K scales. Optionally, the electronic device may randomly select M training scales from the first P scales, or may determine at least one specified scale from the first P scales and add the at least one specified scale to the M training scales, so as to use each specified scale as a training scale, thereby selecting M training scales from the first P scales, and so on; this is not limited in the embodiment of the present invention. Optionally, the at least one specified scale may be set according to experience or according to actual needs, and this is not limited in the embodiment of the present invention.

[0085] In an embodiment of the present invention, the multi-scale image coding model may be an encoder in a VQ-VAE model. In other words, the discrete image encoder may be a multi-scale image coding model. Alternatively, the multi-scale image coding model may be a discrete image encoder. Alternatively, the multi-scale image coding model may be a pre-trained model, or may be configured based on experience or actual needs, which is not limited in this embodiment of the present invention.

[0086] S403 , performing feature extraction on the label text of any training image to obtain a feature extraction result of the label text of any training image, wherein the feature extraction result of the label text of any training image includes label pooling text features of the label text of any training image.

[0087] In an embodiment of the present invention, the electronic device may invoke a target text encoding model to perform feature extraction on the label text of any training image, thereby obtaining a feature extraction result for the label text of any training image. Optionally, the feature extraction result for the label text of any training image may further include a refined text feature of the label text of any training image. The refined text feature of a piece of text data may be determined based on a representation vector of each text recognition result in at least one text recognition result corresponding to the corresponding text data.

[0088] S404, calling the initial text-image autoregressive model, determining the image features of the label text of any training image at each training scale based on the label pooling text features of the label text of any training image; and determining the model loss value of any training image based on the image features of any training image at each training scale and the image features of the label text of any training image at each training scale.

[0089] In an embodiment of the present invention, the electronic device may call an initial text-based autoregressive model to sequentially determine the image features of the label text of any training image at each scale based on the label pooled text features of the label text of any training image, until the image features of the label text of any training image at each training scale are obtained. For example, assuming that the M training scales include the second and fifth scales, after obtaining the image features of the label text of any training image at the fifth scale, the image features of the label text of any training image at each training scale may be obtained.

[0090] Optionally, when determining the model loss value at any training image based on the image features of any training image at each training scale and the image features of the label text of any training image at each training scale, for any training scale among the M training scales, the electronic device may determine the label category of each grid block at any training scale based on the image features of any training image at any training scale (i.e., the label category of each grid block in the image features of any training image at any training scale). In an embodiment of the present invention, the features of each grid block in the image features of any training image at any training scale obtained by the multi-scale image coding model can all be benchmark features in the benchmark feature data (also referred to as a code table), and the category of a grid block (such as a label category or a predicted category) can be determined based on the features of the corresponding grid block, and any benchmark feature in the benchmark feature data can correspond to one category. Based on this, the electronic device can determine the label category of any grid block at any training scale based on the features of any grid block in the image features of any training image at any training scale; illustratively, based on the features of any grid block in the image features of any training image at any training scale, the category corresponding to the features of any grid block in the image features of any training image at any training scale can be determined from the baseline feature data, so as to use the category corresponding to the features of any grid block in the image features of any training image at any training scale as the label category of any grid block at any training scale. Optionally, the label category of each grid block at any training scale can also be called a supervisory signal at any training scale, that is, a supervisory signal of the image features of any training image at any training scale.

[0091] Accordingly, the electronic device can determine the predicted category of each grid block at any training scale based on the image features of the label text of any training image at any training scale, which can also be called the predicted category of each grid block in the image features of the label text of any training image at any training scale; optionally, a text graph autoregressive model can also include at least one classification layer (i.e., classification network layer), such as each training scale can correspond to a classification layer or each scale in K scales can correspond to a classification layer, etc.; the embodiment of the present invention is not limited to this. Among them, the classification layer corresponding to a scale can be used to classify and predict the image features of any text data (such as the label text of any training image) at the corresponding scale. Based on this, the electronic device can determine the predicted category of each grid block at any training scale based on the image features of the label text of any training image at any training scale through the classification layer corresponding to any training scale; for any grid block at any training scale, the electronic device can classify and predict the features of any grid block in the image features of the label text of any training image at any training scale through the classification layer corresponding to any training scale, and obtain the predicted category of any grid block at any training scale; illustratively, the features of any grid block in the image features of the label text of any training image at any training scale can be classified and predicted through the classification layer corresponding to any training scale, and the probability vector of any grid block in the image features of the label text of any training image at any training scale can be obtained, thereby taking the category indicated by the maximum probability in the probability vector of any grid block in the image features of the label text of any training image at any training scale as the predicted category of any grid block at any training scale, and so on; the embodiments of the present invention are not limited to this.

[0092] Furthermore, the electronic device may calculate the model loss value at any training scale based on the label category and prediction category of each grid block at any training scale, so as to determine the model loss value at any training image based on the model loss value at each training scale. Optionally, the electronic device may calculate the model loss value at any training scale by using a cross entropy loss function or a mean square error loss function, etc., which is not limited in this embodiment of the present invention. Optionally, the electronic device may calculate the loss value of any grid block at any training scale based on the label category and prediction category of any grid block at any training scale, and perform a weighted summation (such as a mean operation or a summation operation, etc.) on the loss value of each grid block at any training scale to obtain the model loss value at any training scale. Optionally, the electronic device may perform a weighted summation (such as a mean operation or a summation operation, etc.) on the model loss value at each training scale to obtain the model loss value at any training image, so as to determine the model loss value at any training image.

[0093] In other embodiments, the electronic device may also directly determine the model loss value at any training image based on the image features of any training image at each training scale and the image features of the label text of any training image at each training scale, and the present invention is not limited to this. Exemplarily, for any grid block at any training scale (such as the first grid block, etc.), loss calculation (such as mean square error loss calculation or square error loss calculation, etc.) may be performed on the features of any grid block in the image features of any training image at any training scale and the features of any grid block in the image features of the label text of any training image at any training scale to obtain the loss value of any grid block at any training scale, thereby performing weighted summation on the loss values of each grid block at any training scale to obtain the model loss value at any training scale, and then determining the model loss value at any training image based on the model loss value at each training scale, and so on.

[0094] S405, after obtaining the model loss value under each training image, calculate the loss value of the Wensheng graph autoregression model based on the model loss value under each training image; and optimize the model parameters in the initial Wensheng graph autoregression model in the direction of reducing the loss value of the Wensheng graph autoregression model to obtain the initial Wensheng graph autoregression model after model optimization, and determine the target Wensheng graph autoregression model based on the initial Wensheng graph autoregression model after model optimization.

[0095] The target text-image autoregressive model can support generating a target generated image for the target text data; that is, the target text-image autoregressive model can be called to sequentially determine image features of the target text data at each of K scales based on the target pooled text features, and decode the image features of the target text data at the Kth scale to obtain the target generated image for the target text data. Optionally, after obtaining the target text-image autoregressive model, the target text-image autoregressive model can be saved for subsequent use.

[0096] Optionally, the electronic device may perform a weighted summation (such as a mean operation or a summation operation) on the model loss value under each training image to obtain a VANSI graph autoregressive model loss value (also referred to herein as the VANSI graph autoregressive model loss value of the VANSI graph autoregressive model under the first P scales). The initial VANSI graph autoregressive model supports the generation of image features of a text data at each of the first P scales included in the K scales, that is, the VANSI graph autoregressive model after model training of the initial VANSI graph autoregressive model can support the generation of image features of a text data at each of the first P scales included in the K scales.

[0097] Optionally, when determining a target Sentence Graph Autoregressive Model based on the optimized initial Sentence Graph Autoregressive Model, the electronic device may continue to iteratively train the optimized initial Sentence Graph Autoregressive Model until the Sentence Graph Autoregressive Model at the first P scales reaches a convergence condition, thereby determining a first Sentence Graph Autoregressive Model. Specifically, the Sentence Graph Autoregressive Model at the first P scales that reaches the convergence condition may be used as the first Sentence Graph Autoregressive Model. The Sentence Graph Autoregressive Model at the first P scales may refer to a Sentence Graph Autoregressive Model that supports generating image features for text data at each of the first P scales. Optionally, the Sentence Graph Autoregressive Model at the first P scales may be determined to have reached the convergence condition when the number of iterations for the Sentence Graph Autoregressive Model at the first P scales reaches a first preset iteration number threshold. Alternatively, the Sentence Graph Autoregressive Model at the first P scales may be determined to have reached the convergence condition when the loss value of the Sentence Graph Autoregressive Model at the current iteration is less than a first preset loss threshold, and so on. This is not limited in this embodiment of the present invention. Optionally, the first preset iteration number threshold and the first preset loss threshold can be set based on experience or actual needs, and the embodiment of the present invention does not limit this. The first text-to-image autoregressive model can support generating image features of text data at each of the first P scales.

[0098] Optionally, the process of model training the Wensheng graph autoregressive model may also refer to the process of model training the attention model in the Wensheng graph autoregressive model, that is, optimizing the model parameters in the attention model in the Wensheng graph autoregressive model to achieve model training.

[0099] In one embodiment, P is equal to K. In this case, the first Wensheng graph autoregressive model can be used as the target Wensheng graph autoregressive model to determine the target Wensheng graph autoregressive model.

[0100] In another embodiment, P is less than K. In this case, a second text-generated image autoregressive model can be constructed based on the first text-generated image autoregressive model. The second text-generated image autoregressive model supports generating image features of text data at each scale among K scales. Based on this, the model parameters of the attention model in the first text-generated image autoregressive model can be used as the model parameters of the attention model in the second text-generated image autoregressive model, that is, the model parameters of the attention model in the second text-generated image autoregressive model can be initialized by the model parameters of the attention model in the first text-generated image autoregressive model. Furthermore, for K scales, the second text-graph autoregressive model can be iteratively trained (e.g., selecting H training scales from the K scales for iterative training, where H is a positive integer, etc.) until the text-graph autoregressive model at the K scales reaches convergence conditions, thereby determining a target text-graph autoregressive model. The text-graph autoregressive model at the K scales that reaches convergence conditions can be used as the target text-graph autoregressive model; wherein the text-graph autoregressive model at the K scales can refer to a text-graph autoregressive model that supports generating image features for a text data at each of the K scales. Optionally, the process of iteratively training the second text-graph autoregressive model can be the same as the process of iteratively training the initial text-graph autoregressive model, and this embodiment of the present invention will not be further described. Optionally, the convergence condition of the Vincent graph autoregressive model at the K scales can be determined when the number of iterations of the Vincent graph autoregressive model at the K scales reaches a second preset iteration threshold; alternatively, the convergence condition of the Vincent graph autoregressive model at the K scales can be determined when the loss value of the Vincent graph autoregressive model at the current iteration (here, the loss value of the Vincent graph autoregressive model at the K scales) is less than the second preset loss threshold, and so on; this embodiment of the present invention is not limited to this. Optionally, the second preset iteration threshold and the second preset loss threshold can be set based on experience or actual needs, and this embodiment of the present invention is not limited to this.

[0101] Optionally, the electronic device may use the image-text matching dataset to iteratively train the second text-generated image autoregressive model; or, a high-resolution image-text matching dataset may be obtained, and the resolution of the training images in the high-resolution image-text matching dataset (such as 512×512) may be higher than the resolution of the training images in the image-text matching dataset (such as 256×256), so as to use the high-resolution image-text matching dataset to iteratively train the second text-generated image autoregressive model, and so on; the embodiments of the present invention are not limited to this.

[0102] For example, Figure 5As shown, taking P less than K as an example, assuming that decoding the image features at the Pth scale can obtain a 256×256 image, and decoding the image features at the Kth scale can obtain a 512×512 image, then the initial Wensheng graph autoregressive model can be iteratively trained using the Wensheng graph matching dataset to train the attention model in the Wensheng graph autoregressive model, thereby achieving the generation of training 256×256 images; then after completing the iterative training of the Wensheng graph autoregressive model at the first P scales to obtain the first Wensheng graph autoregressive model, the second Wensheng graph autoregressive model can be iteratively trained to achieve the generation of training 512×512 images, so that after completing the iterative training of the Wensheng graph autoregressive model at K scales, the target Wensheng graph autoregressive model is obtained.

[0103] Optionally, the electronic device may also perform image data set preprocessing on at least one training image before performing model training to update at least one training image; optionally, the image data set preprocessing may include but is not limited to at least one of the following: cropping, data cleaning, etc., which is not limited in the embodiment of the present invention.

[0104] Optionally, the electronic device may also input the test text prompt into the target text encoding model to call the target text encoding model to perform feature extraction on the test text prompt, and obtain the feature extraction result of the test text prompt, thereby determining the image features of the test text prompt at the Kth scale through the target text-image autoregressive model, so as to be decoded by the discrete image decoder; based on this, the model performance of the target text-image autoregressive model can be verified by the generated image corresponding to the test text prompt. Based on this, the embodiment of the present invention can implement a stage training strategy that adapts to the resolution (i.e., scale), that is, first modeling the overall image-text matching relationship with a large batch size on a low-resolution image-text dataset, and then further fine-tuning on a high-resolution image-text dataset to enhance the ability to generate high-quality high-resolution images; and the training of this paradigm can be simplified to a multi-stage training paradigm, which can significantly improve its training efficiency and adaptability.

[0105] An embodiment of the present invention can obtain a graph-text matching dataset, which includes at least one training image and label text for each training image in the at least one training image. Then, for any training image in the at least one training image, M training scales can be selected from the first P scales of K scales, and a multi-scale image coding model can be invoked to perform image coding on the training image, obtaining image features for the training image at each of the M training scales, where P and M are both positive integers. Based on this, feature extraction can be performed on the label text of the training image to obtain feature extraction results for the label text of the training image. The feature extraction results for the label text of the training image include label pooled text features of the label text of the training image. Accordingly, an initial graph-text autoregressive model can be invoked to determine the image features of the label text of the training image at each training scale based on the label pooled text features of the label text of the training image. Furthermore, a model loss value for the training image can be determined based on the image features of the training image at each training scale and the image features of the label text of the training image at each training scale. After obtaining the model loss value under each training image, the loss value of the text graph autoregressive model can be calculated based on the model loss value under each training image; and in the direction of reducing the loss value of the text graph autoregressive model, the model parameters in the initial text graph autoregressive model are optimized to obtain the initial text graph autoregressive model after model optimization, and the target text graph autoregressive model is determined based on the initial text graph autoregressive model after model optimization; wherein, the target text graph autoregressive model supports the target generated image for generating target text data. It can be seen that the embodiment of the present invention can improve the model performance of the target text-graph autoregressive model to improve the quality of the generated image; and, the embodiment of the present invention can use the label pooled text features as the initial image features (i.e., the image features at the first scale) to train the text-graph autoregressive model on this basis to predict new feature maps, that is, to train the attention model in the text-graph autoregressive model to predict new feature maps, which can ensure the global consistency of the image and text, that is, provide universal features, help adapt to new text scenarios, and ensure the consistency between text description and visual output (i.e., generated image); based on this, the embodiment of the present invention can conveniently generate the target generated image of the target text data, effectively improve the efficiency of image generation, and effectively improve the quality of the generated image, that is, the embodiment of the present invention can reduce the time required for image reconstruction while maintaining or improving image quality.

[0106] Based on the description of the related embodiments of the above-mentioned method for obtaining a text graph, an embodiment of the present invention further proposes a text graph device, which may be a computer program (including program code) running in an electronic device; Figure 6As shown, the text image device may include a first acquisition unit 601 and a first processing unit 602. The text image device may execute Figure 1 The Wenshengtu method shown, that is, the Wenshengtu device can run the above units:

[0107] A first acquiring unit 601 is used to acquire target text data;

[0108] A first processing unit 602 is configured to call a target text encoding model, perform feature extraction on the target text data, and obtain a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes a target pooled text feature of the target text data;

[0109] The first processing unit 602 is further configured to call a target text-image autoregressive model to sequentially determine image features of the target text data at each of K scales based on the target pooled text features, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale;

[0110] The first processing unit 602 is further configured to decode the image features of the target text data at the K-th scale to obtain a target generated image of the target text data.

[0111] In one embodiment, when the first processing unit 602 calls the target text-image autoregressive model and sequentially determines the image features of the target text data at each of the K scales based on the target pooled text features, it can be specifically configured to:

[0112] Calling a target text-graph autoregressive model, taking the target pooled text features as image features of the target text data at a first scale, and sequentially determining image features of the target text data at each of k-1 scales based on the image features of the target text data at the first scale, k∈[2,K];

[0113] Performing interpolation processing on the image features of the target text data at the k-1th scale to obtain interpolated image features; and determining a first image feature based on the interpolated image features;

[0114] Performing feature extraction on the first image feature to obtain a second image feature; and determining the image feature of the target text data at the kth scale based on the second image feature;

[0115] The image features of the target text data at the k-th scale are used as the image features of the target text data at the k-1-th scale, and the interpolation processing of the image features of the target text data at the k-1-th scale is iteratively performed until the image features of the target text data at the K-th scale are obtained.

[0116] In another embodiment, when determining the first image feature based on the interpolated image feature, the first processing unit 602 may be specifically configured to:

[0117] Determine the target normalized grid scale;

[0118] Calculating, according to the target normalized grid scale and the interpolated image features, position coding features of each grid block at the k-th scale, respectively, where the scale of the interpolated image features is equal to the k-th scale, and the k-th scale is greater than the k-1-th scale;

[0119] A first image feature is determined based on the position encoding feature of each grid block.

[0120] In another embodiment, the feature extraction result of the target text data further includes target refined text features of the target text data, the target text-graph autoregressive model includes a target attention model, and the target attention model includes at least one cross-attention layer; when the first processing unit 602 extracts the first image features to obtain the second image features, it can be specifically used to:

[0121] The target attention model is called, and based on the target refined text features, feature extraction is performed on the first image features to obtain second image features.

[0122] In another embodiment, when determining the image feature of the target text data at the kth scale based on the second image feature, the first processing unit 602 may be specifically configured to:

[0123] Based on the features of any grid block in the second image features, classify and predict the features of any grid block to obtain a probability vector of the any grid block;

[0124] Determining a prediction category of any grid block based on the probability vector of any grid block;

[0125] Based on the predicted category of any one of the grid blocks, a benchmark feature indicated by the predicted category of any one of the grid blocks is determined from the benchmark feature data, and based on the benchmark feature indicated by the predicted category of any one of the grid blocks, the features of any one of the grid blocks in the image features of the target text data at the kth scale are determined to achieve determination of the image features of the target text data at the kth scale.

[0126] Based on the description of the related embodiments of the above-mentioned method for text graphs, the embodiment of the present invention further proposes another text graph device, which can be a computer program (including program code) running in an electronic device; Figure 7 As shown, the text image device may include a second acquisition unit 701 and a second processing unit 702. The text image device may execute Figure 4 The Wenshengtu method shown, that is, the Wenshengtu device can run the above units:

[0127] A second acquiring unit 701 is configured to acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and a label text of each training image in the at least one training image;

[0128] A second processing unit 702 is configured to select, for any training image among the at least one training image, M training scales from the first P scales of the K scales, and invoke a multi-scale image coding model to perform image coding on the training image to obtain image features of the training image at each of the M training scales, where K, P, and M are all positive integers.

[0129] The second processing unit 702 is further configured to perform feature extraction on the label text of any training image to obtain a feature extraction result of the label text of any training image, wherein the feature extraction result of the label text of any training image includes a label pooling text feature of the label text of any training image;

[0130] The second processing unit 702 is further configured to call the initial text-image autoregressive model to determine, based on the label pooled text features of the label text of the any training image, image features of the label text of the any training image at each training scale; and determine, based on the image features of the any training image at each training scale and the image features of the label text of the any training image at each training scale, a model loss value for the any training image;

[0131] The second processing unit 702 is further configured to, after obtaining the model loss value under each training image, calculate a loss value of a text graph autoregressive model based on the model loss value under each training image; and optimize the model parameters in the initial text graph autoregressive model in a direction of reducing the loss value of the text graph autoregressive model to obtain an optimized initial text graph autoregressive model, and determine a target text graph autoregressive model based on the optimized initial text graph autoregressive model; wherein the target text graph autoregressive model supports a target generated image for generating target text data.

[0132] In one embodiment, when determining the model loss value for any training image based on the image features of the any training image at each training scale and the image features of the label text of the any training image at each training scale, the second processing unit 702 may be specifically configured to:

[0133] For any training scale among the M training scales, determining a label category of each grid block at the any training scale based on image features of the any training image at the any training scale;

[0134] Determining a predicted category of each grid block at any training scale based on image features of the label text of any training image at any training scale;

[0135] Based on the label category and the predicted category of each grid block at any training scale, the model loss value at any training scale is calculated to determine the model loss value at any training image based on the model loss value at each training scale.

[0136] In another embodiment, the initial text graph autoregressive model supports generating image features of text data at each of the first P scales included in the K scales, where P is less than K. When the second processing unit 702 determines the target text graph autoregressive model based on the initial text graph autoregressive model after model optimization, it can be specifically used to:

[0137] Continuing to iteratively train the initial text-graph autoregressive model after model optimization until the text-graph autoregressive model at the first P scales reaches a convergence condition, thereby determining a first text-graph autoregressive model; the first text-graph autoregressive model supports generating image features of a text data at each scale in the first P scales;

[0138] Based on the first text-image autoregressive model, a second text-image autoregressive model is constructed, wherein the second text-image autoregressive model supports generating image features of text data at each of the K scales;

[0139] The second Vincent graph autoregressive model is iteratively trained for the K scales until the Vincent graph autoregressive model at the K scales reaches a convergence condition, so as to determine the target Vincent graph autoregressive model.

[0140] According to one embodiment of the present invention, Figure 6 and Figure 7 Each unit in the illustrated text image device can be individually or completely combined into one or several other units to form a structure, or one (or some) of the units can be further divided into multiple functionally smaller units to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present invention, any text image device can also include other units. In actual applications, these functions can also be assisted by other units and can be implemented by the collaboration of multiple units.

[0141] According to another embodiment of the present invention, the program can be executed by running a program on a general electronic device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), and other processing elements and storage elements. Figure 1 A computer program (including program code) for each step involved in the corresponding method shown in Figure 6 The Wensheng diagram device shown in the embodiment of the present invention and the Wensheng diagram method of the embodiment of the present invention can be implemented; accordingly, the Wensheng diagram method can be implemented by running a general electronic device such as a computer including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM) and other processing elements and storage elements. Figure 4 A computer program (including program code) for each step involved in the corresponding method shown in Figure 7 The computer program can be recorded on a computer storage medium, for example, and loaded into the electronic device through the computer storage medium and run therein.

[0142] Based on the description of the above method embodiment and apparatus embodiment, the exemplary embodiments of the present invention further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when executed by the at least one processor, the computer program causes the electronic device to perform a method according to an embodiment of the present invention.

[0143] Exemplary embodiments of the present invention further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present invention.

[0144] An exemplary embodiment of the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor of a computer, the computer is configured to cause the computer to perform a method according to an embodiment of the present invention.

[0145] refer to Figure 8 , a block diagram of an electronic device 800 that can serve as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0146] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0147] Multiple components within electronic device 800 are connected to I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. Input unit 806 can be any type of device capable of inputting information into electronic device 800. Input unit 806 can receive input numeric or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0148] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the Vincent diagram method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the Vincent diagram method in any other appropriate manner (e.g., by means of firmware).

[0149] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0150] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0151] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0153] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0154] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0155] Furthermore, it should be understood that the above disclosure is only a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A method for generating a Wensheng diagram, characterized in that: include: Acquire target text data, call a target text encoding model, perform feature extraction on the target text data, and obtain a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes a target pooled text feature of the target text data; Calling a target text-image autoregressive model, based on the target pooled text features, sequentially determining the image features of the target text data at each of K scales, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale; The image features of the target text data at the K-th scale are decoded to obtain a target generated image of the target text data.

2. The method according to claim 1, characterized in that The calling of the target text-image autoregressive model, based on the target pooled text features, sequentially determining the image features of the target text data at each of K scales, includes: Calling a target text-graph autoregressive model, taking the target pooled text features as image features of the target text data at a first scale, and sequentially determining image features of the target text data at each of k-1 scales based on the image features of the target text data at the first scale, k∈[2,K]; Performing interpolation processing on the image features of the target text data at the k-1th scale to obtain interpolated image features; and determining a first image feature based on the interpolated image features; Performing feature extraction on the first image feature to obtain a second image feature; and determining the image feature of the target text data at the kth scale based on the second image feature; The image features of the target text data at the k-th scale are used as the image features of the target text data at the k-1-th scale, and the interpolation processing of the image features of the target text data at the k-1-th scale is iteratively performed until the image features of the target text data at the K-th scale are obtained.

3. The method according to claim 2, characterized in that The determining of the first image feature based on the interpolated image feature includes: Determine the target normalized grid scale; Calculating, according to the target normalized grid scale and the interpolated image features, position coding features of each grid block at the k-th scale, respectively, where the scale of the interpolated image features is equal to the k-th scale, and the k-th scale is greater than the k-1-th scale; A first image feature is determined based on the position encoding feature of each grid block.

4. The method according to claim 2, characterized in that The feature extraction result of the target text data further includes target refined text features of the target text data, the target text-graph autoregressive model includes a target attention model, and the target attention model includes at least one cross-attention layer; and extracting the first image features to obtain the second image features includes: The target attention model is called, and based on the target refined text features, feature extraction is performed on the first image features to obtain second image features.

5. The method according to claim 2, characterized in that The determining, based on the second image feature, the image feature of the target text data at the kth scale includes: Based on the features of any grid block in the second image features, classify and predict the features of any grid block to obtain a probability vector of the any grid block; Determining a prediction category of any grid block based on the probability vector of any grid block; Based on the predicted category of any one of the grid blocks, a benchmark feature indicated by the predicted category of any one of the grid blocks is determined from the benchmark feature data, and based on the benchmark feature indicated by the predicted category of any one of the grid blocks, the features of any one of the grid blocks in the image features of the target text data at the kth scale are determined to achieve determination of the image features of the target text data at the kth scale.

6. A method for generating a Wensheng diagram, characterized in that: include: Acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and label text of each training image in the at least one training image; For any training image of the at least one training image, select M training scales from the first P scales of the K scales, call the multi-scale image coding model, perform image coding on the any training image, and obtain image features of the any training image at each of the M training scales, where K, P, and M are all positive integers; Performing feature extraction on the label text of any one of the training images to obtain a feature extraction result of the label text of any one of the training images, wherein the feature extraction result of the label text of any one of the training images includes a label pooling text feature of the label text of any one of the training images; Invoking an initial text-image autoregressive model to determine, based on the label pooled text features of the label text of any training image, image features of the label text of any training image at each training scale; and determining a model loss value for any training image based on image features of any training image at each training scale and image features of the label text of any training image at each training scale; After obtaining the model loss value under each training image, the loss value of the text graph autoregression model is calculated based on the model loss value under each training image; and the model parameters in the initial text graph autoregression model are optimized in the direction of reducing the loss value of the text graph autoregression model to obtain the model-optimized initial text graph autoregression model, and the target text graph autoregression model is determined based on the model-optimized initial text graph autoregression model; wherein the target text graph autoregression model supports a target generated image for generating target text data.

7. The method according to claim 6, characterized in that The determining the model loss value of any training image based on the image features of any training image at each training scale and the image features of the label text of any training image at each training scale includes: For any training scale among the M training scales, determining a label category of each grid block at the any training scale based on image features of the any training image at the any training scale; Determining a predicted category of each grid block at any training scale based on image features of the label text of any training image at any training scale; Based on the label category and the predicted category of each grid block at any training scale, the model loss value at any training scale is calculated to determine the model loss value at any training image based on the model loss value at each training scale.

8. The method according to claim 6, characterized in that The initial text graph autoregressive model supports generating image features of text data at each of the first P scales included in the K scales, where P is less than K. The initial text graph autoregressive model optimized based on the model, determining the target text graph autoregressive model, includes: Continuing to iteratively train the initial text-graph autoregressive model after model optimization until the text-graph autoregressive model at the first P scales reaches a convergence condition, thereby determining a first text-graph autoregressive model; the first text-graph autoregressive model supports generating image features of a text data at each scale in the first P scales; Based on the first text-image autoregressive model, a second text-image autoregressive model is constructed, wherein the second text-image autoregressive model supports generating image features of text data at each of the K scales; The second Vincent graph autoregressive model is iteratively trained for the K scales until the Vincent graph autoregressive model at the K scales reaches a convergence condition, so as to determine the target Vincent graph autoregressive model.

9. A Wensheng diagram device, characterized in that: The device comprises: A first acquiring unit, configured to acquire target text data; a first processing unit, configured to call a target text encoding model, perform feature extraction on the target text data, and obtain a feature extraction result of the target text data, wherein the feature extraction result of the target text data includes a target pooled text feature of the target text data; The first processing unit is further configured to call a target text-image autoregressive model, and determine, based on the target pooled text features, the image features of the target text data at each of K scales, where K is a positive integer; wherein the image features of the target text data at the Kth scale are determined based on the image features of the target text data at the K-1th scale; The first processing unit is further configured to decode the image features of the target text data at the K-th scale to obtain a target generated image of the target text data.

10. A Wensheng diagram device, characterized in that: The device comprises: A second acquisition unit is configured to acquire an image-text matching dataset, where the image-text matching dataset includes at least one training image and a label text of each training image in the at least one training image; a second processing unit, configured to, for any training image among the at least one training image, select M training scales from first P scales of the K scales, and call a multi-scale image coding model to perform image coding on the any training image to obtain image features of the any training image at each of the M training scales, where K, P, and M are all positive integers; The second processing unit is further configured to perform feature extraction on the label text of any training image to obtain a feature extraction result of the label text of any training image, wherein the feature extraction result of the label text of any training image includes a label pooling text feature of the label text of any training image; The second processing unit is further configured to call the initial text-image autoregressive model to determine, based on the label pooled text features of the label text of the any training image, the image features of the label text of the any training image at each training scale; and determine, based on the image features of the any training image at each training scale and the image features of the label text of the any training image at each training scale, the model loss value for the any training image; The second processing unit is further configured to, after obtaining the model loss value under each training image, calculate a loss value of a text graph autoregressive model based on the model loss value under each training image; and optimize the model parameters in the initial text graph autoregressive model in a direction of reducing the loss value of the text graph autoregressive model to obtain an optimized initial text graph autoregressive model, and determine a target text graph autoregressive model based on the optimized initial text graph autoregressive model; wherein the target text graph autoregressive model supports a target generated image for generating target text data.

11. An electronic device, characterized in that: include: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 5; or, when executed by the processor, cause the processor to perform the method according to any one of claims 6 to 8.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 5; or, the computer instructions are used to cause a computer to execute the method according to any one of claims 6 to 8.