Image generation method, display method, device, storage medium and program product
By obtaining the description data related to the target text from the database, dimensionality reduction and text expansion processing are performed, prompt text is divided according to the dimensions of the image content, and target images are generated using a big model, which solves the problem of insufficient accuracy of text generation pictures, and achieves higher generation accuracy and user satisfaction.
Patent Information
- Application Number
- CN202510400074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art is difficult to generate accurate text-generated pictures, especially when converting low-dimensional semantic information into high-dimensional image information in cross-modal tasks, there is a problem of insufficient accuracy.
By obtaining the description data related to the target text from the database, performing dimensionality reduction processing, obtaining multiple concept dimensions, and text expansion processing for these dimensions, dividing the prompt text according to the preset image content dimensions, and generating the target image using a large model.
The accuracy and comprehensiveness of the prompt text are improved, thereby improving the accuracy of generating images, meeting users' personalized needs, and continuously improving picture quality through iterative optimization mechanisms.
Smart Images

Figure CN119919525B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of data processing, and in particular, to a method for generating pictures, a display method, a device, a storage medium, and a program product. Background Art
[0002] With the rapid development of data science and deep learning, the research on Artificial Intelligence Generated Content (AIGC), such as text-to-image generation, has become increasingly in-depth and has been applied in many fields. Taking text-to-image generation as an example, it refers to inputting a text description and having a computer generate one or more pictures related to the description, aiming to establish an interpretable mapping between the image space and the text semantic space, and convert the low-dimensional semantic information of the text into high-dimensional image information, which is a challenging cross-modal task.
[0003] Therefore, how to generate accurate pictures has become an urgent problem to be solved currently. Summary of the Invention
[0004] Embodiments of the present application provide a method for generating pictures, a display method, a device, a storage medium, and a program product to improve the accuracy of picture generation.
[0005] In a first aspect, embodiments of the present application provide a method for generating pictures, including:
[0006] Responding to a picture generation request to determine a target text;
[0007] Obtaining description data related to the target text from a database and performing dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text;
[0008] Performing text augmentation processing on each of the multiple concept dimensions to obtain a text database corresponding to each of the multiple concept dimensions, where the text database includes multiple prompt texts;
[0009] Dividing the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to each of the multiple picture content dimensions;
[0010] Selecting a target prompt text from the prompt texts corresponding to each of the multiple picture content dimensions according to picture content generation requirements;
[0011] Generating a target picture based on the target prompt text by using a first large model.
[0012] In a second aspect, embodiments of the present application provide a display method, including:
[0013] Displaying a user interface;
[0014] Based on the input operation on the user interface, determine the target text, generate an image generation request based on the target text, and send the image generation request to the server, so that the server can obtain the description data related to the target text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; perform text expansion processing on the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, where the text databases include multiple prompt texts; divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain the prompt texts corresponding to the multiple picture content dimensions respectively;
[0015] Obtain and display the multiple picture content dimensions and the corresponding prompt texts in the user interface;
[0016] Based on the user operation on the user interface, generate an operation request and send it to the server, so that the server can determine the picture content generation requirements based on the user operation, and select the target prompt text from the prompt texts corresponding to the multiple picture content dimensions respectively according to the picture content generation requirements; generate a target picture based on the target prompt text by using a first large model;
[0017] Obtain and display the target picture in the user interface.
[0018] In a third aspect, an embodiment of the present application provides an image generation method, including:
[0019] In response to the image generation request, determine the target concept text of mental health;
[0020] Obtain the description data related to the target concept text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target concept text;
[0021] Perform text expansion processing on the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, where the text databases include multiple prompt texts;
[0022] Divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain the prompt texts corresponding to the multiple picture content dimensions respectively;
[0023] Select the target prompt text from the prompt texts corresponding to the multiple picture content dimensions respectively according to the picture content generation requirements;
[0024] Generate a target picture based on the target prompt text by using a first large model; wherein, the target picture is used to analyze the mental health status of the user.
[0025] In a fourth aspect, an embodiment of the present application provides a computing device, including a storage component and a processing component; the storage component stores one or more computer program instructions, and the computer program instructions are called and executed by the processing component, and the processing component executes the one or more computer program instructions to implement the picture generation method described in the first aspect, or the display method described in the second aspect, or the picture generation method described in the third aspect.
[0026] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and the computer program is executed by a computer to implement the picture generation method described in the first aspect, or the display method described in the second aspect, or the picture generation method described in the third aspect.
[0027] In a sixth aspect, an embodiment of the present application provides a computer program product storing a computer program, and the computer program, when executed by a computer, implements the picture generation method described in the first aspect, or the display method described in the second aspect, or the picture generation method described in the third aspect.
[0028] In the solution of the embodiment of the present application, the server can respond to a picture generation request, determine a target text, obtain description data related to the target text from a database and perform dimensionality reduction processing to obtain multiple concept dimensions corresponding to the target text, perform text expansion processing on each of the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, where each database may include multiple prompt texts, divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to the multiple picture content dimensions respectively, and select target prompt texts from the prompt texts corresponding to the multiple picture content dimensions respectively according to the picture content generation requirements, so as to generate a target picture based on the target prompt text using a first large model. By obtaining description data related to the target text from the database and performing dimensionality reduction processing on it to obtain multiple concept dimensions corresponding to the target text, and performing text expansion processing on each of the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, which include multiple prompt texts, the accuracy and comprehensiveness of the prompt texts are improved, and further the accuracy of the target prompt text selected therefrom and the target picture generated based on the target prompt text are improved.
[0029] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0031] Figure 1 Fig. 4 shows a flowchart of an embodiment of a picture generation method provided by the present application;
[0032] Figure 2 Fig. 8 shows a flowchart of an embodiment of a picture evaluation method provided by the present application;
[0033] Figure 3 Fig. 12 shows a flowchart of an embodiment of a display method provided by the present application;
[0034] Figure 4 Fig. 16 shows a system architecture diagram in an actual application;
[0035] Figure 5 Fig. 20 shows a schematic structural diagram of an embodiment of a picture generation device provided by the present application;
[0036] Figure 6 Fig. 24 shows a schematic structural diagram of an embodiment of a display device provided by the present application;
[0037] Figure 7 Fig. 28 shows a schematic structural diagram of an embodiment of a computing device provided by the present application. Detailed implementation manners
[0038] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application.
[0039] In some processes described in the specification, claims and the above accompanying drawings of the present application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, units, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0040] As described in the foregoing background art section, in order to improve the accuracy of generating pictures from text, the inventors have proposed the technical solution of the present application, including determining a target text in response to a picture generation request; obtaining description data related to the target text from a database, and performing dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; respectively performing text expansion processing on the multiple concept dimensions to obtain text databases respectively corresponding to the multiple concept dimensions, where each text database includes multiple prompt texts; dividing the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts respectively corresponding to the multiple picture content dimensions; selecting a target prompt text from the prompt texts respectively corresponding to the multiple picture content dimensions according to the picture content generation requirements; and generating a target picture based on the target prompt text by using a first large model.
[0041] By obtaining description data related to the target text from the database and performing dimensionality reduction processing on it to obtain multiple concept dimensions corresponding to the target text, and respectively performing text expansion processing on the multiple concept dimensions to obtain corresponding text databases each including multiple prompt texts, the accuracy and comprehensiveness of the prompt texts are improved, thereby improving the accuracy of the target prompt text selected therefrom and the target picture generated based on the target prompt text.
[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0043] The technical solution of the embodiment of the present application can be applied to a system architecture including a client and a server, and a connection is established between the client and the server through a network. The network provides a medium for the communication link between the client and the server. The network can include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0044] The client can interact with the server through the network to send a picture generation request or receive prompt texts, target pictures, etc.
[0045] Among them, the client can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The client can be deployed in an electronic device and needs to rely on the device or certain apps in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. For the sake of easy understanding, various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0046] The server side can include servers that provide various services, such as a server that responds to a picture generation request sent by the client, a server that generates a target picture using a large model, etc.
[0047] It should be noted that the server side can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0048] It should be noted that the picture generation method provided in the embodiments of this application is generally executed by the server side, and the display method is generally executed by the client side. Correspondingly, the picture generation device is generally deployed in the server side, and the display device is generally deployed in the client side. However, in other embodiments of this application, the client can also have a similar function to the server side, so as to execute the picture generation method provided in the embodiments of this application. In other embodiments, the picture generation method provided in the embodiments of this application can also be jointly executed by the client and the server side.
[0049] It should be noted that the embodiments of this application may involve the use of user data. In actual applications, user-specific personal data can be used in the solutions described in this article within the scope permitted by applicable laws and regulations, provided that the requirements of the applicable laws and regulations of the country where it is located are met (for example, the user clearly consents, and the user is effectively notified, etc.).
[0050] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0051] It should be noted that the technical solution of the embodiment of this application is applicable to a network virtual environment. The described user generally refers to a "virtual user". A real user can register a user account on the server through registration to obtain a user identity in the network environment.
[0052] As Figure 1 shown, it is a flowchart of an embodiment of a picture generation method provided by this application, which can be executed by the server. This method can include the following steps:
[0053] 101: In response to a picture generation request, determine the target text.
[0054] Among them, the picture generation request can be generated and sent by the client in response to a user operation. The picture generation request includes the target text. The target text can include various forms such as words, phrases, sentences, etc. In a practical application, the target text can be, for example, a mental health concept word, such as social anxiety disorder.
[0055] 102: Obtain description data related to the target text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text.
[0056] Among them, the database can refer to a general database in the technical field corresponding to the target text, and can include a search engine, etc. The description data related to the target text can include description data such as the definition, structure, and characteristics of the target text. Specifically, a preset data acquisition algorithm can be used to obtain the description data from the database. The data acquisition algorithm can include, for example, the Scrapy crawler framework (a fast, high-level screen scraping and web scraping framework suitable for Python, used to scrape web sites and extract structured data from pages), a data scraping algorithm combining BeautifulSoup (a library for parsing HTML (Hyper Text Markup Language) and XML (Extensible Markup Language) documents) and Requests (a library for sending HTTP requests), etc.
[0057] Optionally, various forms of descriptive data related to the target text, such as text, scales, etc., can be obtained from multiple databases. Taking the target text being social anxiety disorder as an example, various forms of descriptive data related to social anxiety disorder, such as definitions, scales, clinical information, and film and television works, can be obtained from multiple databases.
[0058] After obtaining the descriptive data, a preset dimensionality reduction algorithm can be used to perform dimensionality reduction processing on the descriptive data to obtain multiple concept dimensions corresponding to the target text. The dimensionality reduction algorithm can, for example, include PCA technology (principal components analysis), t-SNE (t-distributed Stochastic Neighbor Embedding, a non-linear dimensionality reduction technology), UMAP (Uniform Manifold Approximation and Projection, a non-linear dimensionality reduction technology), etc., and this application does not limit this. Optionally, before performing dimensionality reduction processing, the descriptive data can be preprocessed. Taking the descriptive data in text form as an example, preprocessing operations such as word segmentation, stop word removal, and lemmatization can be performed on it, and then the dimensionality reduction algorithm is used to reduce the feature dimension to simplify the data and improve efficiency.
[0059] The concept dimensions can be used to represent different aspects of the target text. Taking the target text being social anxiety disorder and the obtained descriptive data including descriptive data such as the definition of social anxiety, scales, clinical information, and film and television works as an example, the multiple concept dimensions obtained after dimensionality reduction processing can include three concept dimensions: emotional expression, social dominance, and eye contact. Social dominance can refer to the key position and environment of a certain group or individual in social relations within a certain social structure.
[0060] 103: Perform text expansion processing on each of the multiple concept dimensions to obtain text databases corresponding to each of the multiple concept dimensions. The text databases include multiple prompt texts.
[0061] Specifically, based on the multiple concept dimensions, for each concept dimension, a preset text analysis algorithm and text rewriting algorithm can be used to perform text expansion processing to generate multiple prompt texts corresponding to this concept dimension. These multiple prompt texts constitute the text database corresponding to this concept dimension. The prompt texts can include various forms such as characters, words, and sentences. In an actual application, the prompt texts can be in the form of words.
[0062] Among them, the text analysis algorithms can include, for example, NLTK (Natural Language Toolkit, a Python library providing a large number of natural language processing functions), SpaCy (a Python library providing efficient natural language processing functions such as semantic analysis and dependency parsing), Scikit-learn (a Python library providing a large number of machine learning and deep learning functions), Word2Vec (a neural network-based word vector representation model), TF-IDF (Term Frequency-Inverse Document Frequency, a statistical method for text mining and information retrieval), VADER (Valence Aware Dictionary and sEntimentReasoner, a lexicon- and rule-based sentiment analysis tool), BERT for Sentiment Analysis (a sentiment analysis method based on the BERT (Bidirectional Encoder Representations from Transformers) model), etc. The text expansion algorithms can include, for example, a similar word expansion algorithm based on word vectors, an expansion algorithm based on a synonym dictionary, a dynamic generation algorithm based on context such as the Transformer model, an expansion algorithm based on data augmentation techniques, a fine-tuning algorithm based on a pre-trained language model, etc. The present application does not limit this.
[0063] 104: Divide the multiple prompt texts according to a plurality of preset picture content dimensions to obtain the prompt texts respectively corresponding to the multiple picture content dimensions.
[0064] The picture content dimensions can be used to represent different aspects of the picture content. In this embodiment, the picture content dimensions can include three preset dimensions: spatial environment, person object, and action. According to the preset picture content dimensions, the above-generated multiple prompt texts can be divided by using a preset text recognition algorithm to obtain the prompt texts respectively corresponding to the multiple picture content dimensions. The text recognition algorithms can include, for example, BoW (Bag of Words), Word2Vec, etc. The present application does not limit this.
[0065] 105: Select a target prompt text from the prompt texts respectively corresponding to the multiple picture content dimensions according to the picture content generation requirements.
[0066] Among them, the requirements for generating picture content can be requirements for generating pictures in terms of the dimensions of picture content, which are used to indicate the target pictures for generating content reflecting the specified dimensions of picture content. For example, the requirements for generating picture content can be the dimension of specifying the spatial environment, that is to say, indicating the target pictures for generating content reflecting the spatial environment. At this time, the target prompt text can be selected from the prompt text corresponding to the spatial environment dimension.
[0067] The requirements for generating picture content can be automatically set by the server, and of course, it also supports user settings to meet user needs, which will be described in subsequent embodiments.
[0068] 106: Based on the target prompt text, use the first large model to generate the target picture.
[0069] The large models (Large Model, also known as the foundation model, that is, FoundationModel) involved in the embodiments of the present application, including the first large model and the second large model, the third large model, etc. mentioned below, refer to machine learning models with a large number of parameters and complex structures, which can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc., and are a type of AI (Artificial Intelligence) model. The large model can be implemented using LLM (Large Language Model) or MLM (Multimodal Large Model), etc., and the present application does not limit this.
[0070] Specifically, in this embodiment, the first large model can be an image generation large model such as Dalle-2 (a text-to-image system), Stable Diffusion XL (an image generation model), etc. Inputting the target prompt text into the first large model can generate one or more target pictures.
[0071] In this embodiment, by obtaining the description data related to the target text from the database and performing dimensionality reduction processing on it, multiple concept dimensions corresponding to the target text are obtained, and text expansion processing is respectively performed on multiple concept dimensions to obtain a text database containing multiple prompt texts, thereby improving the accuracy and comprehensiveness of the prompt text, and further improving the accuracy of the target prompt text selected therefrom and the target pictures generated based on the target prompt text.
[0072] In order to further improve the accuracy of the pictures, in some embodiments, the above method may further include:
[0073] Extract the feature vector of the target text;
[0074] Based on the feature vector, find at least one associated text associated with the target text in the target semantic network, and the associated weights respectively corresponding to the at least one associated text; wherein, the target semantic network is pre-constructed based on the dataset in the target field where the target text is located;
[0075] Add at least one associated text to the text database in descending order of the associated weights as hint texts.
[0076] Among them, the target semantic network can be pre-constructed based on the dataset in the target field where the target text is located. For example, if the target text is social anxiety disorder, the target field can be the mental health field, and the corresponding dataset can be collected using data crawling technologies such as the Scrapy crawler framework, and can include multiple mental health concept texts and their respective definitions, symptom manifestations, etc.
[0077] Specifically, natural language processing (NLP) technology can be used to parse the target text to extract the feature vector of the target text. Of course, a feature extraction model or the like can also be used to extract the feature vector of the target text, and there is no limitation on this. Then, the extracted feature vector is input into the target semantic network, and algorithms such as graph neural network (GNN) and knowledge graph technology are used to calculate the relationship between the target text and other texts in the target semantic network, such as calculating the semantic similarity and co-occurrence frequency between the target text and other texts, so as to find at least one associated text associated with the target text and the respective associated weights. The associated weight can represent the degree of association between the target text and the associated text. The higher the associated weight, the deeper the degree of association. In descending order of the associated weights, the at least one associated text can be added to the text database as hint texts. For example, if the target text is social anxiety disorder, the at least one associated text may include anxiety disorder, emotional tension, sweating, etc., and then the associated text can be used as a hint text.
[0078] By constructing the target semantic network in the target field where the target text is located, the association between different texts is established, so that for a new target text, other texts associated with it can be found and the degree of association can be calculated, realizing using the associated text as a hint text, improving the richness and accuracy of the hint text, and further improving the accuracy of image generation.
[0079] In some embodiments, the above method may further include:
[0080] Send at least one associated text and the first selection hint information to the client;
[0081] Determine the target associated text based on a selection operation for at least one associated text.
[0082] At this time, according to the requirements for generating picture content, selecting the target prompt text from the prompt texts corresponding to multiple picture content dimensions respectively may include:
[0083] According to the requirements for generating picture content, select the target prompt text including the target associated text from the prompt texts corresponding to multiple picture content dimensions respectively. Thus, on the basis of enriching the prompt text, the personalized needs of users can be met.
[0084] In practical applications, the above target semantic network can also be updated. Therefore, in some embodiments, the above method may further include:
[0085] Send the target prompt text and updated prompt information to the client;
[0086] Based on the confirmation operation for the updated prompt information, when the update condition is met, update the target semantic network based on the target prompt text.
[0087] Among them, the updated prompt information can be used to prompt the user whether to accept the update of the target semantic network based on the target prompt text. Based on the confirmation operation triggered by the user, the target semantic network can be updated using the target prompt text. Specifically, an update condition can be set in advance. For example, when the number of prompt texts for update reaches a threshold, the target semantic network is updated, and this threshold can be set according to actual needs, such as 1000, 10000, etc.
[0088] Thus, the dynamic update of the target semantic network is realized, improving the comprehensiveness and accuracy of the target semantic network to better adapt to the new target text and providing more comprehensive support for the search of associated texts.
[0089] In order to further improve the accuracy of the picture, in some embodiments, the above method may further include:
[0090] Obtain the prompt text corresponding to the historical text related to the target text from the target platform.
[0091] In this embodiment, a target platform can be provided to support multiple users to upload and share pictures with high accuracy and meeting their own needs based on text. For the sake of easy description, the pictures here are called historical pictures, and the corresponding texts are called historical texts. The target platform can store the historical texts, historical pictures corresponding to the historical texts, and prompt texts uploaded by multiple users respectively. The prompt texts can be generated based on the historical texts or historical pictures.
[0092] Based on this, the hint text corresponding to the historical text related to the target text can be obtained from the target platform. For example, the preset text similarity algorithm can be used to determine the historical text that meets the requirements of similarity with the target text, and the hint text of this historical text can be obtained. Then, continue to execute according to the steps of performing text expansion processing on multiple concept dimensions respectively to obtain text databases corresponding to multiple concept dimensions.
[0093] By providing a target platform to support multi-user information sharing, the hint text corresponding to the historical text related to the target text can be obtained from it, further improving the accuracy of the hint text, thereby improving the accuracy of the target hint text and the target picture.
[0094] To ensure that the generated target picture meets the personalized needs of users, the picture generation requirements can be specified by the user. In some embodiments, the above method may further include:
[0095] Send multiple picture content dimensions and the corresponding hint text to the client for display in the user interface;
[0096] Based on the user operations in the user interface, determine the picture content generation requirements.
[0097] Among them, the user operations may include input, selection, sliding operations, etc., without limitation.
[0098] Optionally, based on the user operations in the user interface, determining the picture content generation requirements may include:
[0099] Based on the weight values respectively set by the user for multiple picture content dimensions, determine the picture content generation requirements.
[0100] Among them, the sum of the weight values corresponding to multiple picture content dimensions is the value 1. The picture content dimension with a larger weight value indicates that the priority or importance of reflecting the picture content is higher than that of the picture content dimension with a smaller weight value. For example, if the weight values set for the three picture content dimensions of the space scene, character image, and action are 70%, 20%, and 10% respectively, then the picture content generation requirements can be a picture that highly reflects the space scene content, moderately reflects the character image content, and lowly reflects the action content.
[0101] By selecting the target hint text according to the picture content generation requirements specified by the user and generating a target picture that meets the requirements, the personalized needs of the user are met and the user experience is improved.
[0102] To further meet the user's needs, in some embodiments, the above method may further include:
[0103] Send multiple candidate picture styles and the second selection hint information to the client;
[0104] Based on the selection operation for any candidate picture style, determine the target picture style.
[0105] Among them, the picture style can refer to the visual style of the picture, which can include cartoon style, abstract style, realistic style, etc., and can be preset according to actual needs.
[0106] Optionally, based on the user's input operation, the picture style input by the user can be used as the target picture style.
[0107] On this basis, using the first large model to generate the target picture can include:
[0108] According to the target picture style, use the first large model to generate the target picture.
[0109] Specifically, machine learning algorithms such as style transfer technology (a technology that combines the content of one picture with the style of another picture using deep learning algorithms) and GAN (Generative Adversarial Networks) can be used to apply the target picture style selected or input by the user to the generated target picture to ensure that the target picture meets the user's aesthetic preferences.
[0110] In practical applications, the generated pictures can also be evaluated. As Figure 2 shown, it is a flowchart of an embodiment of a picture evaluation method provided by this application, which can be executed by the server. This method can include the following steps:
[0111] 2001: In response to the picture generation request, determine the target text.
[0112] 2002: Obtain the description data related to the target text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text.
[0113] 2003: Perform text expansion processing on each of the multiple concept dimensions to obtain text databases corresponding to each of the multiple concept dimensions. The text databases include multiple prompt texts.
[0114] 2004: Divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to each of the multiple picture content dimensions.
[0115] 2005: Select the target prompt text from the prompt texts corresponding to each of the multiple picture content dimensions according to the picture content generation requirements.
[0116] 2006: Based on the target prompt text, use the first large model to generate the target picture.
[0117] The implementation processes of steps 2001 to 2006 are the same as Figure 1 the implementation of steps 101 to 106 in the illustrated embodiment, and will not be elaborated here.
[0118] 2007: Use the second largest model to generate the descriptive text corresponding to the target picture.
[0119] Among them, the second largest model can be a large text generation model such as CogVLM (a powerful open-source vision-language foundation model, which is an image-to-text model), ClipCap (a deep learning-based image description model), BLIP-2 (a multimodal vision-text large language model), etc. Inputting the target picture into the second largest model can generate the descriptive text of the target picture, and the descriptive text can be in various forms such as words, phrases, sentences, etc. In a practical application, the descriptive text can be in the form of phrases.
[0120] 2008: Calculate the correlation score corresponding to the target picture and calculate the logical score corresponding to the target picture based on the descriptive text, the target text, and the target prompt text.
[0121] Specifically, based on the descriptive text, the target text, and the target prompt text, a preset semantic similarity model such as SimCSE (Simple Contrastive Learning of Sentence Embeddings, a contrastive learning model for sentence embeddings) can be used to calculate the correlation score corresponding to the target picture.
[0122] Based on the target picture, a preset vision-language model such as CLIP (Contrastive Language-Image Pre-Training, a multimodal pre-trained neural network) can be used to calculate the logical score of the target picture. Among them, this logical score can be used to characterize the degree to which the target picture conforms to the visual logic of the real world.
[0123] 2009: Evaluate the target picture based on the logical score and the correlation score.
[0124] 2010: Based on the evaluation result, judge whether the target picture meets the evaluation requirements. If the judgment result is no, adjust the text database and the requirements for the picture content, and return to the operation in step 2004 to continue execution; if the judgment result is yes, perform the operation in step 2011.
[0125] 2011: Obtain the target picture that meets the evaluation requirements.
[0126] Among them, the logical score and the correlation score can be normalized and then weighted to calculate the evaluation score of the target picture, and the evaluation score is compared with the preset evaluation score to determine whether the target picture meets the evaluation requirements. For example, when the evaluation score is greater than or equal to the preset evaluation score, it is determined to meet the evaluation requirements; otherwise, it does not meet the evaluation requirements.
[0127] When the evaluation requirements are not met, the text database and the picture content generation requirements can be adjusted. Among them, the adjustment of the text database can include operations such as adding, deleting, and replacing the prompt text. The adjustment of the picture content generation requirements can include adjusting the weight values corresponding to multiple picture content dimensions. Then, the steps of step 2004 can be returned to continue execution until a target picture that meets the evaluation requirements is obtained.
[0128] Optionally, during the above evaluation iteration process, the target prompt text and the updated prompt information selected in each iteration can be sent to the client, and based on the user's confirmation operation for the updated prompt information, when the update conditions are met, the target semantic network is updated based on the target prompt text to further improve the accuracy of the target semantic network.
[0129] In this embodiment, by adopting an evaluation method that combines logic and relevance, and through an automated iterative optimization mechanism, it is ensured that the generated pictures are continuously improved in quality and effect, further improving the accuracy of the pictures.
[0130] In some embodiments, the above method may further include:
[0131] Sending the target picture that meets the evaluation requirements to the client;
[0132] Obtaining user response information for the target picture;
[0133] According to multiple preset scoring dimensions, based on the user response information, using the third large model to score the target picture that meets the evaluation requirements to obtain a scoring result.
[0134] Among them, the user response information can be obtained based on various forms such as text, voice, gestures, and physiological signals. Taking the target text as social anxiety disorder as an example, the multiple preset scoring dimensions can include the user's social anxiety score, valence, arousal, immersion, and picture-text consistency. The social anxiety score can refer to the score obtained by the user in a social anxiety medical test, which can be obtained from the data stored in the corresponding medical database. The valence can refer to the efficacy unit of a substance that causes a biological reaction. The arousal can refer to the degree of physiological and cognitive arousal brought about by the user's emotion when viewing the picture. The immersion can refer to the degree of investment of the user when viewing the picture. The picture-text consistency can refer to the consistency between the picture and the text.
[0135] In this embodiment, the third large model can be a large artificial neural network model. For example, a picture scoring system can be constructed using ResNet50 (Deep Residual Network) as the training model. According to multiple preset scoring dimensions, based on the user response information, an automated score can be given to the target picture that meets the evaluation requirements, and the scoring result of the target picture can be obtained.
[0136] It should be noted that in the foregoing embodiment, the evaluation of the picture is realized by taking the logic of the picture itself and the relevance to the target text as the evaluation dimensions. In this embodiment, it is realized by taking the preset dimensions related to the user's response to the picture as the scoring dimensions.
[0137] In some embodiments, the above method may further include:
[0138] Perform dimensionality reduction processing on the target prompt text and description text corresponding to the target picture that meets the evaluation requirements, and compare the dimensionality-reduced target prompt text with the description text. Adjust the picture content dimension according to the comparison result.
[0139] Among them, the dimensionality reduction processing algorithm has been described in the foregoing embodiment and will not be elaborated here. Adjusting the picture content dimension can include, for example, adding, deleting, replacing, etc. For example, based on the comparison result, it is determined that the generated picture has nothing to do with the picture content dimension of the character image but is related to the brightness of the picture. At this time, the picture content dimension of the character image can be deleted, and the picture content dimension of the brightness can be added.
[0140] At this time, scoring the target picture that meets the evaluation requirements using the third large model based on the user response information according to multiple preset scoring dimensions may include:
[0141] Score the target picture that meets the evaluation requirements using the third large model based on the user response information according to multiple preset scoring dimensions and the adjusted picture content dimension.
[0142] Specifically, according to the adjusted picture content dimension, for each picture content dimension, score the target picture that meets the evaluation requirements using the third large model according to multiple preset scoring dimensions based on the user response information, so as to improve the accuracy of the scoring result.
[0143] The above one or more embodiments have described the technical solution of the present application from the perspective of the server. Next, the technical solution of the present application will be described from the perspective of the client. As Figure 3 shown, it is a flowchart of an embodiment of a display method provided by the present application, which can be executed by the client and may include the following steps;
[0144] 301: Display a user interface.
[0145] 302: Based on an input operation on the user interface, determine a target text, generate an image generation request based on the target text, and send the image generation request to the server. The server is used to obtain description data related to the target text from the database, perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; perform text expansion processing on each of the multiple concept dimensions to obtain text databases corresponding to each of the multiple concept dimensions, where each text database includes multiple prompt texts; divide the multiple prompt texts according to a preset multiple image content dimensions to obtain prompt texts corresponding to each of the multiple image content dimensions.
[0146] 303: Obtain and display multiple image content dimensions and corresponding prompt texts in the user interface.
[0147] 304: Based on a user operation on the user interface, generate an operation request and send it to the server. The server is used to determine an image content generation requirement based on the user operation, and select a target prompt text from the prompt texts corresponding to each of the multiple image content dimensions according to the image content generation requirement; generate a target image using a first large model based on the target prompt text.
[0148] 305: Obtain and display the target image in the user interface.
[0149] In this embodiment, by generating an operation request based on a user operation on the user interface and sending it to the server, the server determines an image content generation requirement based on the user operation, selects a target prompt text according to the image content generation requirement, and generates a target image, which improves the accuracy of the image and meets the personalized needs of users.
[0150] In some embodiments, the above method may further include:
[0151] Obtain and display the target prompt text in the user interface;
[0152] Based on an upload operation, upload the target text, the target image, and the target prompt text to a target platform.
[0153] In some embodiments, the above method may further include:
[0154] Based on a save operation for at least one prompt text, save at least one prompt text locally.
[0155] By displaying the target prompt text in the user interface, it supports the user to upload the target prompt text, the corresponding target text, and the target image to the target platform to achieve multi-user information sharing, and supports the user to save them locally.
[0156] Taking the concept in the field of mental health as an example of the target text, the technical solution of the present application will be described. An embodiment of the present application also provides a method for generating pictures, which may include:
[0157] In response to a picture generation request, determine the target concept text of mental health;
[0158] Obtain the description data related to the target concept text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target concept text;
[0159] Perform text expansion processing on multiple concept dimensions respectively to obtain text databases corresponding to multiple concept dimensions respectively, and the text databases include multiple prompt texts;
[0160] Divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to multiple picture content dimensions respectively;
[0161] According to the picture content generation requirements, select the target prompt text from the prompt texts corresponding to multiple picture content dimensions respectively;
[0162] Based on the target prompt text, use the first large model to generate a target picture; wherein, the target picture can be used to analyze the mental health status of the user.
[0163] In this embodiment, the target concept text may refer to a concept in the field of mental health, such as social anxiety disorder, well-being, stress, etc. By obtaining the description data related to the target concept text from the database and performing dimensionality reduction processing on it, multiple concept dimensions corresponding to the target concept text are obtained. Text expansion processing is performed on multiple concept dimensions respectively to obtain corresponding text databases containing multiple prompt texts, thereby improving the accuracy and comprehensiveness of the prompt texts, and further improving the accuracy of the target prompt text selected therefrom and the target picture generated based on the target prompt text.
[0164] In some embodiments, the above method may further include:
[0165] Send multiple picture content dimensions and corresponding prompt texts to the client for display in the user interface;
[0166] Based on the user operation in the user interface, determine the picture content generation requirements.
[0167] In some embodiments, determining the picture content generation requirements based on the user operation in the user interface may include:
[0168] Based on the weight values set by the user for multiple picture content dimensions respectively, determine the picture content generation requirements.
[0169] In some embodiments, the above method may further include:
[0170] Extract the feature vector of the target concept text;
[0171] Based on the feature vector, find at least one associated text associated with the target concept text in the target semantic network, and the associated weights respectively corresponding to the at least one associated text; wherein, the target semantic network is pre-constructed based on a dataset in the field of mental health;
[0172] Add at least one associated text to the text database in descending order of the associated weights as the prompt text.
[0173] In some embodiments, the above method may further include:
[0174] Send at least one associated text and the first selection prompt information to the client;
[0175] Based on the selection operation for at least one associated text, determine the target associated text;
[0176] Selecting the target prompt text from the prompt texts respectively corresponding to multiple picture content dimensions according to the picture content generation requirements may include:
[0177] Select the target prompt text including the target associated text from the prompt texts respectively corresponding to multiple picture content dimensions according to the picture content generation requirements.
[0178] In some embodiments, the above method may further include:
[0179] Send the target prompt text and the update prompt information to the client;
[0180] Based on the confirmation operation for the update prompt information, when the update condition is met, update the target semantic network based on the target prompt text.
[0181] In some embodiments, the above method may further include:
[0182] Use the second large model to generate the description text corresponding to the target picture;
[0183] Based on the description text, the target concept text, and the target prompt text, calculate the association score corresponding to the target picture;
[0184] Calculate the logic score corresponding to the target picture;
[0185] Evaluate the target picture based on the logic score and the association score;
[0186] When it is determined that the target image does not meet the evaluation requirements based on the evaluation results, the text database and the requirements for the image content are adjusted, and the step of dividing multiple prompt texts according to multiple preset image content dimensions and obtaining the prompt texts corresponding to the multiple image content dimensions is continued until the generated target image meets the evaluation requirements.
[0187] In some embodiments, the above method may further include:
[0188] Obtain the prompt text corresponding to the historical text related to the target concept text from the target platform; wherein, the target platform stores the historical concept texts uploaded by multiple users respectively, the historical images corresponding to the historical concept texts, and the prompt texts.
[0189] In some embodiments, the above method may further include:
[0190] Send multiple candidate image styles and the second selection prompt information to the client;
[0191] Based on the selection operation for any one of the candidate image styles, determine the target image style;
[0192] Using the first large model to generate the target image may include:
[0193] Generate the target image using the first large model according to the target image style.
[0194] In some embodiments, the above method may further include:
[0195] Send the target image that meets the evaluation requirements to the client;
[0196] Obtain the user response information for the target image;
[0197] According to multiple preset scoring dimensions, based on the user response information, use the third large model to score the target image that meets the evaluation requirements to obtain a scoring result.
[0198] In some embodiments, the above method may further include:
[0199] Perform dimensionality reduction processing on the target prompt text and the description text corresponding to the target image that meets the evaluation requirements, and compare the dimensionality-reduced target prompt text with the description text, and adjust the image content dimensions according to the comparison result;
[0200] According to multiple preset scoring dimensions, based on the user response information, using the third large model to score the target image that meets the evaluation requirements may include:
[0201] Based on multiple preset scoring dimensions and the adjusted picture content dimensions, and using the user response information, the third large model is used to score the target pictures that meet the evaluation requirements.
[0202] For ease of understanding, the following takes the target text as a concept in the field of mental health as an example, and combines Figure 4 with the system architecture diagram shown in the following to describe the technical solution of this application in detail. As Figure 4 shown, the system may include a client 41 and a server 42. It should be noted that Figure 4 the presented client 41 and server 42 are only exemplary descriptions and do not limit their implementation forms. For convenience of viewing, the illustration of the client is represented by a device image, such as the computer image in the figure, and the illustration of the server is represented by a device image, such as the server image in the figure.
[0203] Among them, the server 42 may include a data acquisition module 421, a data processing module 422, a prompt text processing module 423, a picture content generation requirement setting module 424, a picture generation module 425, a picture evaluation module 426, and a picture effect evaluation module 427.
[0204] Specifically, the client 41 may, in response to a user input operation, determine that the mental health concept input by the user is, for example, social anxiety disorder, and generate a picture generation request and send it to the server 42.
[0205] The data acquisition module 421 in the server 42 may use data acquisition algorithms such as the Scrapy crawler framework to obtain description data related to social anxiety disorder from a general database in the field of mental health, such as social anxiety scales, clinical information of patients with social anxiety disorder, etc.
[0206] The data processing module 422 may use dimensionality reduction processing algorithms such as PCA technology to perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to social anxiety disorder. For example, it may include three concept dimensions: emotional expression, social dominance, and eye contact. And text analysis algorithms such as NLTK and text expansion algorithms such as similar word expansion algorithms based on word vectors may be used to perform text expansion processing on each concept dimension to obtain a text database corresponding to each concept dimension, and each text database may include multiple prompt texts.
[0207] The prompt text processing module 423 may, according to multiple preset picture content dimensions, such as three picture content dimensions: spatial environment, character image, and behavior action, use text recognition algorithms such as BoW to divide the above multiple prompt texts to obtain prompt texts corresponding to each picture content dimension. And according to the picture content generation requirements, target prompt texts may be selected from the prompt texts corresponding to each picture content dimension.
[0208] Optionally, the server 42 may further include an image content generation requirement setting module 424. The server 42 may also send multiple image content dimensions and corresponding prompt texts to the client 41 for display in the user interface. The image content generation requirement setting module 424 may determine the user-specified image content generation requirements based on the weight values set by the user for multiple image content dimensions respectively, and select target prompt texts from the prompt texts corresponding to multiple image content dimensions respectively according to the image content generation requirements, so as to meet the personalized needs of users and improve the user experience.
[0209] The image generation module 425 may generate a target image based on the target prompt text by using a first large model such as an image generation large model like Dalle-2, StableDiffusion XL, etc. By obtaining the description data related to the target text from the database and performing dimensionality reduction processing on it, multiple concept dimensions corresponding to the target text are obtained, and text expansion processing is performed on multiple concept dimensions respectively to obtain a text database containing multiple prompt texts, thereby improving the accuracy and comprehensiveness of the prompt texts, and further improving the accuracy of the target prompt text selected therefrom and the target image generated based on the target prompt text.
[0210] Optionally, the server 42 may further include an image evaluation module 426. The image evaluation module 426 may use a second large model such as a text generation large model like CogVLM to generate a description text corresponding to the target image, and calculate the association score corresponding to the target image and the logical score corresponding to the target image based on the description text, the target text, and the target prompt text. Based on the logical score and the association score, the target image is evaluated. If it is determined that the target image does not meet the evaluation requirements, the text database and the image content generation requirements may be adjusted, and the steps of dividing the prompt texts and selecting the target prompt texts are returned to continue execution. After multiple iterations, a target image that meets the evaluation requirements is obtained. By adopting an evaluation method that combines logic and relevance, and through an automated iterative optimization mechanism, it is ensured that the generated images are continuously improved in quality and effect, and the accuracy of the images is further improved.
[0211] Optionally, the server 42 may further include a picture effect evaluation module 427. The server 42 may send target pictures that meet the evaluation requirements to the client 41 and obtain user response information for the target pictures in various forms such as voice, gestures, and physiological signals. The picture effect evaluation module 427 may score the target pictures based on the user response information using a third large model such as an artificial neural network large model according to multiple preset scoring dimensions, such as the user's social anxiety score, valence, arousal, immersion, and picture-text consistency, to obtain a multi-dimensional scoring result. Optionally, dimensionality reduction processing may also be performed on the target prompt text and description text corresponding to the target pictures that meet the evaluation requirements, and the dimensionality-reduced target prompt text and description text may be compared. According to the comparison result, the picture content dimension may be adjusted, and based on the user response information, the target pictures that meet the evaluation requirements may be scored using the third large model according to multiple preset scoring dimensions and the adjusted picture content dimension. Based on the target pictures whose scores meet the preset requirements, the mental health status of the user may be analyzed, for example, the social anxiety disorder status of the user may be analyzed, which helps to provide timely diagnosis and treatment for the user.
[0212] By providing a comprehensive intelligent system integrating a series of functions such as text-to-picture generation, picture evaluation, picture effect evaluation, and iterative update, it can accurately identify and judge the needs and preferences of users, generate target pictures with a high degree of compliance, and at the same time output multi-dimensional scores of the target pictures for auxiliary analysis, diagnosis, and treatment of the mental health status of users.
[0213] As Figure 5 shown, it is a schematic structural diagram of an embodiment of a picture generation device provided by the present application. The device may include the following units:
[0214] The first determination unit 501 is configured to determine a target text in response to a picture generation request.
[0215] The first text processing unit 502 is configured to obtain description data related to the target text from a database and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text.
[0216] The text expansion unit 503 is configured to perform text expansion processing on each of the multiple concept dimensions to obtain text databases corresponding to the multiple concept dimensions respectively, and the text databases include multiple prompt texts.
[0217] The division unit 504 is configured to divide the multiple prompt texts according to multiple preset picture content dimensions to obtain prompt texts corresponding to the multiple picture content dimensions respectively.
[0218] An extraction unit 505, configured to select a target prompt text from the prompt texts corresponding to the multiple picture content dimensions according to the requirements for generating picture content;
[0219] A generation unit 506, configured to generate a target picture by using a first large model based on the target prompt text.
[0220] Figure 5 The illustrated picture generation device can be used to implement Figure 1 The illustrated picture generation method, and its implementation principle and technical effects will not be elaborated. Among them, Figure 5 One or more units of the illustrated picture generation device can constitute the server in the above system architecture. For the picture generation device in the above embodiments, the specific manners of operations performed by each unit have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0221] As Figure 6 shown, it is a schematic structural diagram of an embodiment of a display device provided by the present application. The device may include the following units:
[0222] A first display unit 601, configured to display a user interface;
[0223] A first sending unit 602, configured to determine a target text based on an input operation on the user interface, generate a picture generation request based on the target text, and send the picture generation request to the server, so that the server obtains description data related to the target text from a database, performs dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; perform text expansion processing on the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, where the text databases include multiple prompt texts; divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to the multiple picture content dimensions respectively;
[0224] A second display unit 603, configured to obtain and display the multiple picture content dimensions and the corresponding prompt texts in the user interface;
[0225] A second sending unit 604, configured to generate an operation request based on a user operation on the user interface and send it to the server, so that the server determines picture content generation requirements based on the user operation, selects a target prompt text from the prompt texts corresponding to the multiple picture content dimensions respectively; generate a target picture by using a first large model based on the target prompt text;
[0226] A third display unit 605, configured to obtain and display the target picture in the user interface.
[0227] Figure 6 The display device shown can be used to implement Figure 3 the display method shown. The implementation principle and technical effects will not be elaborated here. Among them, Figure 6 one or more units of the display device shown can constitute the client in the above system architecture. For the display device in the above embodiments, the specific ways in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0228] As Figure 7 shown, it is a schematic structural diagram of an embodiment of a computing device provided by the present application. The device may include a storage component 701 and a processing component 702;
[0229] The storage component 701 can be used to store one or more computer program instructions. Among them, the one or more computer program instructions are called and executed by the processing component 702 to implement Figure 1 the picture generation method shown, Figure 2 the picture evaluation method shown, or Figure 3 the display method shown.
[0230] Of course, the above computing device may also include other components, such as an input / output interface, a communication component, etc.
[0231] The input / output interface provides an interface between the processing component and the peripheral interface unit. The above peripheral interface unit may be an output device, an input device, etc. The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0232] It should be noted that when the above computing device is used to implement Figure 1 the picture generation method shown, or Figure 2 the picture evaluation method shown, it can be a physical device or an elastic computing host provided by a cloud computing platform, etc. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can be implemented as a single server or a single terminal device.
[0233] When the above computing device is used to implement Figure 3 the display method shown, it can be implemented as an electronic device. The electronic device can refer to a device used by a user and having functions such as Internet access, computing, communication, etc. required by the user. For example, it can be a mobile phone, a tablet computer, a personal computer, a wearable device, etc. It can be understood that the above electronic device will also necessarily include other components such as a display component, an input / output interface, a communication component, etc., which will not be elaborated.
[0234] In one or more of the above embodiments, the processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above methods. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above methods.
[0235] The storage component is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0236] The display component may be an electroluminescent (EL) element, a liquid crystal display or a micro display with a similar structure, or a direct retinal display or a similar laser scanning display.
[0237] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a computer can implement Figure 1 the picture generation method shown, Figure 2 the picture evaluation method shown, or Figure 3 the display method shown. The computer-readable medium may be included in the computing device described in the above embodiments; or may exist separately without being assembled into the computing device.
[0238] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above.
[0239] An embodiment of the present application also provides a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program can implement Figure 1 the picture generation method shown, Figure 2 the picture evaluation method shown, or Figure 3 the display method shown.
[0240] In such an embodiment, the computer program may be downloaded and installed from the network and / or installed from a removable medium. When the computer program is executed by the processor, it executes various functions defined in the system of the present application.
[0241] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0242] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0243] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0244] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for generating pictures, characterized in that, Including: In response to a picture generation request, determine the target text; Obtain description data related to the target text from a database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; Perform text augmentation processing on each of the multiple concept dimensions to obtain a text database corresponding to each of the multiple concept dimensions, where the text database includes multiple prompt texts; Divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to each of the multiple picture content dimensions; Select target prompt texts from the prompt texts corresponding to each of the multiple picture content dimensions according to the picture content generation requirements; The picture content generation requirements include picture generation requirements for the picture content dimension; Based on the target prompt texts, use a first large model to generate a target picture; Also including: Use a second large model to generate a description text corresponding to the target picture; Based on the description text, the target text, and the target prompt texts, calculate an association score corresponding to the target picture; Calculate a logic score corresponding to the target picture; Based on the logic score and the association score, evaluate the target picture; Based on the evaluation result, if it is determined that the target picture does not meet the evaluation requirements, adjust the text database and the picture content generation requirements, and return to the step of dividing the multiple prompt texts according to a preset multiple picture content dimensions to obtain prompt texts corresponding to each of the multiple picture content dimensions, and continue to execute until the generated target picture meets the evaluation requirements; Send the target picture that meets the evaluation requirements to the client; Obtain user response information for the target picture; Perform dimensionality reduction processing on the target prompt text and the description text corresponding to the target picture that meets the evaluation requirements, and compare the dimensionality-reduced target prompt text with the description text, and adjust the picture content dimension according to the comparison result; According to multiple preset scoring dimensions and the adjusted picture content dimension, based on the user response information, use a third large model to score the target picture that meets the evaluation requirements.
2. The method according to claim 1, wherein Also including: Send the multiple picture content dimensions and the corresponding prompt texts to the client for display in the user interface; Based on the user operation in the user interface, determine the picture content generation requirements.
3. The method according to claim 2, characterized in that The determining the picture content generation requirements based on the user operation in the user interface includes: Based on the weight values respectively set by the user for the multiple picture content dimensions, determine the picture content generation requirements.
4. The method according to claim 1, wherein Also including: Extract the feature vector of the target text; Based on the feature vector, search for at least one associated text associated with the target text and the corresponding association weights of the at least one associated text in a target semantic network; where the target semantic network is pre-constructed based on a data set in the target domain where the target text is located; Add the at least one associated text to the text database in descending order of the association weights as prompt texts.
5. The method according to claim 4, characterized in that, Also including: Send the at least one associated text and the first selection prompt information to the client; Determine the target associated text based on the selection operation for the at least one associated text; The selecting the target prompt text from the prompt texts corresponding to the multiple picture content dimensions according to the picture content generation requirements includes: Select the target prompt text including the target associated text from the prompt texts corresponding to the multiple picture content dimensions according to the picture content generation requirements.
6. The method according to claim 5, wherein Further included is: Send the target prompt text and the update prompt information to the client; Based on the confirmation operation for the update prompt information, when the update condition is met, update the target semantic network based on the target prompt text.
7. The method according to claim 1, characterized in that, Further included is: Obtain the prompt text corresponding to the historical text related to the target text from the target platform; wherein, the target platform stores the historical texts uploaded by multiple users respectively, the historical pictures corresponding to the historical texts, and the prompt texts.
8. The method according to claim 2, wherein Further included is: Send multiple candidate picture styles and the second selection prompt information to the client; Determine the target picture style based on the selection operation for any one of the candidate picture styles; The generating the target picture by using the first large model includes: Generate the target picture by using the first large model according to the target picture style.
9. A display method, characterized in that, Included is: Display the user interface; Based on the input operation on the user interface, determine the target text, generate a picture generation request based on the target text, and send the picture generation request to the server, so that the server can obtain the description data related to the target text from the database, perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target text; perform text augmentation processing on the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, and the text databases include multiple prompt texts; divide the multiple prompt texts according to a preset multiple picture content dimensions to obtain the prompt texts corresponding to the multiple picture content dimensions respectively; Obtain and display the multiple picture content dimensions and the corresponding prompt texts in the user interface; Based on the user operation on the user interface, generate an operation request and send it to the server, so that the server can determine the picture content generation requirements based on the user operation, and select the target prompt text from the prompt texts corresponding to the multiple picture content dimensions according to the picture content generation requirements; the picture content generation requirements include the picture generation requirements for the picture content dimension; generate the target picture by using the first large model based on the target prompt text; generate the description text corresponding to the target picture by using the second large model; Calculate the association score corresponding to the target picture based on the description text, the target text, and the target prompt text; Calculate the logic score corresponding to the target picture; Based on the logical score and the correlation score, evaluate the target picture; when it is determined that the target picture does not meet the evaluation requirements based on the evaluation result, adjust the text database and the picture content generation requirements, and return to continue executing the step of dividing the multiple hint texts according to multiple preset picture content dimensions to obtain the hint texts corresponding to the multiple picture content dimensions respectively, until the generated target picture meets the evaluation requirements. Send the target picture that meets the evaluation requirements to the client; obtain the user response information for the target picture; perform dimensionality reduction processing on the target hint text and the description text corresponding to the target picture that meets the evaluation requirements, and compare the dimensionality-reduced target hint text with the description text, and adjust the picture content dimensions according to the comparison result; according to multiple preset scoring dimensions and the adjusted picture content dimensions, based on the user response information, use the third large model to score the target picture that meets the evaluation requirements; Obtain and display the target picture in the user interface.
10. The method according to claim 9, wherein It also includes: Obtain and display the target hint text in the user interface. Based on the upload operation for the target hint text, upload the target text, the target picture, and the target hint text to the target platform.
11. The method according to claim 9, characterized in that, It also includes: Based on the save operation for at least one hint text, save the at least one hint text locally.
12. A method for generating an image, characterized in that, It includes: In response to the picture generation request, determine the target concept text of mental health. Obtain the description data related to the target concept text from the database, and perform dimensionality reduction processing on the description data to obtain multiple concept dimensions corresponding to the target concept text. Perform text expansion processing on the multiple concept dimensions respectively to obtain text databases corresponding to the multiple concept dimensions respectively, and the text databases include multiple hint texts. Divide the multiple hint texts according to multiple preset picture content dimensions to obtain the hint texts corresponding to the multiple picture content dimensions respectively. Select the target hint text from the hint texts corresponding to the multiple picture content dimensions respectively according to the picture content generation requirements; the picture content generation requirements include the picture generation requirements for the picture content dimensions. Generate the target picture based on the target hint text using the first large model. Generate the description text corresponding to the target picture using the second large model. Calculate the correlation score corresponding to the target picture based on the description text, the target text, and the target hint text. Calculate the logical score corresponding to the target picture. Evaluate the target picture based on the logical score and the correlation score. When it is determined that the target picture does not meet the evaluation requirements based on the evaluation result, adjust the text database and the picture content generation requirements, and return to continue executing the step of dividing the multiple hint texts according to multiple preset picture content dimensions to obtain the hint texts corresponding to the multiple picture content dimensions respectively, until the generated target picture meets the evaluation requirements; Send the target picture that meets the evaluation requirements to the client; Obtain the user response information for the target picture; Perform dimensionality reduction processing on the target prompt text and description text corresponding to the target picture that meets the evaluation requirements, compare the dimensionality-reduced target prompt text with the description text, and adjust the picture content dimension according to the comparison result; Based on the user response information, score the target picture that meets the evaluation requirements using the third large model according to multiple preset scoring dimensions and the adjusted picture content dimension; wherein, the target picture is used to analyze the mental health status of the user.
13. A computing device, characterized in that, Includes a storage component and a processing component; the storage component stores one or more computer program instructions, and the computer program instructions are called and executed by the processing component, and the processing component executes the one or more computer program instructions to implement the picture generation method according to any one of claims 1 to 8, or the display method according to any one of claims 9 to 11, or the picture generation method according to claim 12.
14. A computer-readable storage medium, characterized in that, Stores a computer program, and the computer program is executed by a computer to implement the picture generation method according to any one of claims 1 to 8, or the display method according to any one of claims 9 to 11, or the picture generation method according to claim 12.
15. A computer program product, characterized in that, Stores a computer program, and when the computer program is executed by a computer, it implements the picture generation method according to any one of claims 1 to 8, or the display method according to any one of claims 9 to 11, or the picture generation method according to claim 12.
Citation Information
Patent Citations
Content text generation method and device and music comment text generation method
CN112115718A
Picture generation method and device, electronic equipment and storage medium
CN117221645A
Image generation method and data processing method for image generation
CN117409109A