Method and apparatus for generating image, electronic device, and readable storage medium
By providing a prompt input interface in image processing applications, users can generate a first prompt, and the system can automatically generate a second prompt. This solves the problem of users' inability to accurately express their intentions, improves the efficiency and accuracy of image generation, and enhances the user experience.
Patent Information
- Application Number
- PCT/CN2024/096317
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Users often struggle to accurately express their intentions using existing image processing applications, resulting in generated images that do not meet their needs. Furthermore, repeatedly adjusting prompts reduces efficiency and resource utilization.
The system provides a prompt input interface, where users generate the first prompt, and the system automatically generates the second prompt, which includes descriptions of the person's characteristics and the scene's features, reducing user operations and improving accuracy and efficiency.
By automatically generating secondary prompts, users are reduced from making repeated adjustments, and the generated images better match the user's intent, thus improving the efficiency and accuracy of image generation and enhancing the user experience.
Smart Images

Figure CN2024096317_04122025_PF_FP_ABST
Abstract
Description
Image generation method and device, electronic device, and readable storage medium TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to an image generation method and device, an electronic device, and a readable storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, a user can realize image processing through increasingly simple operations. For example, the user can realize beautification of a person in an image and adjustment of a background through an image beautification function provided in some application.
[0003] SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter.
[0005] According to some embodiments of the present disclosure, an image generation method is provided, including: displaying a prompt information input interface; in response to obtaining first prompt information, displaying the first prompt information on the prompt information input interface, wherein the first prompt information includes first description information of a target image; displaying second prompt information generated based on the first prompt information, wherein the second prompt information includes second description information of a feature of a person in the target image and third description information of a picture feature other than the feature of the person; and in response to a user triggering an image generation function, displaying a target image generated based on a to-be-processed image and the second prompt information.
[0006] According to some other embodiments of the present disclosure, an image generation device is provided, including: a first display module configured to display a prompt information input interface; a second display module configured to, in response to obtaining first prompt information, display the first prompt information on the prompt information input interface, wherein the first prompt information includes first description information of a target image; a third display module configured to, in response to displaying second prompt information generated based on the first prompt information, wherein the second prompt information includes second description information of a feature of a person in the target image and third description information of a picture feature other than the feature of the person; and a fourth display module configured to, in response to a user triggering an image generation function, display a target image generated based on a to-be-processed image and the second prompt information.
[0007] According to some other embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor, configured to store instructions, when the instructions are executed by the processor, causing the processor to execute an image generation method of any one of the embodiments of the present disclosure.
[0008] According to still another embodiment of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, which when executed by a processor, performs the image generation method of any one of the embodiments of the present disclosure.
[0009] According to yet another embodiment of the present disclosure, a computer program product is provided, comprising instructions, which when executed by a processor, implement the image generation method of any one of the embodiments of the present disclosure.
[0010] According to still another embodiment of the present disclosure, a computer program is provided, comprising instructions, which when executed by a processor, implement the image generation method of any one of the embodiments of the present disclosure.
[0011] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference made to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0012] The preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. The accompanying drawings are provided to provide further understanding of the present disclosure, and together with the following detailed description, form a part of the description of the present disclosure and serve to explain the present disclosure. It should be understood that the accompanying drawings only relate to some embodiments of the present disclosure, and do not limit the present disclosure. In the drawings:
[0013] FIG. 1 shows a flowchart of an image generation method according to some embodiments of the present disclosure;
[0014] FIG. 2 shows a schematic diagram of a prompt information input interface according to some embodiments of the present disclosure;
[0015] FIG. 3 shows a schematic diagram of a prompt information input interface according to some other embodiments of the present disclosure;
[0016] FIG. 4A shows a schematic diagram of a prompt information input interface according to yet some other embodiments of the present disclosure;
[0017] FIG. 4B shows a schematic diagram of a prompt information input interface according to still some other embodiments of the present disclosure;
[0018] FIG. 5A shows a schematic diagram of a prompt information input interface according to yet some other embodiments of the present disclosure;
[0019] FIG. 5B shows a schematic diagram of a prompt information input interface according to still some other embodiments of the present disclosure;
[0020] FIG. 6 shows a schematic diagram of an image generation apparatus according to some embodiments of the present disclosure;
[0021] FIG. 7 shows a schematic diagram of an electronic device according to some embodiments of the present disclosure;
[0022] Figure 8 shows a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure.
[0023] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0024] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0025] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0026] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".
[0027] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.
[0028] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the terms "first", "second", and the like are not intended to imply a given order or any other manner of given order in time, space, ranking, or any other manner.
[0029] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that unless otherwise explicitly specified in the context, it should be understood as "one or more".
[0030] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0031] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined by any suitable means from the present disclosure which is clear to those skilled in the art.
[0032] A user can achieve image processing through a specified function provided by some image processing application (APP). Especially for some images containing people, the user can use the specified function in the application to beautify, for example, modify the clothes, actions, change the background, add some materials, etc. However, these specified functions are usually fixed and the same for all users, and the user often cannot get a satisfactory image.
[0033] With the emergence of large models, applications based on large models can generate images or process images according to some prompt information (Prompt). However, it is difficult for ordinary users to accurately express their intentions through input prompt information and for the large model to accurately understand the intention. Therefore, processing images based on user input prompt information is also difficult to meet the user's demand to generate accurate images, and the user may need to repeatedly adjust the prompt information, reducing the efficiency of image processing and wasting computing resources. If some candidate prompt information is provided for the user to choose, these candidate prompt information is limited, which will also cause the user to be unable to get a satisfactory image, and the user will stack multiple pieces of prompt information by selecting, which will make the large model unable to accurately understand and generate accurate images. The process of selecting prompt information by the user also reduces the efficiency of image processing.
[0034] Based on the above idea, the present disclosure provides an image generation method, which supports generating second prompt information based on the first prompt information generated by the user, reduces the number of times and operation complexity of the user repeatedly adjusting the prompt information, and generates second prompt information content that is more rich and can more accurately and fully describe the target image, thereby improving the efficiency and accuracy of image generation, making the generated image more in line with the user's intention, and improving the user's experience.
[0035] FIG. 1 is a flowchart of some embodiments of the image generation method of the present disclosure. As shown in FIG. 1, the method of the embodiment includes steps S102-S108.
[0036] In step S102, a prompt information input interface is displayed.
[0037] The user can trigger the display of the prompt information input interface in various ways. In some embodiments, the prompt information input interface is displayed in response to the user's operation on the to-be-processed image.
[0038] The to-be-processed image can or can not include an image of a person. In the case where the to-be-processed image includes an image of a person, the person can be a real person or a generated virtual person. The to-be-processed image is an image that has been authorized for application.
[0039] In some embodiments, in response to the user triggering an image processing function, an image determination page is displayed, and in response to the user's selection or shooting operation on the to-be-processed image, the prompt information input interface is displayed. The user can select the to-be-processed image from the stored images or take an image as the to-be-processed image. The to-be-processed image can also be displayed in the prompt information input interface.
[0040] In some embodiments, in response to the user triggering an image processing function, an image determination page is displayed, and in response to the user's selection or shooting operation on the to-be-processed image, the prompt information input interface is displayed. The user can select the to-be-processed image from the stored images or take an image as the to-be-processed image. The to-be-processed image can also be displayed in the prompt information input interface.
[0041] In some embodiments, the prompt information input interface is displayed in response to the user's interaction with the intelligent agent. For example, the user can mention the information of adjusting the to-be-processed image in the chat process with the intelligent agent, and the prompt information input interface can be displayed.
[0042] The triggering of the prompt information input interface can be in various ways, not limited to the above-mentioned examples.
[0043] In step S104, in response to obtaining the first prompt information, the first prompt information is displayed in the prompt information input interface.
[0044] For example, the first prompt information includes first description information of the target image. The first description information can be description information of a subject of the target image, and can be used to describe a display effect of the target image. For example, the first description information includes description information of a feature of a person in the target image and / or description information of an image style. The feature of the person can include at least one of an appearance, an action, a profession, and a relationship with an object other than the person, without being limited to the examples. For example, the appearance includes at least one of a hairstyle, a costume, and an accessory. The profession of the person can affect the appearance features such as the costume of the person. The relationship of the person with the object other than the person is, for example, holding a cat, driving a car, and the like. The description information of the feature of the person can be description information of any feature of the person, without being limited to the examples. For example, the description information of the image style is a Japanese anime style, a Barbie style, and the like, without being limited to the examples.
[0045] The first prompt information can also include description information of a picture feature other than the person and the image style, such as description information of at least one of an environment, a color, a composition, a light, and an atmosphere. For example, the environment includes a background, surroundings, and the like, the color includes low saturation, bright, and the like, the composition includes a golden section, a wide angle, a bird's eye view, and the like, the light includes top light, diffuse reflection, and the like, and the atmosphere includes under the sun, foggy, horror, and the like, without being limited to the examples.
[0046] The user can generate the first prompt information by inputting and / or selecting operations, and the first prompt information can be displayed in the first display area of the prompt information input interface.
[0047] In step S106, second prompt information generated based on the first prompt information is displayed.
[0048] For example, the second prompt information includes second description information of a feature of a person in the target image and third description information of a picture feature other than the feature of the person. The third description information can be description information of at least one of an environment, a color, a composition, a light, an atmosphere, and a style. The second prompt information is generated by adjusting the first prompt information, and is more accurate, clear, and easy for a model to understand the target image than the first prompt information. The second prompt information can be generated by expanding, shortening, or rewriting the first prompt information. Expanding means adding description information on the basis of the first prompt information or the first description information, shortening means shortening the length of the first prompt information or the first description information while keeping the semantics unchanged, and rewriting means modifying all or part of the content of the first prompt information or the first description information in a manner other than expanding and shortening.
[0049] The second prompt information can be displayed in the first display area, or can be displayed in various forms of interfaces such as a floating layer, a pop-up window, and a panel, without being limited to the examples.
[0050] In step S108, in response to the user's triggering of the image generation function, the target image generated based on the image to be processed and the second prompt information is displayed.
[0051] For example, based on the second prompt, the features of the person and the scene in the image to be processed are adjusted to generate the target image. The generated target image can be displayed on the prompt input interface or on a new interface; there is no restriction on this.
[0052] In the method of the above embodiments, a first prompt message can be generated and displayed by user operation. The first prompt message includes a first description of the target image by the user. Then, a second prompt message is automatically generated based on the first prompt message. The second prompt message includes a second description of the features of the person in the target image and a third description of the image features other than the features of the person. This can more accurately and fully describe the target image. Users can obtain more accurate and detailed second description information by inputting a brief first prompt message. Furthermore, it reduces the process of users repeatedly modifying or selecting prompt messages, improves the efficiency and accuracy of target image generation, enhances user experience, and enables users to obtain target images that better meet their needs.
[0053] The method described in the above embodiments can be executed on the client side. The client can send a first prompt message to the server, and the server can generate a second prompt message based on the first prompt message and return it to the client for display. The server can generate a target image based on the image to be processed and the second prompt message, and return it to the client for display.
[0054] The following describes the prompt information input interface and the process of obtaining the first prompt information in conjunction with some embodiments.
[0055] In some embodiments, the first prompt information is obtained in response to the user entering first prompt information in the input area of the prompt information input interface.
[0056] Figure 2 shows the prompt information input interface, including an input area 201, which can display a virtual keyboard, etc. The prompt information input interface may also include a first display area 202, where the user-inputted prompt information can be displayed. Before the user inputs the first prompt information, preset guidance information can be displayed in the first display area 202 to guide the user in inputting the first prompt information. For example, the preset guidance information could be "Describe the desired effect, intelligently optimize and generate".
[0057] In some embodiments, a recommended control is displayed on the prompt information input interface, and multiple candidate prompt messages are displayed in response to the user's triggering of the recommended control.
[0058] In some embodiments, multiple candidate prompts are displayed, and in response to the user's selection of a first candidate prompt from the multiple candidate prompts, the first candidate prompt is obtained as the first prompt.
[0059] For example, multiple candidate prompts are displayed in the second display area of the prompt input interface. In response to the user's selection of the first candidate prompt from among the multiple candidate prompts, the first candidate prompt is selected as the first prompt. Multiple candidate prompts can also be displayed in other forms of interfaces such as pop-ups and panels, and are not limited to the examples given.
[0060] As shown in Figure 2, a recommendation control 203 is displayed on the prompt input interface. In response to the user's triggering of the recommendation control 203, as shown in Figure 3, multiple candidate prompts are displayed in the second display area 301, such as "Doctor with fluffy curly hair" or "Bald artist." Alternatively, only one candidate prompt can be displayed. The second display area can be expanded or collapsed. When expanded, multiple candidate prompts and a hidden control 302 are displayed; when the user triggers the hidden control 302, the multiple candidate prompts are hidden. When the second display area is collapsed, guidance information (e.g., keyword recommendations) and an expand control (which can serve as a recommendation control) can be displayed.
[0061] Candidate suggestions can also be called keywords. For example, each candidate suggestion includes a fourth descriptive information such as the characteristics of a person or the style of an image.
[0062] Multiple candidate prompts can also be displayed using other formats such as overlays, pop-ups, half-screen pages, and panels. Multiple candidate prompts can also be displayed directly in the second display area without needing to set up recommendation controls or for users to trigger them. There are various methods for displaying multiple candidate prompts, not limited to the examples mentioned above.
[0063] In some embodiments, in response to a user's selection of a first candidate prompt among multiple candidate prompts, the first candidate prompt is displayed in the display area of the multiple candidate prompts with a preset effect.
[0064] For example, each of the multiple candidate prompts includes a fourth descriptive information about the person's characteristics or image style. For example, in response to the user's selection of a first candidate prompt among the multiple candidate prompts, the first candidate prompt is displayed as the first prompt in the first display area, and the first candidate prompt is also displayed in the second display area with a preset effect.
[0065] The preset effect indicates that the first candidate suggestion is selected. For example, the preset effect may be highlighted, have a background color, or be underlined. Displaying the first candidate suggestion in the first display area and simultaneously displaying it in the second display area with the preset effect improves the display effect, makes it easier for users to view the selected suggestion, avoids repetitive and useless operations, and improves operational efficiency.
[0066] In some embodiments, in response to a user inputting initial prompt information in the input area of the prompt information input interface, the initial prompt information is displayed in the prompt information input interface, and multiple candidate prompt information is displayed. In response to the user selecting a first candidate prompt information from the multiple candidate prompt information, a combination of the initial prompt information and the first candidate prompt information is obtained as the first prompt information.
[0067] For example, in response to a user entering initial prompt information in the input area of the prompt information input interface, the initial prompt information is displayed in the first display area of the prompt information input interface, and multiple candidate prompt information is displayed in the second display area of the prompt information input interface. In response to the user's selection of the first candidate prompt information among the multiple candidate prompt information, a combination of the initial prompt information and the first candidate prompt information is obtained as the first prompt information.
[0068] Users can first enter an initial prompt message, then select a first candidate prompt message, and can also modify both the initial prompt message and the first candidate prompt message. In response to user modification of at least one of the initial prompt message and the first candidate prompt message, the modified combination of the initial prompt message and the first candidate prompt message is displayed as the first prompt message in the first display area.
[0069] The display of multiple candidate prompts can also be triggered by the user interacting with the recommendation control, which will not be elaborated upon here. The order of displaying multiple candidate prompts and inputting the initial prompt is not fixed; the user can input the initial prompt first and then trigger the display of multiple candidate prompts, or vice versa.
[0070] In some embodiments, multiple candidate prompts are displayed. In response to the user's selection of a first candidate prompt among the multiple candidate prompts, the first candidate prompt is displayed on the prompt input interface. In response to the user's adjustment of the first candidate prompt, the adjusted prompt is obtained as the first prompt.
[0071] For example, multiple candidate prompts are displayed in the second display area of the prompt input interface. In response to the user's selection of the first candidate prompt among the multiple candidate prompts, the first candidate prompt is displayed in the first display area of the prompt input interface. In response to the user's adjustment of the first candidate prompt, the adjusted prompt is obtained as the first prompt.
[0072] The display of multiple candidate prompts can also be triggered by the user activating the recommendation control, which will not be elaborated here.
[0073] In some embodiments, user adjustment of the first candidate prompt information includes at least one of adding supplementary prompt information and modifying the first candidate prompt information based on the first candidate prompt information.
[0074] For example, in response to adding supplementary prompts based on the first candidate prompt, the first candidate prompt is displayed in the display area of multiple candidate prompts with a preset effect. In response to modifying the first candidate prompt, the first candidate prompt is removed from the display area of multiple candidate prompts and displayed with a preset effect. Modifying the first candidate prompt includes deleting some content from it, or modifying some or all of its content. If the user adds supplementary prompts, the first candidate prompt remains selected; if the user modifies the first candidate prompt, it is no longer selected. This linkage between the display effects of candidate prompts in different areas improves the display effect, makes it easier for users to view selected prompts, avoids repetitive and useless operations, and improves operational efficiency.
[0075] The first candidate prompt can be modified by inputting information or by selecting a second candidate prompt to replace it. The selection and modification of candidate prompts, as well as user input, can be combined arbitrarily to obtain the first prompt.
[0076] In some embodiments, in response to a user adding supplementary prompt information based on the first candidate prompt information, the first candidate prompt information and the supplementary prompt information are displayed on the prompt information input interface; in response to a user replacing the first candidate prompt information with a second candidate prompt information from a plurality of candidate prompt information, the second candidate prompt information and the supplementary prompt information are displayed on the indicated prompt information input interface as the adjusted prompt information.
[0077] For example, in response to a user adding supplementary prompt information based on the first candidate prompt information, the supplementary prompt information is displayed in the first display area; in response to a user replacing the first candidate prompt information with the second candidate prompt information from multiple candidate prompt information, the second candidate prompt information and the supplementary prompt information are displayed in the first display area as the adjusted prompt information.
[0078] Users can add supplementary suggestions based on the initial suggestion. For example, a user selects "doctor with fluffy curly hair" as the initial suggestion and then adds "wearing glasses" as a supplementary suggestion. The user can then select a second suggestion to replace the initial suggestion. For instance, if the user selects "wearing a school uniform" as the second suggestion to replace "doctor with fluffy curly hair," the final suggestion will be "wearing a school uniform and glasses."
[0079] The maximum number of candidate prompts that a user can select can be configured. If the number of candidate prompts selected by the user exceeds the maximum and the user selects another candidate prompt, the latest selected prompt replaces the earliest selected prompt. For example, the maximum number can be 1, 2, etc. For instance, if the maximum number is 1, and the user has already selected "doctor with fluffy curly hair" as the first candidate prompt, and then selects "wearing a school uniform," then "doctor with fluffy curly hair" will be replaced with "wearing a school uniform" in the first display area.
[0080] The selected candidate prompts in the second display area are displayed with a preset effect. If the selected candidate prompt is replaced, the replaced candidate prompt is displayed with the preset effect. If the selected candidate prompt is modified by user input, it is no longer displayed with the preset effect. This notifies the user of the candidate prompts they have already selected, preventing duplicate selections and the generation of invalid prompts.
[0081] In some embodiments, in response to the user's selection of a third candidate prompt, and while existing prompts are displayed in the first display area, the third candidate prompt is added after the existing prompts in a preset format.
[0082] Existing prompts include prompts entered by the user and / or candidate prompts selected and modified through input. For example, a preset format includes multiple prompts separated by preset punctuation marks. A preset punctuation mark is, for example, a comma. If other specific punctuation marks and / or separators (such as spaces) appear after existing prompts, they can be automatically changed to preset punctuation marks. For example, if a period, comma, or space appears after existing prompts, it can be automatically changed to a comma. If a comma appears after existing prompts, the third candidate prompt is directly added after the existing prompts for display.
[0083] In the above embodiments, users can generate the first prompt information through a combination of various operation methods such as input and selection, which facilitates the user's operation. Furthermore, by providing candidate prompt information, the user's operation efficiency can be improved.
[0084] The following describes the method for generating and displaying candidate prompt information with reference to some embodiments.
[0085] Multiple candidate prompts can be retrieved from a database and displayed. In some embodiments, multiple candidate prompts are retrieved from the database in response to the opening of the prompt input page. Multiple candidate prompts can be retrieved each time the prompt input page is accessed, and the candidate prompts retrieved each time can be different. The retrieved multiple candidate prompts fill the second display area (the display area for multiple candidate prompts).
[0086] For example, candidate prompts can be sorted in the database from highest to lowest usage frequency, and then selected according to this sorting. Alternatively, other metrics can be used for sorting, such as sorting by the number of times each candidate prompt has been displayed from lowest to highest. The sorting and selection methods for candidate prompts are not limited to the examples given above.
[0087] In some embodiments, in response to the image to be processed including a person, multiple candidate prompts are determined based on at least one of a first feature of the person, a second feature of a person in a user-posted image, and an image style of the posted image, and the multiple candidate prompts are displayed in a second display area; and / or in response to the image to be processed not including a person, multiple candidate prompts are determined based on at least one of a second feature of a person in a user-posted image and an image style of the image to be processed, and the multiple candidate prompts are displayed in a second display area.
[0088] It can identify whether the image to be processed contains people, and if so, identify the first characteristic of the person. Images previously published by the user are examples of images generated and published by the user with their authorization. The people in the user-published images may be the same as or different from the people in the image to be processed. The first and second characteristics include, for example, appearance, actions, and relationship to objects other than people. The image style of the published images may be, for example, anime style, Chinese style, etc.
[0089] For example, based on at least one of the following: a first feature of a person in the image to be processed, a second feature of a person in an image posted by the user, and the image style of the posted image, the user's preference for the image can be determined, and then multiple candidate prompts matching that preference can be generated.
[0090] The method described above generates multiple candidate prompts based on at least one of the following: a first feature of a person in the image to be processed, a second feature of a person in an image published by the user, and the image style of the published image. Different candidate prompts can be displayed for different users, providing personalized recommendations, improving the effectiveness of the candidate prompts, and enhancing the efficiency and experience of user selection.
[0091] In some embodiments, a switching control is displayed, and in response to the user's triggering of the switching control, multiple candidate prompt messages are displayed after the switch.
[0092] For example, a switching control is displayed in the display area of multiple candidate prompts (e.g., the second display area), and multiple candidate prompts are displayed in response to the user's triggering of the switching control.
[0093] As shown in Figure 3, the toggle control 303 "Change Group" can be displayed. The multiple candidate prompts after the switch are different from the multiple candidate prompts before the switch. Through the toggle control, users can quickly view multiple candidate prompts, providing them with more choices, improving operational efficiency and user experience.
[0094] The following describes, with reference to some embodiments, the display method and generation method related to the second description information.
[0095] The second prompt is generated based on the first prompt and conforms to the target evaluation metrics. Target evaluation metrics can be used to construct prompts that make the large model easier to understand and more accurately represent the target image. For example, target evaluation metrics include the type of descriptive information, the target expression method, and its length. For example, the type of descriptive information can be configured to include descriptive information about the features of a person and descriptive information about other scene features besides the person's features. Another example is that the type of descriptive information can be configured to include descriptive information about the person's clothing, actions, and other features, as well as descriptive information about scene features such as environment and lighting. Yet another example is that the type of descriptive information can be configured to include the first descriptive information and descriptive information that expands upon the main features in the first descriptive information; the main features, for example, are the features of a person. For example, the target expression method can be configured to meet requirements such as non-repetition, conciseness, coherence, and clarity.
[0096] The second prompt message can be directly generated and replaced by the first prompt message displayed in the first display area, or it can be displayed in other areas of the prompt message input interface or in interfaces such as pop-ups, overlays, and panels. In response to the user's confirmation operation, the second prompt message is displayed in the first display area.
[0097] In some embodiments, in response to a user's triggering of an adjustment function, a second prompt message generated based on the first prompt message is displayed.
[0098] For example, a control displaying adjustment functions can display a second prompt message in response to a user's triggering action on the control. As shown in Figure 4A, for example, if the user selects the first prompt message "Winged White Dragon Horse," the adjustment function control 401 "Help Me Optimize" can be displayed below the first prompt message. If there is no information in the first display area, or if the first display area contains the second prompt message, the adjustment function control is not displayed. The input area can be hidden in response to a user's triggering action on the adjustment function control.
[0099] In some embodiments, one or more third prompt messages generated based on the first prompt message are displayed; in response to the user's selection of a target prompt message among the one or more third prompt messages, a second prompt message determined based on the target prompt message and the first prompt message is displayed.
[0100] For example, in response to a user's triggering of an adjustment function, one or more third prompts generated based on the first prompt are displayed. The triggering method for generating the second prompt can be referenced, and will not be elaborated upon here.
[0101] One or more third prompts can be generated based on the first prompt. After the user selects a target prompt from the one or more third prompts, the target prompt can be combined with the first prompt to obtain a second prompt, or the target prompt can be used directly as the second prompt. If the target prompt is supplementary descriptive information to the first prompt, it is combined with the first prompt to obtain the second prompt; otherwise, the target prompt is used directly as the second prompt. One or more third prompts can be displayed in the third display area of the prompt input interface, and can be displayed in various forms of interfaces such as overlays, pop-ups, and panels, not limited to the examples given.
[0102] In some embodiments, one or more third prompt messages generated based on the first prompt message and the target adjustment method are displayed; in response to the user's selection of the target prompt message among the one or more third prompt messages, a second prompt message determined based on the target prompt message and the first prompt message is displayed.
[0103] For example, target adjustment methods include expanding, abbreviating, or rewriting. The target adjustment method is determined based on the initial prompt information.
[0104] After the user inputs the first prompt, one or more generated third prompts are displayed. Based on the user's selection, a second prompt is then determined and displayed. Through these two interactive processes, the second prompt can be determined more accurately, resulting in a more accurate generated image. Furthermore, this provides the user with multiple options, facilitating operation and improving efficiency.
[0105] In some embodiments, in response to the target adjustment method being expansion, one or more third prompt messages include supplementary description information generated based on the first description information; or in response to the target adjustment method being abbreviation, one or more third prompt messages include description information generated from the abbreviation of the first description information; or in response to the target adjustment method being rewrite, one or more third prompt messages include description information generated from the rewrite of the first description information.
[0106] Expanding, abbreviating, and rewriting are three different types of modification methods for the first prompt information, determined based on the first prompt information, which can generate the third prompt information more accurately.
[0107] In some embodiments, during the generation of one or more third prompt messages, a first guidance message is displayed, wherein the first guidance message is used to indicate that one or more third prompt messages are being generated.
[0108] If the second prompt message is generated directly, the first guiding message can also be displayed to indicate that the second prompt message is being generated. As shown in Figure 4B, during the generation of one or more third prompt messages, the first guiding message 402 can be displayed in the prompt message input interface. For example, the first guiding message is "Optimizing prompt words". If multiple candidate prompt messages are displayed in the second display area of the prompt message input interface, and one or more third prompt messages are displayed in the third display area, to avoid conflicts, during the generation of one or more third prompt messages, the multiple candidate prompt messages are hidden, and the second display area is in a collapsed state.
[0109] In some embodiments, a second guidance message is displayed, wherein the second guidance message is used to indicate the target adjustment method corresponding to one or more third prompt messages.
[0110] The second guiding information can be displayed simultaneously with one or more third prompts. For example, the second guiding information can be displayed within the third display area. As shown in Figure 5A, the first prompt is "Cheers to everyone," and multiple third prompts are displayed in the third display area 502, along with the second guiding information 501 "Try adding modifiers for a better effect," indicating that the target adjustment method is expansion. If the target adjustment method is abbreviation, for example, the second guiding information could be "The prompt is too long; we've optimized it for you." If the target adjustment method is rewriting, for example, the second guiding information could be "The prompt can be even more perfect; we've polished it for you." The specific content of the second guiding information can be set according to actual needs and is not limited to the examples given.
[0111] In some embodiments, an added control is displayed; in response to the user's triggering of the added control, one or more new third prompts generated based on the first prompt and the target adjustment method are displayed.
[0112] As shown in Figure 5A, an add control 503 "Load More" is displayed in the third display area. When the user triggers the add control, more third-level prompts can be displayed. The size of the third display area is dynamically adjusted according to the number of third-level prompts. The position and font of the third-level prompts can also be dynamically adjusted based on their quantity. For example, when displaying more third-level prompts, the third display area increases, the original content moves upwards, and the new third-level prompts are displayed.
[0113] For example, in response to a user's triggering of a recommendation control, multiple candidate suggestions can be displayed, while one or more third-party suggestions can be hidden.
[0114] After a user selects a target prompt from one or more third prompts, a second prompt can be generated, for example, displayed in the first display area. As shown in Figure 5B, if the user selects "Wearing a suit, smiling brightly, under bright lights," it will be displayed as selected in the third display area (preset effect), and this information will be added after "The person toasting," generating a second prompt, which will then be displayed in the first display area 504. Alternatively, after selecting a target prompt, the user may choose not to display one or more third prompts.
[0115] In some embodiments, in response to not adjusting the first prompt information, a third guidance information is displayed on the prompt information input interface, wherein the third guidance information is used to indicate that the first prompt information is used directly to generate the target image without adjustment.
[0116] The primary prompt message is not adjusted; the system is in a state where an image can be generated directly. Multiple candidate prompt messages and one or more third or second prompt messages do not need to be displayed. For example, the third prompt message could be "The prompt is perfect; an image can be generated directly."
[0117] In some embodiments, in response to the failure to generate the third prompt message, a fourth guidance message is displayed to indicate that the third prompt message generation failed. For example, the fourth prompt message is "Optimization failed, please try again".
[0118] The following describes in detail how one or more third prompt messages are generated, using some examples. If the second prompt message is generated directly, the method for generating the third prompt message can also be referred to, and will not be repeated here.
[0119] In some embodiments, a machine learning model is used to: parse the first prompt information and determine the parsing result; determine whether to adjust the first prompt information based on the parsing result; in response to adjusting the first prompt information, determine the target adjustment method based on the parsing result; and generate one or more third prompt information based on the first prompt information and the target adjustment method.
[0120] For example, the machine learning model could be an LLM (Large Language Model), but it's not limited to the examples given. Preset prompts and initial prompts can be input into the machine learning model, enabling it to parse the initial prompt based on the preset prompts, determine whether to adjust it, and, if so, determine the target adjustment method and generate one or more third prompts that meet the target evaluation metrics. For example, preset prompts might include information such as the machine learning model's role, task, and constraints.
[0121] In the above embodiments, a machine learning model is used to parse and understand the first prompt information. Based on the parsing results, it is determined whether the first prompt information should be adjusted. If adjustment is required, the target adjustment method is determined, and one or more third prompt information messages are generated. The method described in the above embodiments can flexibly adjust the first prompt information according to different situations, thereby making the generated third prompt information more accurate.
[0122] In some embodiments, the parsing result includes semantic information of the first description information. Based on the semantic information of the first description information, it is determined whether the first description information only includes description information of the features of the person or description information of the image style. In response to the first description information only including description information of the features of the person or description information of the image style, the target adjustment method is determined to be expansion.
[0123] Machine learning models can understand the semantics of the initial descriptive information to determine whether it is only used to describe the features of a person or the style of an image. If so, some additional descriptive information is needed, and the target adjustment method is determined to be expansion. Image style can be understood as a type of image feature.
[0124] In some embodiments, in response to the first description information including only the description information of the character's features, the target adjustment method is to expand and determine to expand the description information of the character's features. One or more third description information includes supplementary description information of the character's features that matches the description information of the character's features and supplementary description information of the screen features that matches the character's features. The description information of the character's features and the supplementary description information of the character's features are combined as the second description information, and the supplementary description information of the screen features is used as the third description information.
[0125] In some embodiments, in response to the first description information only including description information of the character's features, the target mode is to expand, and it is determined that the description information of the character's features will not be expanded, one or more third description information includes supplementary description information of screen features that match the description information of the character's features, wherein the description information of the character's features is used as the second description information, and the supplementary description information of screen features is used as the third description information.
[0126] For example, machine learning can be used to determine whether the description information of a person's features includes description information of a specific type of feature, such as clothing or actions. If it does not include description information of a specific type of feature, then it is determined that the description information of the person's features should be expanded.
[0127] Machine learning can be used to generate supplementary descriptive information about the person's features and the scene's features that match the descriptive information about the person's features. For example, if the first descriptive information is "a person toasting," then based on the person's action features, matching features such as clothing, facial features, surrounding environment, lighting, and atmosphere can be determined.
[0128] In some embodiments, in response to the first description information including only the description information of scene features other than people, the target adjustment method is determined to be expansion.
[0129] If the first descriptive information does not include descriptive information about the characteristics of a person, one or more third descriptive information pieces include supplementary descriptive information about the characteristics of the person that matches the descriptive information about the image features. It may also be further determined whether to expand the descriptive information about the image features; if it is determined that the descriptive information about the image features should be expanded, then one or more third descriptive information pieces include supplementary descriptive information about the image features that matches the descriptive information about the image features.
[0130] In some embodiments, in response to the first description information including only image style description information, the target adjustment method is expansion, and one or more third description information includes supplementary description information of character features that match the image style generated based on the image style description information, and supplementary description information of scene features that match the image style, wherein the supplementary description information of character features serves as the second description information, and the image style description information and the supplementary description information of scene features serve as the third description information.
[0131] If the initial descriptive information only includes image style descriptions, such as anime style or Barbie style, then it is necessary to expand on the character's features and specific visual characteristics. Machine learning models can be used to generate supplementary descriptive information on the character's features and visual characteristics that match the image style.
[0132] In the above embodiments, based on the semantic information of the first description information, the type of description information contained in the first description information is determined, and then the missing description information is supplemented. The supplemented content matches the existing content, making the generated third prompt information more accurate and in line with the user's intent. Furthermore, the third prompt information contains richer content, making the generated image more rich and harmonious.
[0133] In some embodiments, in response to the first description information including description information of the character's features or description information of the image style, and description information of screen features other than the character's features, it is determined whether the description information of the character's features or description information of the image style matches the description information of screen features other than the character's features; in response to the description information of the character's features or description information of the image style not matching the description information of screen features other than the character's features, it is determined that the target adjustment method for the first prompt information is rewriting.
[0134] If the descriptions of various features in the first description information do not match, or if there are contradictions or conflicts, the target adjustment method can be determined to be rewriting. For example, based on the description information of a person's features, updated description information of the image features that matches the description information of the person's features can be generated; or, based on the description information of the image style, updated description information of the person's features and / or updated description information of the image features that matches the description information of the image style can be generated.
[0135] In the above embodiments, when there are multiple mismatched descriptive information in the first descriptive information, the target adjustment method is determined to be rewriting, so that the generated third prompt information is clearer and more accurate, thereby making the generated target image more harmonious and accurate.
[0136] In some embodiments, in response to the inclusion of semantically ambiguous words in the first description information, the target adjustment method for the first prompt information is determined to be rewriting.
[0137] Based on the semantics of the first description information, words with unclear meanings in the first description information can be modified.
[0138] In some embodiments, the parsing result includes the length of the first description information, and determines whether the length of the first description information exceeds a preset length; in response to the length of the first description information exceeding the preset length, the target adjustment method is determined to be an abbreviation.
[0139] The length of the first descriptive information can be determined based on the number of words or characters. For example, using a machine learning model, the first descriptive information can be abbreviated to generate one or more third descriptive information entries that do not exceed a preset length and have the same semantic meaning as the first descriptive information. These one or more third descriptive information entries include descriptions of character features and descriptions of visual features.
[0140] In some embodiments, in response to the length of the first description information exceeding a preset length, and the first description information including description information of character features and image features, the target adjustment method is determined to be abbreviation; in response to the length of the first description information exceeding a preset length, and the first description information including only description information of character features or image features, the target adjustment method is determined to be rewrite.
[0141] Based on the length and type of the first descriptive information, it is determined whether to abbreviate or rewrite it. For example, if the target adjustment method is rewriting, on the one hand, the first descriptive information needs to be shortened, and on the other hand, missing descriptive information about the characteristics of the characters or the features of the scene needs to be supplemented, so that the generated third descriptive information is more concise, accurate and richer in content.
[0142] In the method of the above embodiments, the first descriptive information with a length exceeding a preset length can be shortened, so that the model can more accurately understand the second prompt information when generating the target image, thereby generating a more accurate target image.
[0143] In some embodiments, the parsing result includes the expression of the first descriptive information, determining whether the expression of the first descriptive information conforms to the target expression; in response to the expression of the first descriptive information not conforming to the target expression, determining the target adjustment method as rewriting, wherein one or more third prompt messages conform to the target expression.
[0144] For example, the target expression can be configured to meet requirements such as non-repetition, conciseness, coherence, and clarity. The machine learning model can understand the initial descriptive information to determine whether it conforms to the target expression; if not, it can rewrite the descriptive information to conform to the target expression.
[0145] The method described in the above embodiments can determine whether the expression of the first descriptive information conforms to the target expression. If it does not conform, it is rewritten to the descriptive information that conforms to the target expression. In this way, the generated one or more third prompts are more accurate and concise, which makes it easier for the model to understand when generating the target image, thereby generating a more accurate target image.
[0146] Whether expanded, abbreviated, or rewritten, the generated one or more third-party prompts all conform to the constraints of the target expression method, preset length, and descriptive information including character and visual features. Therefore, the constraints can be configured by configuring preset prompts for generating one or more third-party prompts. By inputting the preset prompts and the first prompts into a machine learning model, the machine learning model can automatically generate one or more third-party prompts.
[0147] If the second prompt message is generated directly, one or more of the third prompt messages in the above embodiments can be replaced with the second prompt message. The generation method is similar and will not be described again here.
[0148] The user selects one or more target prompts from the third prompts, generating and displaying a second prompt. If the user confirms that the second prompt can trigger the image generation function, a target image is generated. As shown in Figure 5B, a generation control 505 is displayed. Clicking this control 505 triggers the generation of the target image. After generation, the target image can be displayed in the prompt input interface or a new interface.
[0149] In practical applications, the initial prompts input or selections by users are typically brief, containing only the keywords expressing the desired effect (the initial prompt), such as the characteristics of a person or the style of an image. Modifiers need to be added to these keywords, expanding on the main features of the image, such as descriptions of actions, clothing, and environment. Elements that enhance the image's beauty and harmony, such as descriptions of atmosphere, lighting, color, and composition, can also be added. Furthermore, the generated second prompts must adhere to a specific format, conforming to conciseness, coherence, and clarity. This ensures that the resulting target image is richer, more accurate, and aligns with the user's intent. If the initial prompt is of high quality and can be directly used to generate the target image, a second prompt can be omitted, improving interaction efficiency while maintaining image quality. For excessively long, erroneous, or contradictory initial prompts, abbreviations and rewrites can be used to generate high-quality second prompts, ensuring the final target image is accurate, harmonious, and better reflects the user's intent.
[0150] This disclosure also provides an image generation apparatus, which will be described below with reference to FIG6.
[0151] Figure 6 is a structural diagram of some embodiments of the image generation apparatus of this disclosure. As shown in Figure 6, the image generation apparatus 60 of this embodiment includes: a first display module 610, a second display module 620, a third display module 630, and a fourth display module 640.
[0152] The first display module 610 is configured to display a prompt information input interface.
[0153] The second display module 620 is configured to display the first prompt information on the prompt information input interface in response to obtaining the first prompt information, wherein the first prompt information includes the first description information of the target image.
[0154] The third display module 630 is configured to respond to displaying a second prompt information generated based on the first prompt information, wherein the second prompt information includes second descriptive information of the features of the person in the target image and third descriptive information of the image features other than the features of the person.
[0155] The fourth display module 640 is configured to display a target image generated based on the image to be processed and the second prompt information in response to a user's triggering of the image generation function.
[0156] In some embodiments, the third display module 630 is configured to display one or more third prompt messages generated based on the first prompt message and the target adjustment method; and in response to the user's selection of the target prompt message in one or more third prompt messages, to display second prompt messages determined based on the target prompt message and the first prompt message.
[0157] In some embodiments, in response to the target adjustment method being expansion, one or more third prompt messages include supplementary description information generated based on the first description information; or in response to the target adjustment method being abbreviation, one or more third prompt messages include description information generated from the abbreviation of the first description information; or in response to the target adjustment method being rewrite, one or more third prompt messages include description information generated from the rewrite of the first description information.
[0158] In some embodiments, the third display module 630 is further configured to perform at least one of the following: displaying first guidance information during the generation of one or more third prompt messages, wherein the first guidance information is used to indicate that one or more third prompt messages are being generated; and displaying second guidance information, wherein the second guidance information is used to indicate the target adjustment method corresponding to one or more third prompt messages.
[0159] In some embodiments, the third display module 630 is further configured to display an added control; in response to a user's triggering of the added control, to display one or more new third prompt messages generated based on the first prompt message and the target adjustment method.
[0160] In some embodiments, the second display module 620 is configured to obtain the first prompt information in response to a user inputting first prompt information in the input area of the prompt information input interface; or to display multiple candidate prompt information and obtain the first candidate prompt information as the first prompt information in response to a user selecting the first candidate prompt information from the multiple candidate prompt information.
[0161] In some embodiments, the second display module 620 is configured to, in response to a user inputting initial prompt information in the input area of the prompt information input interface, display the initial prompt information in the prompt information input interface, display multiple candidate prompt information, and, in response to the user selecting a first candidate prompt information among the multiple candidate prompt information, obtain a combination of the initial prompt information and the first candidate prompt information as the first prompt information; or display multiple candidate prompt information, and, in response to the user selecting a first candidate prompt information among the multiple candidate prompt information, display the first candidate prompt information in the prompt information input interface, and, in response to the user adjusting the first candidate prompt information, obtain the adjusted prompt information as the first prompt information.
[0162] In some embodiments, the second display module 620 is configured to display the first candidate prompt and the supplementary prompt in the prompt input interface in response to the user adding supplementary prompt based on the first candidate prompt; and to display the second candidate prompt and the supplementary prompt in the prompt input interface as the adjusted prompt as in response to the user replacing the first candidate prompt by selecting the second candidate prompt from among multiple candidate prompts.
[0163] In some embodiments, the second display module 620 is further configured to perform at least one of the following: in response to a user's selection of a first candidate prompt among a plurality of candidate prompts, displaying the first candidate prompt with a preset effect in the display area of the plurality of candidate prompts, wherein each candidate prompt includes fourth descriptive information about the characteristics of a person or the style of an image; displaying a switching control, and in response to a user's triggering of the switching control, displaying the plurality of candidate prompts after the switch.
[0164] In some embodiments, the third display module 630 is further configured to utilize a machine learning model to: parse the first prompt information and determine the parsing result; determine whether to adjust the first prompt information based on the parsing result; in response to adjusting the first prompt information, determine a target adjustment method based on the parsing result; and generate one or more third prompt information messages based on the first prompt information and the target adjustment method.
[0165] In some embodiments, the third display module 630 is further configured to display third guidance information on the prompt information input interface in response to not adjusting the first prompt information, wherein the third guidance information is used to indicate that the first prompt information is used directly to generate the target image without adjustment.
[0166] In some embodiments, the parsing result includes semantic information of the first description information, and the third display module 630 is configured to use a machine learning model to determine whether the first description information only includes description information of the features of a person or description information of the image style based on the semantic information of the first description information; in response to the first description information only including description information of the features of a person or description information of the image style, the target adjustment method is determined to be expansion.
[0167] In some embodiments, in response to the first description information including only the description information of a person's features, the target adjustment method is to expand and determine to expand the description information of the person's features. One or more third description information includes supplementary description information of the person's features that matches the description information of the person's features and supplementary description information of the scene features that matches the description information of the person's features, wherein the description information of the person's features and the supplementary description information of the person's features are combined as the second description information, and the supplementary description information of the scene features is the third description information; and / or in response to the first description information including only the description information of the person's features, the target method is to expand and determine not to expand the description information of the person's features. One or more third description information includes supplementary description information of the scene features that matches the description information of the person's features, wherein the description information of the person's features is the second description information, and the supplementary description information of the scene features is the third description information.
[0168] In some embodiments, in response to the first description information including only image style description information, the target adjustment method is expansion, and one or more third description information includes supplementary description information of character features that match the image style generated based on the image style description information, and supplementary description information of scene features that match the image style, wherein the supplementary description information of character features serves as the second description information, and the image style description information and the supplementary description information of scene features serve as the third description information.
[0169] In some embodiments, the third display module 630 is configured to utilize a machine learning model to: in response to the first description information including description information of the features of a person or description information of the image style, and description information of screen features other than the features of the person, determine whether the description information of the features of the person or description information of the image style matches the description information of screen features other than the features of the person; in response to the description information of the features of the person or description information of the image style not matching the description information of screen features other than the features of the person, determine that the adjustment method for the target is rewriting.
[0170] In some embodiments, the parsing result includes the length of the first description information, and the third display module 630 is configured to use a machine learning model to: determine whether the length of the first description information exceeds a preset length; and in response to the length of the first description information exceeding the preset length, determine that the target adjustment method is an abbreviation.
[0171] In some embodiments, the parsing result includes the expression of the first descriptive information, and the third display module 630 is configured to use a machine learning model to: determine whether the expression of the first descriptive information conforms to the target expression; and in response to the first descriptive information not conforming to the target expression, determine the target adjustment method as rewriting, wherein one or more third descriptive information conforms to the target expression.
[0172] It should be noted that the above-described units (modules) are logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of both. In actual implementation, the above-described units can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.). Furthermore, the units shown in the accompanying drawings with dashed lines indicate that these units may not actually exist, and the operations / functions they perform can be implemented by the processing circuitry itself.
[0173] In addition, although not shown, the device may also include a memory that can store various information generated by the device and its constituent units during operation, programs and data used for operation, data to be transmitted by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory may also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in a manner known in the art, such as including communication components such as antenna arrays and / or radio frequency links, various types of interfaces, communication units, etc. These will not be described in detail here. Furthermore, the device may also include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, etc. These will not be described in detail here.
[0174] Some embodiments of this disclosure also provide an electronic device. Figure 7 shows a block diagram of some embodiments of the electronic device of this disclosure. For example, in some embodiments, the electronic device 70 can be various types of devices, such as mobile terminals including, but not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 70 may include a display panel for displaying data and / or execution results utilized in the scheme according to this disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a planar panel, but also a curved panel, or even a spherical panel.
[0175] As shown in FIG. 7, the electronic device 70 of this embodiment includes a memory 71 and a processor 72 coupled to the memory 71. It should be noted that the components of the electronic device 70 shown in FIG. 7 are merely exemplary and not limiting; the electronic device 70 may also have other components depending on the actual application requirements. The processor 72 can control other components in the electronic device 70 to perform desired functions.
[0176] In some embodiments, memory 71 is used to store one or more computer-readable instructions. When processor 72 executes the computer-readable instructions, the computer-readable instructions are executed by processor 72 to implement the method according to any of the above embodiments. For specific implementations and related explanations of the various steps of the method, please refer to the above embodiments; repeated details will not be elaborated here.
[0177] For example, processor 72 and memory 71 can communicate with each other directly or indirectly. For example, processor 72 and memory 71 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. Processor 72 and memory 71 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0178] For example, processor 72 can be embodied in various suitable processors, processing devices, such as central processing unit (CPU), graphics processing unit (GPU), network processor (NP), etc.; it can also be digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The central processing unit (CPU) can be an x86 or ARM architecture, etc. For example, memory 71 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Memory 71 can include, for example, system memory, which stores, for example, the operating system, application programs, boot loader, database, and other programs. Various application programs and various data can also be stored in the storage medium.
[0179] Furthermore, according to some embodiments of this disclosure, various operations / processes according to this disclosure, implemented via software and / or firmware, can install programs constituting the software from a storage medium or network onto a computer system with a dedicated hardware architecture, such as the computer system (or electronic device) 80 shown in FIG. 8. When various programs are installed, the computer system is capable of performing various functions, including those described above. FIG. 8 is a block diagram illustrating an example structure of a computer system that may be employed in an embodiment of this disclosure.
[0180] In Figure 8, the Central Processing Unit (CPU) 801 performs various processes based on a program stored in the Read-Only Memory (ROM) 802 or a program loaded from the Storage Section 808 into the Random Access Memory (RAM) 803. The RAM 803 also stores data required as needed when the CPU 801 performs various processes. The CPU is merely exemplary and can be other types of processors, such as the various processors described above. The ROM 802, RAM 803, and Storage Section 808 can be various forms of computer-readable storage media, as described below. It should be noted that although the ROM 802, RAM 803, and Storage Section 808 are shown separately in Figure 8, one or more of them can be combined or located in the same or different memories or storage modules.
[0181] CPU 801, ROM 802 and RAM 803 are interconnected via bus 804. Input / output interface 805 is also connected to bus 804.
[0182] The following components are connected to the input / output interface 805: input section 806, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 807, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 808, including hard disks, magnetic tapes, etc.; and communication section 809, including network interface cards such as LAN cards, modems, etc. The communication section 809 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the computer system 80 shown in Figure 8 communicate via bus 804, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0183] As needed, drive 810 is also connected to input / output interface 805. Removable media 811, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 810 as needed, so that computer programs read from them can be installed into storage section 808 as needed.
[0184] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or from a storage medium such as removable media 811.
[0185] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the CPU 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0186] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0187] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0188] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods of any of the above embodiments. For example, the instructions may be embodied in computer program code.
[0189] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0191] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0192] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0193] According to some embodiments of this disclosure, an image generation method is provided, comprising: displaying a prompt information input interface; in response to obtaining first prompt information, displaying the first prompt information on the prompt information input interface, wherein the first prompt information includes first descriptive information of a target image; displaying second prompt information generated based on the first prompt information, wherein the second prompt information includes second descriptive information of features of a person in the target image and third descriptive information of image features other than features of the person; and in response to a user triggering an image generation function, displaying a target image generated based on an image to be processed and the second prompt information.
[0194] In some embodiments, displaying a second prompt message generated based on the first prompt message includes: displaying one or more third prompt messages generated based on the first prompt message and the target adjustment method; and, in response to the user's selection of a target prompt message among the one or more third prompt messages, displaying a second prompt message determined based on the target prompt message and the first prompt message.
[0195] In some embodiments, in response to the target adjustment method being expansion, one or more third prompt messages include supplementary description information generated based on the first description information; or in response to the target adjustment method being abbreviation, one or more third prompt messages include description information generated from the abbreviation of the first description information; or in response to the target adjustment method being rewrite, one or more third prompt messages include description information generated from the rewrite of the first description information.
[0196] In some embodiments, the image generation method further includes at least one of the following: displaying first guidance information during the generation of one or more third prompts, wherein the first guidance information is used to indicate that one or more third prompts are being generated; and displaying second guidance information, wherein the second guidance information is used to indicate the target adjustment method corresponding to one or more third prompts.
[0197] In some embodiments, the image generation method further includes displaying an add control; in response to a user triggering the add control, displaying one or more new third prompts generated based on a first prompt and a target adjustment method.
[0198] In some embodiments, obtaining the first prompt information includes: obtaining the first prompt information in response to the user inputting the first prompt information in the input area of the prompt information input interface; or displaying multiple candidate prompt information, and obtaining the first candidate prompt information as the first prompt information in response to the user selecting the first candidate prompt information from the multiple candidate prompt information.
[0199] In some embodiments, obtaining the first prompt information includes: in response to a user inputting initial prompt information in the input area of the prompt information input interface, displaying the initial prompt information in the prompt information input interface, displaying multiple candidate prompt information, and in response to the user selecting a first candidate prompt information from the multiple candidate prompt information, obtaining a combination of the initial prompt information and the first candidate prompt information as the first prompt information; or displaying multiple candidate prompt information, in response to the user selecting a first candidate prompt information from the multiple candidate prompt information, displaying the first candidate prompt information in the prompt information input interface, and in response to the user adjusting the first candidate prompt information, obtaining the adjusted prompt information as the first prompt information.
[0200] In some embodiments, in response to a user's adjustment of the first candidate prompt information, obtaining the adjusted prompt information as the first prompt information includes: in response to a user adding supplementary prompt information based on the first candidate prompt information, displaying the first candidate prompt information and the supplementary prompt information on the prompt information input interface; in response to a user replacing the first candidate prompt information with a second candidate prompt information from a plurality of candidate prompt information, displaying the second candidate prompt information and the supplementary prompt information on the indicated prompt information input interface as the adjusted prompt information.
[0201] In some embodiments, the image generation method further includes at least one of the following: in response to a user's selection of a first candidate prompt among a plurality of candidate prompts, displaying the first candidate prompt with a preset effect in the display area of the plurality of candidate prompts, wherein each candidate prompt includes fourth descriptive information about the characteristics of a person or the style of an image; displaying a switching control, and in response to a user's triggering of the switching control, displaying the plurality of candidate prompts after switching.
[0202] In some embodiments, displaying one or more third prompts generated based on the first prompt and the target adjustment method includes using a machine learning model to: parse the first prompt and determine the parsing result; determine whether to adjust the first prompt based on the parsing result; in response to adjusting the first prompt, determine the target adjustment method based on the parsing result; and generate one or more third prompts based on the first prompt and the target adjustment method.
[0203] In some embodiments, the image generation method further includes displaying third guidance information on the prompt information input interface in response to not adjusting the first prompt information, wherein the third guidance information is used to indicate that the first prompt information is not adjusted and is directly used to generate the target image.
[0204] In some embodiments, the parsing result includes semantic information of the first description information, and determining the target adjustment method based on the parsing result includes: determining whether the first description information only includes description information of the features of a person or description information of the image style based on the semantic information of the first description information; and determining the target adjustment method as expansion in response to the first description information only including description information of the features of a person or description information of the image style.
[0205] In some embodiments, in response to the first description information including only the description information of a person's features, the target adjustment method is to expand and determine to expand the description information of the person's features. One or more third description information includes supplementary description information of the person's features that matches the description information of the person's features and supplementary description information of the scene features that matches the description information of the person's features, wherein the description information of the person's features and the supplementary description information of the person's features are combined as the second description information, and the supplementary description information of the scene features is the third description information; and / or in response to the first description information including only the description information of the person's features, the target method is to expand and determine not to expand the description information of the person's features. One or more third description information includes supplementary description information of the scene features that matches the description information of the person's features, wherein the description information of the person's features is the second description information, and the supplementary description information of the scene features is the third description information.
[0206] In some embodiments, in response to the first description information including only image style description information, the target adjustment method is expansion, and one or more third description information includes supplementary description information of character features that match the image style generated based on the image style description information, and supplementary description information of scene features that match the image style, wherein the supplementary description information of character features serves as the second description information, and the image style description information and the supplementary description information of scene features serve as the third description information.
[0207] In some embodiments, determining the target adjustment method based on the parsing result further includes: in response to the first description information including description information of the character's features or description information of the image style, and description information of the screen features other than the character's features, determining whether the description information of the character's features or description information of the image style matches the description information of the screen features other than the character's features; in response to the description information of the character's features or description information of the image style not matching the description information of the screen features other than the character's features, determining that the target adjustment method is rewriting.
[0208] In some embodiments, the parsing result includes the length of the first description information, and determining the target adjustment method based on the parsing result includes: determining whether the length of the first description information exceeds a preset length; in response to the length of the first description information exceeding the preset length, determining that the target adjustment method is an abbreviation.
[0209] In some embodiments, the parsing result includes the expression of the first descriptive information, and determining the target adjustment method based on the parsing result includes: determining whether the expression of the first descriptive information conforms to the target expression method; in response to the first descriptive information not conforming to the target expression method, determining the target adjustment method as rewriting, wherein one or more third descriptive information conforms to the target expression method.
[0210] According to other embodiments of this disclosure, an image generation apparatus is provided, comprising: a first display module configured to display a prompt information input interface; a second display module configured to display the first prompt information on the prompt information input interface in response to acquiring the first prompt information, wherein the first prompt information includes first descriptive information of a target image; a third display module configured to display second prompt information generated based on the first prompt information in response to displaying the second prompt information, wherein the second prompt information includes second descriptive information of features of a person in the target image and third descriptive information of image features other than the features of the person; and a fourth display module configured to display a target image generated based on the image to be processed and the second prompt information in response to a user triggering the image generation function.
[0211] According to further embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor for storing instructions, which, when executed by the processor, cause the processor to perform an image generation method according to any embodiment of the present disclosure.
[0212] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, performs an image generation method according to any embodiment of the present disclosure.
[0213] According to further embodiments of the present disclosure, a computer program product is provided, comprising: instructions that, when executed by a processor, implement the image generation method of any embodiment of the present disclosure.
[0214] According to further embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, implement the image generation method of any embodiment of the present disclosure.
[0215] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0216] Many specific details are set forth in the description provided herein. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description.
[0217] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0218] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. An image generation method, comprising: Display prompt message input interface; In response to obtaining the first prompt information, the first prompt information is displayed on the prompt information input interface, wherein the first prompt information includes first description information of the target image; Display a second prompt based on the first prompt, wherein the second prompt includes a second description of the features of the person in the target image and a third description of the image features other than the features of the person; In response to the user's triggering of the image generation function, the target image generated based on the image to be processed and the second prompt information is displayed.
2. The image generation method according to claim 1, wherein, The second prompt information generated based on the first prompt information includes: Display one or more third prompts generated based on the first prompt and the target adjustment method; In response to the user's selection of a target prompt from one or more of the third prompts, the second prompt is displayed based on the target prompt and the first prompt.
3. The image generation method according to claim 2, wherein: In response to the target adjustment method being expansion, the one or more third prompt messages include supplementary description information generated based on the first description information; or In response to the target adjustment method being an abbreviation, the one or more third prompt messages include descriptive information generated from the abbreviation of the first descriptive information; or In response to the target adjustment method being rewritten, the one or more third prompt messages include description information generated by rewriting the first description information.
4. The image generation method according to claim 2 or 3, further comprising at least one of the following: During the generation of the one or more third prompt messages, first guidance information is displayed, wherein... The first guidance information is used to indicate that one or more third prompt messages are being generated; Display second guidance information, wherein the second guidance information is used to indicate the target adjustment method corresponding to the one or more third prompts.
5. The image generation method according to any one of claims 2-4, further comprising: Display added controls; In response to the user's triggering of the added control, one or more new third prompt messages are displayed based on the first prompt message and the target adjustment method.
6. The image generation method according to any one of claims 1-5, wherein, The acquisition of the first prompt information includes: In response to the user inputting the first prompt information in the input area of the prompt information input interface, the first prompt information is obtained; or Display multiple candidate prompts, and in response to the user's selection of a first candidate prompt among the multiple candidate prompts, obtain the first candidate prompt as the first prompt.
7. The image generation method according to any one of claims 1-5, wherein, The acquisition of the first prompt information includes: In response to the user inputting initial prompt information in the input area of the prompt information input interface, the initial prompt information is displayed on the prompt information input interface, and multiple candidate prompt information is displayed. In response to the user selecting a first candidate prompt information from the multiple candidate prompt information, a combination of the initial prompt information and the first candidate prompt information is obtained as the first prompt information; or Display multiple candidate prompts, and in response to the user's selection of the first candidate prompt among the multiple candidate prompts, display the first candidate prompt on the prompt input interface, and in response to the user's adjustment of the first candidate prompt, obtain the adjusted prompt as the first prompt.
8. The image generation method according to claim 7, wherein, The step of responding to the user's adjustment of the first candidate prompt information and obtaining the adjusted prompt information as the first prompt information includes: In response to the user adding supplementary prompt information based on the first candidate prompt information, the first candidate prompt information and the supplementary prompt information are displayed on the prompt information input interface; In response to the user replacing the first candidate prompt information by selecting the second candidate prompt information from the plurality of candidate prompt information, the second candidate prompt information and the supplementary prompt information are displayed on the prompt information input interface as the adjusted prompt information.
9. The image generation method according to any one of claims 6-8, further comprising at least one of the following: In response to the user's selection of the first candidate prompt among the plurality of candidate prompts, the first candidate prompt is displayed in the display area of the plurality of candidate prompts with a preset effect, wherein, Each of the multiple candidate prompts includes a fourth descriptive information about the character's features or image style; Display a switching control, and in response to the user's triggering of the switching control, display multiple candidate prompt messages after the switch.
10. The image generation method according to any one of claims 2-9, wherein, The display of one or more third prompts generated based on the first prompt and the target adjustment method includes utilizing a machine learning model: The first prompt message is parsed to determine the parsing result; Based on the analysis results, determine whether to adjust the first prompt message; In response to adjusting the first prompt information, the target adjustment method is determined based on the parsing result; The one or more third prompt messages generated based on the first prompt message and the target adjustment method.
11. The image generation method according to claim 10, further comprising: In response to not adjusting the first prompt information, a third guidance information is displayed on the prompt information input interface, wherein the third guidance information is used to indicate that the first prompt information is not adjusted and is directly used to generate the target image.
12. The image generation method according to claim 10 or 11, wherein, The parsing result includes the semantic information of the first descriptive information, and determining the target adjustment method based on the parsing result includes: Based on the semantic information of the first description information, determine whether the first description information only includes description information of the characteristics of the person or description information of the image style; In response to the first description information being either a description information that only includes the characteristics of the person or a description information that includes the image style, the target adjustment method is determined to be expansion.
13. The image generation method according to any one of claims 3-12, wherein, In response to the first descriptive information only including descriptive information of the character's features, the target adjustment method is to expand and determine the expansion of the descriptive information of the character's features. The one or more third descriptive information includes supplementary descriptive information of the character's features that matches the descriptive information of the character's features and supplementary descriptive information of the scene features that matches the character's features. The descriptive information of the character's features and the supplementary descriptive information of the character's features are combined as the second descriptive information, and the supplementary descriptive information of the scene features is the third descriptive information; and / or In response to the first descriptive information only including descriptive information of the character's features, the target method is expansion, and it is determined that the descriptive information of the character's features will not be expanded. The one or more third descriptive information includes supplementary descriptive information of the scene features that match the descriptive information of the character's features, wherein the descriptive information of the character's features serves as the second descriptive information, and the supplementary descriptive information of the scene features serves as... This refers to the third descriptive information.
14. The image generation method according to any one of claims 3-13, wherein, In response to the first description information including only image style description information, the target adjustment method is expansion, and the one or more third description information includes supplementary description information of the character's features that match the image style generated based on the image style description information, and supplementary description information of the scene features that match the image style, wherein the supplementary description information of the character's features serves as the second description information, and the image style description information and the supplementary description information of the scene features serve as the third description information.
15. The image generation method according to claim 12, wherein, The step of determining the target adjustment method based on the analysis result also includes: In response to the first description information including description information of the characteristics of the person or description information of the image style, and description information of the image features other than the characteristics of the person, it is determined whether the description information of the characteristics of the person or description information of the image style matches the description information of the image features other than the characteristics of the person. If the description information of the character's features or the description information of the image style does not match the description information of the screen features other than the character's features, the adjustment method for the target is determined to be rewriting.
16. The image generation method according to any one of claims 10-15, wherein, The parsing result includes the length of the first descriptive information, and determining the target adjustment method based on the parsing result includes: Determine whether the length of the first description information exceeds a preset length; In response to the fact that the length of the first description information exceeds the preset length, the target adjustment method is determined to be an abbreviation.
17. The image generation method according to any one of claims 10-16, wherein, The parsing result includes the expression method of the first descriptive information, and determining the target adjustment method based on the parsing result includes: Determine whether the expression method of the first descriptive information conforms to the target expression method; In response to the fact that the expression of the first descriptive information does not conform to the target expression, the target adjustment method is determined to be rewriting, wherein the one or more third descriptive information conforms to the target expression.
18. An image generation apparatus, comprising: The first display module is configured to display a prompt information input interface; The second display module is configured to display the first prompt information on the prompt information input interface in response to obtaining the first prompt information, wherein the first prompt information includes first description information of the target image; The third display module is configured to respond to displaying a second prompt message generated based on the first prompt message. The second prompt information includes second descriptive information about the features of the person in the target image and third descriptive information about the image features other than the features of the person. The fourth display module is configured to display the target image generated based on the image to be processed and the second prompt information in response to the user's triggering of the image generation function.
19. An electronic device comprising: processor; as well as A memory coupled to the processor is used to store instructions that, when executed by the processor, cause the processor to perform the image generation method as described in any one of claims 1-17.
20. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the program implements the image generation method according to any one of claims 1-17.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN114266840A
Manuscript generation method and related device, electronic equipment and storage medium
CN117033567A
Prompt word expanding and writing method and device, storage medium and electronic equipment
CN117573913A
Image generation method and device, electronic equipment and storage medium
CN117853600A
Method and device for generating image based on text, electronic equipment and storage medium
CN118037896A