Interaction method, apparatus, electronic device, and storage medium
By segmenting images and editing the generated prompts, the problem of ordinary users struggling to create accurate prompts is solved, thus achieving auxiliary effects in image or video generation.
Patent Information
- Application Number
- CN202411897527.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2026-06-26
AI Technical Summary
Ordinary users lack a deep understanding of AI technology, making it difficult for them to construct accurate and effective generated prompts. This results in the generated images or videos not matching user expectations, limiting the popularization and efficient application of generation technology.
By segmenting the first image, multiple sub-images are obtained, and corresponding generation prompts are generated based on the sub-images to help users understand and edit these prompts to generate the target image or video.
It helps users understand the scope of the generated prompts, demonstrates how the prompts are expressed, and assists users in constructing accurate and effective prompts to achieve the purpose of image or video generation.
Smart Images

Figure CN122289430A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to an interaction method, device, electronic device, and storage medium. Background Technology
[0002] With the rapid development of technology, image and video generation technology has become a significant achievement in today's digital field. Utilizing advanced algorithm models, it is possible to generate corresponding images or videos based on specific input instructions or parameters. This technology has brought innovative opportunities to many fields.
[0003] However, when using generative techniques to generate images or videos, users typically need to input specific prompts to guide the generative model in producing images or videos that meet their needs. But for ordinary users without programming backgrounds or familiarity with AI technology, their lack of in-depth understanding of AI capabilities makes it difficult to accurately express their creative needs, hindering the creation of precise and effective prompts. When the prompts are inaccurate or incomplete, the final image or video content generated by the generative model will not match the user's expectations, severely limiting the widespread adoption and efficient application of generative techniques. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides an interaction method, device, electronic device, and storage medium.
[0005] Firstly, this disclosure provides an interaction method, including:
[0006] The first image is segmented to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements;
[0007] Based on the first image, a first text is obtained to describe the image content of the first image;
[0008] Based on the sub-image, the first text used to describe the image content of the first image is decomposed to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; the generated prompt information is used to describe the image content of the sub-image corresponding to it;
[0009] Based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image, a target image or target video is generated.
[0010] Secondly, this disclosure also provides an interactive device, including:
[0011] A segmentation module is used to segment a first image to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements;
[0012] The text determination module is used to obtain first text describing the image content of the first image based on the first image;
[0013] The relationship determination module is used to decompose the first text used to describe the image content of the first image based on the sub-image, to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; the generated prompt information is used to describe the image content of the sub-image corresponding to it;
[0014] The generation module is used to generate a target image or target video based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image.
[0015] Thirdly, this disclosure also provides an electronic device, the electronic device comprising:
[0016] One or more processors;
[0017] Storage device for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the interaction method as described above.
[0019] Fourthly, this disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the interaction method described above.
[0020] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0021] The technical solution provided in this disclosure involves segmenting a first image to obtain multiple sub-images. Each first image includes multiple elements, and different sub-images correspond to different elements. Based on the first image, first text describing the image content of the first image is obtained. Based on the sub-images, the first text describing the image content of the first image is decomposed to obtain multiple generated prompt messages, and a correspondence between the generated prompt messages and the sub-images. The generated prompt messages are used to describe the image content of their corresponding sub-images. Based on the editing results of the generated prompt messages and the correspondence between the generated prompt messages and the sub-images, a target image or target video is generated. By adopting the technical solution provided in this disclosure, users can understand the scope of the generated prompt messages' effect on the first image. Furthermore, by using the generated prompt messages as examples, the presentation method of the generated prompt messages is demonstrated to users, which can help users understand how to construct accurate and effective generated prompt messages and assist users in completing the image or video generation. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart of an interaction method provided in an embodiment of this disclosure;
[0025] Figure 2 A schematic diagram of a first image provided for an embodiment of this disclosure;
[0026] Figures 3-7 A schematic diagram illustrating several target pages provided in embodiments of this disclosure;
[0027] Figure 8 This is a schematic diagram of the structure of an interactive device according to an embodiment of the present disclosure;
[0028] Figure 9 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0029] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0030] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0031] Figure 1 This flowchart illustrates an interaction method provided in an embodiment of the present disclosure. This embodiment is applicable to situations where a user is assisted in understanding how to generate prompt information when settings are enabled on a client side. The method can be executed by an interactive device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a terminal, including but not limited to smartphones, PDAs, tablets, wearable devices with displays, desktop computers, laptops, all-in-one computers, and smart home devices. Alternatively, this embodiment can be applied to situations where a user is assisted in understanding how to generate prompt information when settings are enabled on a server side. This method can also be executed by an interactive device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a server.
[0032] like Figure 1 As shown, the method may specifically include:
[0033] S110. The first image is segmented to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements.
[0034] The first image is a user-specified image that can be used to demonstrate to the user the scope and / or effect of the generated prompt information on the first image. In some scenarios, it can be further configured to allow users to modify the first image by modifying or supplementing the generated prompt information. The first image can specifically be an image generated by the user, an image taken by the user, an image uploaded by the user, or an image selected by the user from the network. The first image can also be a frame from a video.
[0035] Elements, for example, can refer to individuals or units in the first image that possess clear semantic information, are identifiable, and may be countable or uncountable. These elements, through their respective attribute characteristics, collectively constitute the visual content of the image and are the basic units of image processing and analysis. In this application, an element can be the smallest unit that can be regenerated or modified. Specific elements in the first image may include people, buildings, trees, vehicles, animals, sky, or grass, etc., in the first image. For example, in... Figure 2 The first image given includes elements such as a person, a puppy, a frisbee, grass, mountains, and sky.
[0036] There are various specific implementation methods for this step, and this application does not limit them. For example, the implementation method of this step includes: segmenting the first image using an image segmentation model to obtain multiple sub-images, each sub-image corresponding to one element. For example, for... Figure 2 The first image is segmented to obtain six sub-images. These six sub-images correspond to a person, a puppy, a frisbee, grass, a mountain, and the sky, respectively.
[0037] S120. Based on the first image, obtain the first text used to describe the image content of the first image.
[0038] The first text can be, for example, text information obtained by understanding a first image. For instance, the first image can be input into a model with image understanding capabilities to obtain the first text.
[0039] S130. Based on the sub-image, the first text used to describe the image content of the first image is decomposed to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; the generated prompt information is used to describe the image content of the corresponding sub-image.
[0040] Typically, the first text includes descriptive statements about each element in the first image. Since sub-images correspond to elements, this step essentially involves disassembling the first text to determine which element each statement describes, thereby establishing the correspondence between the descriptive statements and the sub-images. The descriptive statement corresponding to a particular element can serve as a generation prompt for generating a sub-image that includes that element. In other words, the generation prompt is the sub-text obtained by disassembling the first text.
[0041] For example, assuming based on Figure 2 The first text derived from the first image is: A person wearing a gray long-sleeved shirt and khaki trousers, with black and white sneakers, is bending over and extending their right hand, seemingly throwing a yellow frisbee to a dog. The dog is white, wearing a black collar, with its mouth open, its eyes fixed on the frisbee, appearing very excited and focused, as if preparing to catch it. The background is a vast natural landscape, with rolling mountains covered in green vegetation in the distance, and a clear blue sky dotted with white clouds. By deconstructing this first text, we can see that "a person wearing a gray long-sleeved shirt and khaki trousers, with black and white sneakers, is bending over and extending their right hand" describes the element of "person" in the first image. In other words, "a person wearing a gray long-sleeved shirt and khaki trousers, with black and white sneakers, is bending over and extending their right hand" can serve as a cue for generating the element of "person" in the first image.
[0042] S140. Based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-images, generate the target image or target video.
[0043] The editing result of the generated prompt information can be, for example, editing of all the generated prompt information obtained in S130, or editing of some of the generated prompt information. Specifically, it can include modifying the text content of the generated prompt information, and / or determining a reference image associated with the target generated prompt information.
[0044] If an image needs to be generated, the implementation method of this step may include: in response to the modification operation of the target generation prompt information in multiple generation prompt information, generating a new sub-image corresponding to the target generation prompt information based on the modified target generation prompt information; replacing the sub-image corresponding to the target generation prompt information in the first image with the new sub-image to obtain the target image.
[0045] If video generation is required, this step can be implemented by: responding to a modification operation on the target generation prompt information among multiple generation prompts, generating a new sub-image corresponding to the target generation prompt information based on the modified target generation prompt information; replacing the sub-image corresponding to the target generation prompt information in the first image with the new sub-image to obtain a third image; and generating the target video based on the third image. Alternatively, the first image may be an image frame from the original video, and other image frames in the original video may be adjusted based on the third image to obtain the target video.
[0046] The above technical solution involves segmenting a first image to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements; based on the first image, first text describing the image content of the first image is obtained; based on the sub-images, the first text describing the image content of the first image is decomposed to obtain multiple generated prompt messages, and the correspondence between the generated prompt messages and the sub-images; the generated prompt messages are used to describe the image content of their corresponding sub-images; based on the editing results of the generated prompt messages and the correspondence between the generated prompt messages and the sub-images, a target image or target video is generated. By adopting the technical solution provided in this disclosure, on the one hand, users can understand the scope of the generated prompt messages on the first image; on the other hand, by using the generated prompt messages as examples, the method of expressing the generated prompt messages is demonstrated to users, which can help users understand how to construct accurate and effective generated prompt messages and assist users in completing the image or video generation.
[0047] Based on the above technical solution, the method may optionally include: displaying the correspondence between the generated prompt information and the sub-image on the target page.
[0048] The target page could be a page that displays the correspondence between generated prompts and sub-images. By displaying this correspondence on the target page, users can clearly understand the relationship between generated prompts and sub-images, helping them quickly understand which sub-image will change if a certain generated prompt is modified, and which generated prompt needs to be modified if a sub-image needs to be modified.
[0049] There are various ways to implement this step, and this application does not limit the specific implementation. For example, in some embodiments, a connecting line can be set up on the target page to link the generated prompt information and its corresponding sub-image. This connecting line allows the user to clearly understand the correspondence between the generated prompt information and the sub-image.
[0050] In other embodiments, the specific implementation method of this step may include: displaying a first image in a first area of the target page; displaying generation prompt information in a second area of the target page; and adjusting the display state of the generation prompt information corresponding to the target element in the second area to a preset state in response to the selection operation of the target element in the first image.
[0051] The target element could be, for example, an element selected by the user from multiple elements included in the first image. In some scenarios, if the user selects a target element, it means that the user may subsequently want to regenerate or modify the target element.
[0052] Adjusting the display state of the generated prompt information corresponding to the target element in the second region to a preset state means highlighting the generated prompt information corresponding to the target element in the second region. This preset state can be achieved by changing the color and brightness of the generated prompt information corresponding to the target element, adding a border to the generated prompt information corresponding to the target element, changing the border color, changing the background color, etc. By setting the display state of the generated prompt information corresponding to the target element in the second region to a preset state in response to the selection operation of the target element in the first image, users can clearly understand the correspondence between the generated prompt information and the sub-image.
[0053] Furthermore, in the second area of the target page, a generation prompt message may be displayed, which may include:
[0054] Based on the preset hierarchical structure division rules, the generated prompt information is sorted out to obtain the hierarchical structure of the generated prompt information; the preset hierarchical structure division rules include a first level and a second level. The first level includes an image description layer and an action description layer; the second level of the image description layer and / or action description layer includes at least one of the main element, the object element, and the background; the hierarchical structure of the generated prompt information is displayed in the second area of the target page.
[0055] The preset hierarchical structure division rule can be, for example, a preset framework for classifying and organizing generated prompt information, which organizes the generated prompt information through different levels. There can be multiple preset hierarchical structure division rules, and this application does not limit them. For example, the preset hierarchical structure division rule includes a first level and a second level. The first level includes an image description layer and an action description layer; the second level of the image description layer and / or the action description layer includes at least one of a subject element, an object element, and a background. A subject element can, for example, refer to an element that occupies a primary position in the first image and undertakes the main action or behavior. It is the core part of the first image and is usually the focus of the user's attention. An object element can, for example, be relative to the subject element and is the object of the subject element's action. It assists the subject in the scene, helping to showcase the subject's actions, state, or relationships. A background can, for example, refer to the elements remaining in the first image after the subject and object elements, usually used to provide spatial and environmental support for the subject and object elements, creating atmosphere, suggesting location and time, etc.
[0056] The hierarchical structure of the generated prompts can be, for example, a hierarchical information organization form obtained by sorting through multiple generated prompts from the first article based on preset hierarchical structure division rules. It presents the content clearly and logically by classifying it according to preset hierarchical levels.
[0057] For example, if the preset hierarchical structure division rule includes a first level and a second level, the first level includes an image description layer and an action description layer; the second level of the image description layer and / or the action description layer includes at least one of a subject element, an object element, and a background. Based on this preset hierarchical structure division rule, after sorting through the multiple generated prompts obtained from the first text in the previous example, the hierarchical structure of the generated prompts is obtained. The hierarchical structure of the generated prompts is as follows:
[0058] First level - Image description level:
[0059] Second level - Main elements
[0060] Subject 1: The figure is wearing a gray long-sleeved shirt, khaki trousers, and black and white sneakers.
[0061] Subject 2: A white dog wearing a black collar.
[0062] Second level – Object elements
[0063] Object 1: The yellow frisbee.
[0064] Second level – Background:
[0065] Long shot 1: White clouds and a clear blue sky.
[0066] View 2: Rolling mountains and green vegetation.
[0067] Medium shot 1: Green grassland.
[0068] First level – Action description layer:
[0069] Second level - Main elements
[0070] Subject 1: The figure bends over and extends his right hand, making a throwing motion with a frisbee.
[0071] Subject 2: The dog has its mouth open, its eyes fixed on the frisbee, ready to catch it.
[0072] Second level – Object elements
[0073] Object 1: A yellow frisbee moves through the air.
[0074] Second level – Background:
[0075] Long shot 1: White figures moving slowly in the sky.
[0076] View 2: Rolling mountains and green vegetation.
[0077] Medium shot 1: Grass sways in the wind on the grassland.
[0078] For example, see Figure 3 In this target page, the first area is located on the right side of the page, and the first image is displayed within this first area. The second area is located on the left side of the page, and the hierarchical structure for generating the prompt information is displayed within this second area. Figure 3 In this design, the hierarchical structure of the generated prompts uses an expandable and collapsible interactive design. When a user selects an element, the corresponding generated prompt will be highlighted. If the generated prompt for that element was originally in a collapsed state, selecting the element will switch the display state of the generated prompt to an expanded state.
[0079] For example, see [link to previous article] Figure 3 Suppose the user selects the puppy in the first image (e.g., clicks on the area occupied by the puppy). In the first area, the puppy's edges are in a preset state, indicating that the puppy is currently the target element. In the second area, there are two generated prompts corresponding to the puppy: one under the "Image Layer - Subject Layer - Subject 2" entry, specifically "A white dog wearing a black collar," and the other under the "Action Layer - Subject Layer - Subject 2" entry, specifically "The dog has its mouth open, its eyes fixed on the frisbee, ready to catch it." Both of these generated prompts are displayed in an expanded state and are in a preset state.
[0080] The above technical solution displays the correspondence between generated prompts and sub-images on the target page. This approach helps users understand the scope of the generated prompts' influence on the first image and, by using the generated prompts as examples, demonstrates how to express them, thus assisting users in understanding how to construct accurate and effective generated prompts.
[0081] Based on the above technical solution, optionally, generating a target image based on the editing result of the generated prompt information and the correspondence between the generated prompt information and the sub-image may further include: responding to the modification operation of the target generated prompt information in multiple generated prompt information, generating a new sub-image corresponding to the target generated prompt information based on the modified target generated prompt information; and replacing the sub-image corresponding to the target generated prompt information in the first image with the new sub-image to obtain the target image.
[0082] It is important to emphasize that some sub-images in the target image have changed compared to the first image. These changes are because the original sub-images corresponding to the target information generated before the modification were replaced with newly generated images.
[0083] Furthermore, the modification operations for the target-generated prompt information may include: modifying the text content of the target-generated prompt information, and / or determining a reference image associated with the target-generated prompt information.
[0084] For cases where the text content of the target-generated prompt information is modified, for example, in Figure 3 Based on this, if the user modifies the generation prompt information under the "Image Layer - Subject Layer - Subject 2" item in the second area—specifically, changing "white dog wearing a black collar" to "the dog's fur is brown and white, wearing a black collar and harness"—and then clicks the "Generate" option, see [link to documentation]. Figure 4 In the first area, the target image is displayed. Figure 3 Compared to the first image in the middle, Figure 4 The puppy in the first target image is different. All other elements are the same.
[0085] If the modification operation of the target-generated prompt information includes determining a reference image associated with the target-generated prompt information; optionally, based on the modified target-generated prompt information, generating a new sub-image corresponding to the target-generated prompt information includes: generating a new sub-image corresponding to the target-generated prompt information based on the target-generated prompt information and the reference image associated with the target-generated prompt information.
[0086] The reference image associated with the target generation prompt can be, for example, an image that needs to be input into the image generation model along with the target generation prompt. In this scenario, the relationship between the target generation prompt and the reference image can be understood as a binding relationship. This relationship ensures that when performing image generation, the target generation prompt and the reference image are input into the image generation model as a whole, and they interact with each other, jointly affecting the final generated image.
[0087] In practice, there are various methods for determining the reference image associated with the target-generated prompt information, and this application does not limit this method. For example, in some embodiments, a reference image import option associated with the target-generated prompt information is displayed on the target page. If the user triggers this reference image import option (e.g., clicks or drags the option), an image upload page is displayed. The image upload page receives an image specified by the user and stored in local storage space, uploads the specified image, and uses the uploaded image as the reference image associated with the target-generated prompt information. For example, see [link to example]. Figure 3 The second area, under the "Image Layer - Subject Layer - Subject 2" item, includes a "Reference" option in the generated prompt information display box. This "Reference" option is the reference image import option. If the user clicks the "Reference" option, this generated prompt information will be used as the target generated prompt information, and an image upload page will be displayed. Users can upload images through the image upload page. The image uploaded by the user through the image upload page will be the reference image associated with the target generated prompt information. Subsequently, if the user clicks... Figure 3 The "Generate" option will generate a new sub-image based on the text information displayed in the display box of the generation prompt information under the "Image Layer - Subject Layer - Subject 2" item in the second area and the uploaded image. The generated new sub-image will then replace the sub-image in the first image that corresponds to the generation prompt information, thus obtaining the target image.
[0088] In other embodiments, a reference image selection option associated with the target generated prompt information is displayed on the target page. If the user triggers an operation on this reference image selection option, an image selection page is displayed. The image selection page includes multiple images, which are images stored on a web server, not in local storage. In response to an image selection operation on the image selection page, the selected image is used as the reference image associated with the target generated prompt information. For example, see [link to example]. Figure 3The second area, under the "Image Layer - Subject Layer - Subject 2" item, includes a "Tips" option in the generated prompt information display box. This "Tips" option is the reference image option. If the user clicks the "Tips" option, the generated prompt information will be used as the target prompt information, and see [link / reference]. Figure 5 The image selection page is displayed. Users can select images through this page. The images on the image selection page can be, for example, images stored on a web server. The image selected by the user on the image selection page is used as a reference image associated with the target generated prompt information. If the user subsequently clicks... Figure 3 The "Generate" option will generate a new sub-image based on the text information displayed in the display box of the generation prompt information under the "Image Layer - Subject Layer - Subject 2" item in the second area and the image selected in the image selection page. The generated new sub-image will then replace the sub-image in the first image that corresponds to the generation prompt information, thus obtaining the second target image.
[0089] By configuring responses to modifications to the target generated prompt among multiple generated prompts, a new sub-image corresponding to the target generated prompt is generated based on the modified target generated prompt. This new sub-image then replaces the sub-image in the first image corresponding to the target generated prompt, resulting in the target image. Essentially, this approach, while displaying the correspondence between generated prompts and sub-images on the target page, allows users to modify and adjust the first image by altering the generated prompts. This benefits users by providing a deeper understanding of how to control the image generation model to generate specific images, and also improves the efficiency of modifying the first image to meet the user's needs.
[0090] It should also be noted that, in practice, the sub-image generated based on the modified target-generated prompt information may differ from the sub-image corresponding to the original target-generated prompt information in terms of the area it occupies and its outline. This is especially true when the generated prompt information describing the action of an element is modified. For example, a sub-image may show a cat curled up, occupying a relatively small and regular circular outline area. The corresponding generated prompt information describing the action of this sub-image is "resting cat." If this generated prompt information is modified to "jumping cat," the cat's posture in the newly generated sub-image will completely change. To show the cat jumping, it may stretch its limbs, the area it occupies may increase, and the outline will no longer be a regular circle but an irregular shape based on the cat's jumping posture.
[0091] In practice, when it is necessary to regenerate a sub-image of the first image, one can modify the generation prompt information corresponding to the sub-image and set a reference image associated with the generation prompt information; or only modify the generation prompt information corresponding to the sub-image; or only set a reference image associated with the generation prompt information.
[0092] Based on the above technical solution, the method may optionally include: in response to the operation of supplementing the generated prompt information, determining the newly supplemented generated prompt information and the area of effect of the newly supplemented generated prompt information in the first image; based on the newly supplemented generated prompt information, regenerating the image in the area of effect of the newly supplemented generated prompt information in the first image to obtain a second image; and replacing the image in the area of effect of the newly supplemented generated prompt information in the first image with the second image to obtain a target image.
[0093] For example, see Figure 6 The "Add Layer" section in the second area assists users in supplementing the generated prompt information. This "Add Layer" includes a region selection tool and a prompt information input box. Users can enter text and / or set reference images in the prompt information input box; the text entered and / or the reference images set will be used as the newly added prompt information. Users can use the region selection tool to specify the area where the newly added prompt information will apply. For example, Figure 6 In the example, the user sets the newly generated supplementary prompt message to "a kitten," and the scope of this prompt message is area 1. Then the user clicks... Figure 6 The "Generate" option will regenerate the image in region 1 based on "a kitten". The final target image is... Figure 7 The first region is given. (Comparison) Figure 6 and Figure 7 Compared to Figure 6 The first image in the series, Figure 7 The target image shows a white kitten in region 1.
[0094] By setting a newly added generation prompt, the image in the area of effect of the newly added generation prompt in the first image is regenerated to obtain the second image; the second image is then used to replace the image in the area of effect of the newly added generation prompt in the first image to obtain the target image. This can help users modify the first image by adding new element types and meet their diverse image modification needs.
[0095] Based on the above technical solution, optionally, the target image can be used as the updated first image, and the step of segmenting the first image to obtain multiple sub-images can be performed again.
[0096] The purpose of this setup is to use the target image as the updated first image and then execute the technical solution provided in this application again, thereby updating the correspondence between the generated prompt information and the sub-images. For example, see [link to example]. Figure 7 The target image is used as the updated first image, and the technical solution provided in this application is executed again. The updated generated prompt information includes generated prompt information corresponding to the kitten. Specifically, in Figure 7 In the image, the generation prompts for describing the kitten's appearance are displayed under the "Image Layer - Subject Layer - Subject 3" item in the second area; the generation prompts for describing the kitten's actions are displayed under the "Action Layer - Subject Layer - Subject 3" item in the second area. This setup facilitates the continuous modification of the first image through multiple rounds of editing.
[0097] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0098] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0099] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0100] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0101] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0102] Figure 8 This is a schematic diagram of the structure of an interactive device according to an embodiment of this disclosure. The interactive device provided in this embodiment can be configured in a client or in a server. See also Figure 8 The interactive device specifically includes:
[0103] The segmentation module 310 is used to segment the first image to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements;
[0104] The text determination module 320 is used to obtain first text describing the image content of the first image based on the first image;
[0105] The relationship determination module 330 is used to decompose the first text used to describe the image content of the first image based on the sub-image to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; the generated prompt information is used to describe the image content of the sub-image corresponding to it;
[0106] The generation module 340 is used to generate a target image or target video based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image.
[0107] Furthermore, the device also includes a relationship display module, used to display the correspondence between the generated prompt information and the sub-image on the target page.
[0108] Furthermore, the target page includes a first area and a second area; the relationship display module is used for:
[0109] The first image is displayed in the first area of the target page;
[0110] The generation prompt information is displayed in the second area of the target page;
[0111] In response to the selection operation of the target element in the first image, the display state of the generated prompt information corresponding to the target element in the second region is adjusted to a preset state.
[0112] Furthermore, the relationship display module is used for:
[0113] Based on a preset hierarchical structure division rule, the generated prompt information is sorted out to obtain the hierarchical structure of the generated prompt information; the preset hierarchical structure division rule includes a first level and a second level, the first level includes an image description layer and an action description layer; the second level of the image description layer and / or the action description layer includes at least one of a subject element, an object element, and a background.
[0114] The hierarchical structure of the generated prompt information is displayed in the second area of the target page.
[0115] Furthermore, the device also includes a modification module for:
[0116] In response to a modification operation on the target generated prompt information among the plurality of generated prompt information, a new sub-image corresponding to the target generated prompt information is generated based on the modified target generated prompt information;
[0117] The target image is obtained by replacing the sub-image in the first image that corresponds to the target-generated prompt information with the new sub-image.
[0118] Furthermore, the modification operation of the target-generated prompt information includes: modifying the text content of the target-generated prompt information, and / or determining a reference image associated with the target-generated prompt information;
[0119] If the modification operation of the target generation prompt information includes determining a reference image associated with the target generation prompt information; the modification module is used to: generate a new sub-image corresponding to the target generation prompt information based on the target generation prompt information and the reference image associated with the target generation prompt information.
[0120] Furthermore, the module is modified for:
[0121] In response to the operation of supplementing the generated prompt information, the newly supplemented generated prompt information and the area of effect of the newly supplemented generated prompt information in the first image are determined;
[0122] Based on the newly added generation prompt information, the image in the area where the newly added generation prompt information is applied in the first image is regenerated to obtain the second image;
[0123] The target image is obtained by replacing the image in the area of the newly added generation prompt information in the first image with the second image.
[0124] Furthermore, the device also includes an update module for:
[0125] The target image is used as the updated first image, and the step of segmenting the first image to obtain multiple sub-images is performed again.
[0126] The interactive device provided in this disclosure can execute the steps performed by the client or server in the interactive method provided in this disclosure, and has the functions of execution steps and beneficial effects, which will not be described in detail here.
[0127] Figure 9 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 9 The diagram illustrates a structural schematic suitable for implementing the electronic device 1000 in the embodiments of this disclosure. The electronic device 1000 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0128] like Figure 9 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003 to implement the interaction method as described in the embodiments of this disclosure. The RAM 1003 also stores various programs and information required for the operation of the electronic device 1000. The processing device 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0129] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to exchange information with other devices wirelessly or via wired communication. Although Figure 9 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the interactive methods as described above. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.
[0131] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include information signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated information signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0132] In some implementations, clients and servers may communicate using any known or future network protocol, such as HTTP (Hypertext Transfer Protocol), and may interconnect with any form or medium of digital information communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any known or future network.
[0133] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0134] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0135] The first image is segmented to obtain multiple sub-images; the first image includes multiple elements; different sub-images correspond to different elements;
[0136] Based on the first image, a first text is obtained to describe the image content of the first image;
[0137] Based on the sub-image, the first text used to describe the image content of the first image is decomposed to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; the generated prompt information is used to describe the image content of the sub-image corresponding to it;
[0138] Based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image, a target image or target video is generated.
[0139] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.
[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] According to one or more embodiments of this disclosure, this disclosure provides an electronic device, including:
[0146] One or more processors;
[0147] Memory, used to store one or more programs;
[0148] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the interaction methods provided in this disclosure.
[0149] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements an interaction method as described in any of the present disclosure.
[0150] This disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the interaction method described above.
[0151] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0152] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interaction method, characterized in that, include: The first image is segmented to obtain multiple sub-images; The first image includes multiple elements; different sub-images correspond to different elements; Based on the first image, a first text is obtained to describe the image content of the first image; Based on the sub-image, the first text used to describe the image content of the first image is decomposed to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; The generated prompt information is used to describe the image content of the corresponding sub-image; Based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image, a target image or target video is generated.
2. The method according to claim 1, characterized in that, Also includes: The target page displays the correspondence between the generated prompt information and the sub-image.
3. The method according to claim 2, characterized in that, The target page includes a first region and a second region; displaying the correspondence between the generated prompt information and the sub-image on the target page includes: The first image is displayed in the first area of the target page; The generation prompt information is displayed in the second area of the target page; In response to the selection operation of the target element in the first image, the display state of the generated prompt information corresponding to the target element in the second region is adjusted to a preset state.
4. The method according to claim 3, characterized in that, The provision of generating prompt information, displayed in the second area of the target page, includes: Based on a preset hierarchical structure division rule, the hierarchical structure of the generated prompt information is obtained; the preset hierarchical structure division rule includes a first level and a second level, the first level includes an image description layer and an action description layer; the second level of the image description layer and / or the action description layer includes at least one of a subject element, an object element, and a background. The hierarchical structure of the generated prompt information is displayed in the second area of the target page.
5. The method according to claim 1, characterized in that, The step of generating a target image based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image further includes: In response to a modification operation on the target generated prompt information among the plurality of generated prompt information, a new sub-image corresponding to the target generated prompt information is generated based on the modified target generated prompt information; The target image is obtained by replacing the sub-image in the first image that corresponds to the target-generated prompt information with the new sub-image.
6. The method according to claim 5, characterized in that, The modification operation of the target-generated prompt information includes: modifying the text content of the target-generated prompt information, and / or determining a reference image associated with the target-generated prompt information; If the modification operation of the target generation prompt information includes determining a reference image associated with the target generation prompt information; the step of generating a new sub-image corresponding to the target generation prompt information based on the modified target generation prompt information includes: generating a new sub-image corresponding to the target generation prompt information based on the target generation prompt information and the reference image associated with the target generation prompt information.
7. The method according to claim 1, characterized in that, Also includes: In response to the operation of supplementing the generated prompt information, the newly supplemented generated prompt information and the area of effect of the newly supplemented generated prompt information in the first image are determined; Based on the newly added generation prompt information, the image in the area where the newly added generation prompt information is applied in the first image is regenerated to obtain the second image; The target image is obtained by replacing the image in the area of the newly added generation prompt information in the first image with the second image.
8. The method according to any one of claims 5-7, characterized in that, Also includes: The target image is used as the updated first image, and the step of segmenting the first image to obtain multiple sub-images is performed again.
9. An interactive device, characterized in that, include: The segmentation module is used to segment the first image into multiple sub-images; The first image includes multiple elements; different sub-images correspond to different elements; The text determination module is used to obtain first text describing the image content of the first image based on the first image; The relationship determination module is used to decompose the first text used to describe the image content of the first image based on the sub-image, to obtain multiple generated prompt information and the correspondence between the generated prompt information and the sub-image; The generated prompt information is used to describe the image content of the corresponding sub-image; The generation module is used to generate a target image or target video based on the editing results of the generated prompt information and the correspondence between the generated prompt information and the sub-image.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.