Page interaction method and device, electronic equipment, storage medium and program product
By receiving trajectory information and text prompts in the interactive area of the page, and generating images using a multimodal image generation model, the problem of inaccurate control over image generation details in existing technologies is solved, thereby improving image generation quality and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing text-to-image technology struggles to precisely control the details of the generated images, resulting in low-quality images.
By receiving and displaying trajectory information and text prompts in the interactive area of the target page, a multimodal image generation model is used to generate an image that matches the trajectory information and text prompts.
It enables precise control over the generated images, improves the quality of image generation, and optimizes the user experience.
Smart Images

Figure CN121879649A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a page interaction method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] With the rapid development of AI-generated content (AIGC), text-to-image (TTO) technology has made significant breakthroughs, enabling users to generate images based on text descriptions. However, relying solely on text prompts makes it difficult to precisely control details such as the position, shape, and color distribution of objects in the generated image. While current image-to-image technologies can generate new images by locally modifying the original image or adding new objects, they still struggle to accurately control the details of the resulting image. In short, both text-to-image and image-to-image methods currently struggle to accurately control the details of the generated image, resulting in low-quality images. Summary of the Invention
[0003] This disclosure provides a page interaction method, apparatus, electronic device, storage medium, and program product to solve at least one of the aforementioned technical problems. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a page interaction method is provided, comprising: An interactive content area is displayed on the target page, which is used to receive and display trajectory information; In the content interaction area, text prompts are received and displayed. These text prompts are used to indicate the generation requirements of the image generation operation, and the generation requirements are semantically related to the trajectory information. When the target page receives an image generation instruction, a first image is displayed. The first image is the result of the image generation operation and its semantics are consistent with those indicated by the trajectory information.
[0004] In one exemplary embodiment, the trajectory information indicates at least one of the following semantics: the shape semantics of the target object, the visual element semantics of the target object, the direction of change of the visual information of the target object, the direction of change of the visual effect influencing factor, and the influence area of the visual effect influencing factor; The target object is an object within the spatial region indicated by the trajectory information.
[0005] In one exemplary embodiment, the target object is characterized by at least one of the following: the outline formed by the text prompt information and the trajectory information.
[0006] In one exemplary embodiment, the method further includes: A second image is displayed in the content interaction area; The trajectory information is displayed on the second image; The first image is the result of performing the image generation operation based on the second image.
[0007] In one exemplary embodiment, the first image is the result of performing the image generation operation on a target region in the second image, where the target region is the spatial region indicated by the trajectory information.
[0008] In one exemplary embodiment, the trajectory information characterizes the outline of the target object, and the text prompt information is used to describe at least one of the following: the target object, the style of the target object, and whether the target object belongs to the background; Alternatively, the trajectory information may have a direction, and the text prompt information may describe the semantics corresponding to the direction indicated by the trajectory information.
[0009] In one exemplary embodiment, the trajectory information is used to indicate the color of the target object by its own color, and the method further includes: Display color selection controls; When the color selection control is triggered, the selected color is used to display the trajectory information.
[0010] In one exemplary embodiment, the method further includes: Display freely drawable controls; When the free-drawing control is triggered, the drawn trajectory information is received and displayed in the content interaction area.
[0011] In one exemplary embodiment, the method further includes: Display shape selection controls; When the shape selection control is triggered, the selected shape is displayed in the content interaction area; In response to editing of the shape, the trajectory information is displayed based on the editing result.
[0012] According to a second aspect of the present disclosure, a page interaction device is provided, comprising: The interactive area display module is configured to display a content interactive area on the target page, wherein the content interactive area is used to receive and display trajectory information; An interaction module is configured to execute in the content interaction area, receiving and displaying text prompts, the text prompts indicating the generation requirements of the image generation operation, the generation requirements being semantically related to the trajectory information; and, When the target page receives an image generation instruction, a first image is displayed. The first image is the result of the image generation operation and its semantics are consistent with those indicated by the trajectory information.
[0013] In one exemplary embodiment, the trajectory information indicates at least one of the following semantics: the shape semantics of the target object, the visual element semantics of the target object, the direction of change of the visual information of the target object, the direction of change of the visual effect influencing factor, and the influence area of the visual effect influencing factor; The target object is an object within the spatial region indicated by the trajectory information.
[0014] In one exemplary embodiment, the target object is characterized by at least one of the following: the outline formed by the text prompt information and the trajectory information.
[0015] In one exemplary implementation, the interaction module is configured to execute: A second image is displayed in the content interaction area; The trajectory information is displayed on the second image; The first image is the result of performing the image generation operation based on the second image.
[0016] In one exemplary embodiment, the first image is the result of performing the image generation operation on a target region in the second image, where the target region is the spatial region indicated by the trajectory information.
[0017] In one exemplary embodiment, the trajectory information characterizes the outline of the target object, and the text prompt information is used to describe at least one of the following: the target object, the style of the target object, and whether the target object belongs to the background; Alternatively, the trajectory information may have a direction, and the text prompt information may describe the semantics corresponding to the direction indicated by the trajectory information.
[0018] In one exemplary implementation, the trajectory information is used to indicate the color of the target object by its own color, and the interaction module is configured to execute: Display color selection controls; When the color selection control is triggered, the selected color is used to display the trajectory information.
[0019] In one exemplary implementation, the interaction module is configured to execute: Display freely drawable controls; When the free-drawing control is triggered, the drawn trajectory information is received and displayed in the content interaction area.
[0020] In one exemplary implementation, the interaction module is configured to execute: Display shape selection controls; When the shape selection control is triggered, the selected shape is displayed in the content interaction area; In response to editing of the shape, the trajectory information is displayed based on the editing result.
[0021] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the page interaction method described above.
[0022] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the page interaction method as described above.
[0023] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the page interaction method described above.
[0024] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: The page interaction method, apparatus, electronic device, storage medium, and program product disclosed herein are applied to a target page, which is a page for generating and displaying an image. The target page can be a page in an application for image creation.
[0025] The target page displays an interactive content area that can receive and display trajectory information, as well as text prompts. These text prompts indicate the generation requirements for the image generation operation, and these requirements are semantically related to the trajectory information. This design ensures that the first image generated based on the text prompts conforms to the semantics indicated by the trajectory information. It fully utilizes the synergy between trajectory information and text prompts, resulting in a first image that meets both the requirements of the text prompts in the text modality and the semantics of the trajectory information in the image modality. This achieves the technical goal of precisely controlling the image generation quality based on multimodal information. Furthermore, the acquisition of both trajectory and text information is relatively simple. Users can input trajectory information through simple interactions, such as drawing, to achieve the goal of precisely controlling the quality of the generated image using trajectory information combined with text prompts, thereby improving image generation effects and optimizing the user experience.
[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0028] Some of the accompanying drawings in this disclosure are in color. Because this disclosure involves the representation of color, to ensure clarity of expression, while retaining the color drawings, corresponding grayscale drawings are also provided for review purposes.
[0029] Figure 1 This is a schematic diagram of an implementation environment according to an exemplary embodiment.
[0030] Figure 2 This is a flowchart illustrating a page interaction method according to an exemplary embodiment.
[0031] Figure 3 This is a schematic diagram of a content interaction area according to an exemplary embodiment.
[0032] Figure 4 This is a schematic diagram of a page interaction process according to an exemplary embodiment. Figure 1 .
[0033] Figure 5 This is a schematic diagram of a page interaction process according to an exemplary embodiment. Figure 2 .
[0034] Figure 6 This is a schematic diagram illustrating the interactive effect according to an exemplary embodiment.
[0035] Figure 7 This is a block diagram of a page interaction device according to an exemplary embodiment.
[0036] Figure 8 This is a block diagram illustrating an electronic device for page interaction according to an exemplary embodiment.
[0037] Figure 9 This is another block diagram illustrating an electronic device for page interaction according to an exemplary embodiment. Detailed Implementation
[0038] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0039] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0041] This disclosure is primarily applicable to image creation applications, which allow users to create images, such as generating new images or redrawing or editing existing images.
[0042] However, these applications in related technologies can only generate or redraw images using text-to-image or image-to-image methods alone, and cannot combine multimodal information to guide image generation. Therefore, the control over the generated image is not precise. In redrawing scenarios, these applications allow users to define their own regions and redraw within those regions. However, these defined regions only serve as masks to limit the redrawing area, providing only spatial information related to the redrawing area. They cannot provide semantic information such as color and shape for the redrawing, which leads to low-quality redrawing results.
[0043] In related technologies, users find it difficult to precisely control the quality of generated images through simple interactions, resulting in inaccurate semantic control over the geometry and color of objects in the generated images. In view of this, this disclosure provides a simple interaction method that allows users to control a multimodal image generation model to generate high-quality images that meet expectations through simple "trajectory information (doodles) + text prompts (typing)". This multimodal image generation model fully utilizes multimodal information from both the image modality (trajectory information) and the text modality (text prompts), improving the precision of semantic control over the generated image and the objects within it, thereby enabling the generation of high-quality images.
[0044] Please see Figure 1 The illustration shows an implementation environment provided by an embodiment of the present disclosure. The implementation environment may include at least one page interaction terminal 110 and an information acquisition server 120, wherein the page interaction terminal 110 and the information acquisition server 120 can communicate with each other via a network.
[0045] Specifically, the page interaction terminal 110 interacts with the user through interaction with the information acquisition server 120. Specifically, the page interaction terminal 110 can display a content interaction area on the target page, which is used to receive and display trajectory information; in the content interaction area, it receives and displays text prompts, which are used to indicate the generation requirements of an image generation operation, and the generation requirements are semantically related to the trajectory information; when the target page receives an image generation instruction, it displays a first image, which is the execution result of the image generation operation, and the first image is semantically consistent with the trajectory information.
[0046] The page interaction terminal 110 can communicate with the information acquisition server 120 based on a browser / server (B / S) mode or a client / server (C / S) mode. The page interaction terminal 110 may include physical devices such as smartphones, tablets, laptops, digital assistants, smart wearable devices, in-vehicle terminals, and servers, and may also include software running on the physical device, such as applications. The operating system running on the page interaction terminal 110 in this embodiment may include, but is not limited to, Android, iOS, Linux, and Windows.
[0047] The information acquisition server 120 and the page interaction terminal 110 can establish and display a communication connection through wired or wireless means. The information acquisition server 120 may include a stand-alone server, a distributed server, or a server cluster composed of multiple servers, wherein the server may be a cloud server.
[0048] Please refer to Figure 2 The diagram illustrates a page interaction method flowchart in an exemplary embodiment of this disclosure. The execution subject of this method can be the aforementioned page interaction terminal. Please refer to [link / reference] for details. Figure 2 The method may include: S210. Display a content interaction area on the target page, wherein the content interaction area is used to receive and display trajectory information.
[0049] In this disclosure, the target page can be a page provided in an application for image creation, which displays an interactive content area that serves as an entry point for users to input, edit, and view images.
[0050] This disclosure does not limit the trajectory information, which can present any geometric shape, such as a straight line, curve, polyline, or other complex shape. For example, it can be closed or open, and can include points, lines, surfaces, or any combination thereof. This disclosure does not limit the input method of trajectory information. For example, users can freely draw in the content interaction area using their fingers, a mouse, or a stylus to create personalized trajectory information through drawing. Alternatively, users can generate trajectory information through voice commands, gesture recognition, or other input devices and import this trajectory information into the content interaction area. In some embodiments, the content interaction area can also provide several trajectory templates, allowing users to select a suitable template for quick drawing, thereby improving interaction efficiency. Users can also edit the results drawn based on the templates to obtain personalized trajectory information. These templates can cover common geometric shapes or specific pattern styles to meet the needs of different scenarios. Furthermore, the content interaction area also supports editing and adjusting the generated trajectory information, such as modifying visual elements like color, thickness, and transparency, or performing operations such as cropping, scaling, and rotating the trajectory to achieve more accurate trajectory expression.
[0051] In one exemplary implementation, please refer to Figure 3 This diagram illustrates a content interaction area in an exemplary embodiment of the present disclosure. Figure 3 (a) The content interaction area 310 is an area capable of real-time sensing of multimodal input, integrating multiple modal inputs such as text and images into a single interaction area. The content interaction area 310 can display a drawing editing control 320. Clicking the drawing editing control 320 allows drawing trajectory information, which can then be displayed. Figure 3 (b) in the content interaction area 310.
[0052] S220. In the content interaction area, receive and display text prompt information, the text prompt information being used to indicate the generation requirements of the image generation operation, the generation requirements being semantically related to the trajectory information.
[0053] The text prompts in this disclosure are used to indicate the generation requirements of the image generation operation. These requirements can be used to define the image's theme, style, color scheme, or other specific conditions related to image generation. These requirements help users express their needs more precisely, thereby improving the quality and efficiency of image generation. For example, the text prompts may include descriptive text about the image, such as "modern minimalist interior design" or "abstract pattern containing blue and green gradients." Furthermore, the generation requirements can be further refined, such as specifying the image's resolution, aspect ratio, or whether specific visual elements are required.
[0054] The text prompts in this disclosure indicate generation requirements that are semantically related to trajectory information. They serve as a bridge to guide the multimodal image generation model in generating corresponding images based on trajectory information. These text prompts can not only be understood by the multimodal image generation model, but also guide it to understand the semantics of the trajectory information drawn by the user. This enables the multimodal image generation model to achieve true "image + text" bidirectional understanding, allowing it to not only "understand" spoken words, but also "understand" the user's sketch intent, thereby improving the quality of the generated images.
[0055] For example, if the trajectory information describes the outlines of three blue penguins on the canvas of the content interaction area, the text prompts can further clarify the style of these penguins, such as "Three blue penguins, cartoon style, with a snowy world background." In this way, the multimodal image generation model can combine trajectory information and text prompts to generate images that better meet user needs. Furthermore, this two-way understanding mechanism can be extended to more complex scenarios, such as architectural design and product prototyping, helping users quickly transform abstract ideas into concrete visual representations. At the same time, this method also supports multiple iterative optimizations; users can gradually improve the generated results by adjusting the trajectory information or modifying the text prompts, thereby achieving a more efficient creation process.
[0056] This disclosure does not limit the input method for text prompts; users can provide text prompts through various means, such as manual input, speech-to-text conversion, or selection from preset templates. This flexibility ensures adaptability to different use cases, meeting needs from rapid sketching to detailed design.
[0057] In one exemplary implementation, please refer to Figure 3 (b) In Figure 3(b) The content interaction area 310 can be used to input text prompts. Figure 3 (b) shows the outlines of three blue penguins, and the text prompt is “Three blue penguins, in a cartoon style, against a snowy world background.”
[0058] S230. Upon receiving an image generation instruction on the target page, a first image is displayed, wherein the first image is the result of the image generation operation and the semantics of the first image are consistent with those indicated by the trajectory information.
[0059] When a user inputs an image generation command on the target page, a multimodal image generation model can be triggered to generate a first image based on the text prompts and trajectory information, and then display the first image on the target page. In an exemplary implementation, please refer to... Figure 3 ,exist Figure 3 (b) also shows an image content generation control 330, which can be triggered to display the generated first image. Figure 3 (c) A color illustration of the first image is shown. Figure 3 (d) Show a grayscale diagram of the first image.
[0060] The page interaction method disclosed herein displays a content interaction area on the target page. This area can receive and display trajectory information, and also receive and display text prompts. The text prompts indicate the generation requirements of the image generation operation, and these requirements are semantically related to the trajectory information. This design ensures that the first image generated based on the text prompts conforms to the semantics indicated by the trajectory information, fully utilizing the synergy between the trajectory information and the text prompts. This results in the generated first image meeting both the requirements of the text prompts in the text modality and the semantics of the trajectory information in the image modality, achieving the technical objective of precisely controlling the image generation quality based on multimodal information. Furthermore, the acquisition methods for both trajectory and text information are relatively simple. Users can input trajectory information through simple interactions, such as drawing, to achieve the goal of precisely controlling the quality of the generated image using the combination of trajectory information and text prompts, thereby improving the image generation effect and optimizing the user experience.
[0061] This disclosure utilizes a multimodal image generation model to perform image generation operations. This model fully understands the semantics within trajectory information and creates images according to the generation requirements indicated by text prompts. This disclosure does not limit the training methods or model structures of the multimodal image generation model; related technologies can be used without posing any obstacle to implementation. For example, a pre-trained model combining visual and language understanding capabilities can be employed. Such models are trained on a large amount of cross-modal data, effectively capturing semantic features in trajectory information and aligning them with text prompts. These features can then be fine-tuned according to the specific needs of this disclosure. Alternatively, diffusion models or generative adversarial networks can be used as the skeleton of the multimodal image generation model. These models perform well in image generation tasks and can achieve the fusion of multimodal information by adjusting input conditions. Another option is a unified model based on the Transformer architecture, which can simultaneously process multiple types of data input and strengthen the correlation between different modalities through an attention mechanism, thereby generating high-quality images that meet the requirements.
[0062] In some implementations, a large-scale cross-modal dataset can be used to train the model on top of a pre-selected model. For example, a dataset containing doodle trajectories and corresponding descriptive text can be used to train the model, enabling it to learn the association between trajectory and semantics. Another approach is to design a two-stream network structure, with one stream processing trajectory information and the other processing textual prompts, finally combining the information from both modalities through fusion. Furthermore, attention mechanisms can be introduced, allowing the model to dynamically focus on key parts of the trajectory information and combine them with textual prompts to generate more suitable images. These methods can effectively improve the quality and accuracy of multimodal image generation. In some implementations, the multimodal image generation model used can jointly analyze trajectory information and textual prompts using deep learning techniques to generate high-quality images that meet user needs.
[0063] In one exemplary embodiment, the trajectory information indicates at least one of the following semantics: the shape semantics of the target object, the visual element semantics of the target object, the direction of change of the visual information of the target object, the direction of change of the visual effect influencing factor, and the influence area of the visual effect influencing factor; wherein, the target object is an object within the spatial area indicated by the trajectory information.
[0064] The target object refers to the object within the spatial region defined by the trajectory information. The target object can be located within this spatial region. If the trajectory information defines a penguin on an existing image, then the penguin is the target object. If the trajectory information defines a rectangle on a blank canvas, and a small pavilion is drawn inside that rectangle, then the small pavilion is the target object.
[0065] The target object is characterized by at least one of the following: the outline formed by the text prompt information and the trajectory information. For example, if the outline of the trajectory information forms a pair of large hands, then the pair of large hands is the target object. The target object can also be defined by text prompt information; for example, the text prompt information could be "Draw a duck in the circled area," then the "duck" is the target object. In some implementations, the target object can not only be a concrete object but also an abstract concept, such as a color block or a lighting effect. For example, if the trajectory information defines an area and an arrow is drawn in that area, and the text prompt information describes it as "Add a warm ray of sunlight to the circled area and render the effect of sunlight shining in the direction of the arrow," then this ray of sunlight is the target object. In this case, the target object is no longer limited to a concrete object shape but is represented as an abstract visual effect. Users can freely express the target object they want to draw through trajectory information, text prompt information, or a combination of both, thereby meeting diverse creative needs. This flexibility not only enhances the user experience but also opens up broader possibilities for the application of multimodal image generation technology.
[0066] Shape semantics describes the geometric features of a target object, such as its outline and geometric shape, like a circle, rectangle, or polygon. For example, if the target object is a cat, the trajectory information can form the outline of the cat, and that outline is the shape semantics of the cat.
[0067] Visual element semantics further refines the visual characteristics of the target object, including color, transparency, etc. If the trajectory information is the outline of a black and white cat, then the visual element semantics of the cat can include: a black and white cat. If the trajectory information represents a princess wearing a veil, with the princess's lines being opaque black lines and the veil being made of semi-transparent red lines, then the princess's visual element semantics can include: wearing opaque black clothes and a semi-transparent red veil.
[0068] The direction of visual information changes indicates the dynamic trend of a target object in terms of visual appearance. For example, a gradient effect from light to dark, or a transition from static to dynamic. For instance, the trajectory information could be an arrow pointing from the upper right to the lower left, with text indicating a gradual darkening in the direction of the arrow. Or, the trajectory information could be an arrow pointing from left to right, with text indicating that a child in the image is running in the direction of the arrow.
[0069] This disclosure does not limit the factors influencing visual effects; for example, they may include light, contrast, saturation, etc. These factors can significantly affect visual effects, thereby changing the appearance of the target object in the generated image. The trajectory information can indicate the direction of change of these visual effect influencing factors through directional arrows. For example, if the trajectory information is an upward arrow, and the text prompt indicates that the arrow represents the direction of brightness change in the image, then the generated image should present a brightness gradient effect from bottom to top. Similarly, if the trajectory information is a leftward arrow, and the text prompt indicates that the arrow represents the direction of saturation change in the image, then the generated image should present a saturation gradient effect from right to left. Changes in visual effect influencing factors are not limited to a single dimension; they can also be combinations of multiple factors. For example, the trajectory information may include multiple arrows, each indicating the direction of change in brightness and contrast, thereby achieving more complex and subtle visual effects in the generated image. This multi-dimensional variation can better meet users' needs for image adjustments, enhancing the flexibility and creativity of page interaction.
[0070] This disclosure can also limit the area of influence of visual effect influencing factors through trajectory information. For example, the trajectory information can include a closed curve that defines the area in the image that needs adjustment. If a text prompt indicates that the area within the closed curve needs increased brightness, then only that area in the generated image will show an increased brightness, while the area outside the curve remains unchanged. This area-limiting method allows users to more precisely control the local effects of the image, thereby achieving more refined editing needs. Furthermore, the shape and size of the affected area can also be dynamically adjusted based on the trajectory information. This approach not only enhances the intuitiveness of page interaction but also provides users with greater operational freedom, enabling them to complete complex image processing tasks more efficiently.
[0071] The difference between trajectory information and masks in related technologies lies in that trajectory information not only limits the area for image generation operations (e.g., limiting image generation to a defined area), but also guides the image generation process through rich semantics expressed by its visual information, such as color and shape. This semantics, combined with text prompts, can indicate numerous details, including the shape semantics of the target object, the semantics of its visual elements, the direction of change of its visual information, the direction of change of visual effect influencing factors, and the area of influence of these factors. This allows for precise control over the details of the generated image. Furthermore, the trajectory information drawing process is very simple, enabling users to participate in image generation in a more intuitive and flexible way. Users do not need professional image processing knowledge; they can achieve fine-tuning of image effects simply by drawing a trajectory. This approach not only lowers the operational threshold but also greatly enhances the user experience. In addition, trajectory information has a wide range of applications, meeting different levels of needs in both artistic creation and technical image processing. The combination of trajectory information and text prompts further expands its functional boundaries, making complex image generation tasks more efficient and creative.
[0072] In one exemplary embodiment, the trajectory information represents the outline of the target object, and the text prompt information is used to describe at least one of the following: the target object, the style of the target object, and whether the target object belongs to the background; or, the trajectory information has a direction, and the text prompt information describes the semantics corresponding to the direction indicated by the trajectory information.
[0073] In practical applications, this method of combining trajectory information and text prompts significantly improves the flexibility and accuracy of creative work. For example, users can describe the outline of a target object using trajectory information, and then use text prompts to further clarify the object's specific attributes. If the outline represents "a red apple," the text prompts can further specify the apple's details, style, and whether it belongs to the foreground or background. For instance, users can quickly generate an image that meets their needs with a simple description such as "a realistic red apple located in the foreground." This method not only saves time but also avoids tedious steps, making the creative process more intuitive and efficient.
[0074] For example, users can intuitively express directional information through arrows representing trajectory information, while text prompts supplement the semantic meaning of that direction, such as whether it's the direction of lighting, movement, or saturation change. For instance, a user can draw an upward arrow and input "light shining from below, creating a dramatic lighting effect," thus generating the desired visual effect based on the arrow's direction and the text prompt. This approach not only accurately conveys the user's creative intent but also presents complex design requirements in a more intuitive way. By combining directional information with semantic descriptions, users can easily achieve fine-grained control over image details, making it easier to adjust the dynamic trends of elements or optimize the overall layout. This flexible and intuitive operation greatly improves creative efficiency while providing users with more freedom to express themselves.
[0075] For example, users can describe the type of change a subject undergoes (such as "from young to old") using text and indicate the direction or extent of the change with arrows, allowing for dynamic control over the target object's state. For instance, a user can draw a downward arrow on a person's face and input "gradually aging," which will generate an image showing the aging process. This approach greatly facilitates the design of dynamic effects.
[0076] For example, users can quickly sketch the outline of a background by drawing, and then use text prompts to indicate details such as style and atmosphere. For instance, a user can draw an area with random lines and enter "dreamy starry sky style," and a corresponding background image will be generated based on this information. This method saves time while ensuring harmony between the background and foreground elements.
[0077] Overall, this multimodal interaction method deeply integrates visual and semantic information, significantly reducing operational complexity, enhancing creative flexibility, and improving the quality and diversity of generated results. Both professional designers and ordinary users can benefit from it, gaining a more efficient and intuitive creative experience.
[0078] In one exemplary embodiment, the method further includes: A second image is displayed in the content interaction area; The trajectory information is displayed on the second image; The first image is the result of performing the image generation operation based on the second image.
[0079] For example, the second image could be a basic portrait, with the trajectory information represented as an arc extending from the top of the head to the chin, accompanied by the text prompt "gradually increasing light and shadow contrast." Based on this input, the first image is generated as a portrait with a clear light and shadow gradient effect, where the contrast between light and shadow on the face is more prominent, enhancing the sense of three-dimensionality. This method, by combining intuitive trajectory drawing with semantic description, allows users to achieve high-quality image editing without complex operations. It lowers the professional skill threshold, improves creative efficiency, and ensures the artistic and personalized expression of the generated results.
[0080] For another example, the second image could be a simple landscape painting, featuring a calm lake and distant mountains. The trajectory information is represented by multiple ripple lines radiating outwards from the center of the lake, accompanied by the text prompt "Add dynamic water ripple effect." Based on this input, the first image is generated as a vivid landscape painting, with the lake surface exhibiting a clear ripple effect, and sunlight shining on the ripples creating shimmering spots, making the entire image more dynamic and vibrant. This method, by combining intuitive trajectory drawing with semantic description, allows users to easily achieve complex effect adjustments. It simplifies the operation process, enhances the user's creative experience, and ensures the realism and visual impact of the generated results.
[0081] In one exemplary embodiment, the first image is the result of performing the image generation operation on a target area in the second image, where the target area is the spatial region indicated by the trajectory information. This embodiment implements a function to partially redraw the second image. For example, the second image can be an interior design drawing. A blank wall can be circled, and an arc extending outward from the center of the wall can be drawn. Adding the text prompt "Add decorative painting and warm lighting effects" generates the first image based on these inputs. The wall, as the target area, is redrawn, resulting in a beautiful decorative painting appearing on the originally blank wall, surrounded by soft warm lighting, making the entire space more inviting and layered. This method not only allows for precise positioning and editing of the target area but also significantly reduces the learning curve for users accustomed to complex design tools. Simultaneously, through intuitive operation, users can quickly verify creative ideas, thereby improving design efficiency and achieving higher-quality visual expression.
[0082] Please refer to Figure 4 It illustrates an interactive process of an exemplary embodiment of this disclosure. Figure 1 .exist Figure 4In (a), a second image can be received in the content interaction area 410. The user inputs the second image, i.e., image 1, through interaction with the content interaction area 410. This second image shows a table with a chef standing behind it, whose hands are placed on the table. The user can click on this second image to... Figure 4 (b) Display a doodle editing control with the function description "Draw Trajectory". By triggering the doodle editing control, the user can draw the outline of a cake between the chef's hands on the table in the second image, and draw an arrow from the upper right to the lower left of the second image. The outline of the cake and the arrow are both part of the trajectory information entered by the user. The user can enter the text prompt "Help me light the cake according to the color and direction of the arrow, and then add a cake on the table according to my drawing @Image 1" in the content interaction area 410. After the user clicks the image generation control 420, the first image can be generated and displayed. Figure 4 (c) Display the first image.
[0083] In one exemplary embodiment, when using the page interaction method provided in this disclosure, a user can doodle on a second image to form a doodle layer and input text commands (e.g., "a furry monster"). When the target page receives the image generation command, it can perform the following operations using a multimodal image generation model: Image encoding: The second image is encoded into latent spatial features using the VAE (Variational Autoencoder) inside the multimodal image generation model.
[0084] Graffiti encoding: Extracts a binary mask of the area indicated by the trajectory information, which is used to indicate the redrawn area.
[0085] Conditional extraction: The lightweight convolutional network inside the multimodal image generation model is used to process the shape and color features indicated by the trajectory information and generate control vectors.
[0086] Text encoding: An encoder using a multimodal image generation model converts text instructions into text embedding vectors.
[0087] Image generation is performed using a diffusion model within a multimodal image generation model. During denoising, this latent spatial feature, control vector, and text embedding vector are used to guide denoising, yielding the latent features corresponding to the first image. In the denoising process, the model is constrained to match the edge contours and color characteristics of the graffiti. For example, if the graffiti is a blue circle, the model will tend to represent blue in the target object through texture or lighting and fill the circular area.
[0088] Based on the latent features generated by the model, the first image is obtained by decoding using its internal VAE. During the VAE decoding process, the redrawn areas can be fused with the non-redrawn areas of the second image at the pixel level to ensure that the edge transitions of the decoded first image are natural.
[0089] In this implementation, the multimodal image generation model directly obtains the generation result by simultaneously fusing the color, shape, and text prompt information of the graffiti in a single inference. This end-to-end, one-step generation method has a high real-time advantage.
[0090] In one exemplary embodiment, the trajectory information is used to indicate the color of the target object through its own color. The method further includes: displaying a color selection control; and, when the color selection control is triggered, displaying the trajectory information using the selected color. In this way, users can intuitively perceive the color attributes of the target object based on the color changes of the trajectory information, thereby improving the intuitiveness of the interaction and operational efficiency. Simultaneously, displaying the color selection control provides users with a flexible selection mechanism, making color adjustment more convenient and accurate.
[0091] Please refer to Figure 5 It illustrates an interactive process of an exemplary embodiment of this disclosure. Figure 2 .exist Figure 5 In (a), a second image, namely image 1, can be received in the content interaction area 510. The user inputs the second image through interaction with the content interaction area 510. The second image shows a middle-aged man wearing glasses and a shirt. By clicking on the second image, the user can... Figure 5 (b) Display the graffiti editing control "Drawing Trajectory". Triggering the graffiti editing control displays a color selection control 520. Interacting with this control allows drawing red and blue arrows, both of which represent user-drawn trajectory information. The red arrow points from the lower left to the upper right, and the blue arrow points from the upper right to the lower left. Interacting with the color selection control also allows drawing a pair of blue wings, which also represent user-drawn trajectory information.
[0092] In one exemplary embodiment, the method further includes: displaying a free-drawing control; and, when the free-drawing control is triggered, receiving and displaying the drawn trajectory information in the content interaction area. In practical applications, users can activate the free-drawing function by clicking the free-drawing control on the target page. After selecting the free-drawing control, users can freely draw lines, graphics, or other trajectory information. This trajectory information can be a hand-drawn sketch, marked key areas, or a quick sketch of a creative graphic. In this way, users can express their design intentions more flexibly. Compared with the traditional fixed template drawing method, this method greatly improves the freedom and personalization of operation. In addition, since the trajectory information can be displayed in real time in the content interaction area, users can instantly view the drawing effect and make adjustments, thereby improving work efficiency and accuracy.
[0093] In one exemplary embodiment, the method further includes: displaying a shape selection control; displaying the selected shape in the content interaction area when the shape selection control is triggered; and displaying trajectory information based on the editing result in response to editing of the shape. For example, a shape selection control can be provided on the target page, allowing the user to select preset shapes such as rectangles, circles, or triangles after clicking the control. Once a shape is selected, it is immediately displayed in the content interaction area, and the user can edit the shape by dragging, scaling, rotating, or editing the control points of these shapes, and the editing result is used as trajectory information or part of the trajectory information. This method not only simplifies the drawing process of complex graphics but also allows users to flexibly adjust the style and position of shapes according to their needs. In this way, users do not need to manually draw precise geometric shapes, thus reducing the difficulty of operation, and is especially suitable for scenarios requiring rapid design completion. At the same time, since the editing results are provided instantly, users can more intuitively confirm the effect and make fine adjustments, which further improves design efficiency and accuracy.
[0094] Please refer to Figure 6 This diagram illustrates the interactive effect of an exemplary embodiment of the present disclosure. Figure 5 Based on the trajectory information and the second image used, the user can further provide text prompts such as "Please light according to the direction of the arrow, and the light color should be the same as the arrow color to generate a pair of blue wings for the task in the picture." Figure 5 The trajectory information and second image used, along with the text prompts provided by the user, can be used to generate... Figure 6 The effect in, among which Figure 6 (a) Displaying color effects, Figure 6(b) Demonstrating grayscale effect. Obviously, the outline of the blue wings in the generated image matches the trajectory information, and the blue light shines from the upper right to the lower left, while the red light shines from the lower left to the upper right. The areas of effect of the red and blue light match the areas where the red and blue arrows are located, respectively, thus achieving precise control of the image generation effect.
[0095] The page interaction method disclosed herein eliminates the need for users to wait for results like in a blind box. Instead, users can precisely control the posture, orientation, and color of generated objects (such as virtual characters and props) through doodle trajectory information, significantly improving creative freedom and accuracy. Furthermore, compared to professional tools that require complex interactions, this disclosure encapsulates complex input control conditions within intuitive "drawing" actions, making it easy for ordinary users to learn and lowering the operational threshold. Moreover, this disclosure provides users with a multimodal interactive experience. The multimodal image generation model achieves true "image + text" bidirectional understanding, not only "understanding" spoken words but also "reading" the user's sketch intent, enhancing interactivity and fun.
[0096] Figure 7 This is a block diagram illustrating a page interaction device according to an exemplary embodiment. (Refer to...) Figure 7 The device includes: The interactive area display module 710 is configured to display a content interactive area on the target page, wherein the content interactive area is used to receive and display trajectory information; Interaction module 720 is configured to execute in the content interaction area, receiving and displaying text prompts, the text prompts indicating the generation requirements of the image generation operation, the generation requirements being semantically related to the trajectory information; and, When the target page receives an image generation instruction, a first image is displayed. The first image is the result of the image generation operation and its semantics are consistent with those indicated by the trajectory information.
[0097] In one exemplary embodiment, the trajectory information indicates at least one of the following semantics: the shape semantics of the target object, the visual element semantics of the target object, the direction of change of the visual information of the target object, the direction of change of the visual effect influencing factor, and the influence area of the visual effect influencing factor; The target object is an object within the spatial region indicated by the trajectory information.
[0098] In one exemplary embodiment, the target object is characterized by at least one of the following: the outline formed by the text prompt information and the trajectory information.
[0099] In one exemplary implementation, the interaction module 720 is configured to perform: A second image is displayed in the content interaction area; The trajectory information is displayed on the second image; The first image is the result of performing the image generation operation based on the second image.
[0100] In one exemplary embodiment, the first image is the result of performing the image generation operation on a target region in the second image, where the target region is the spatial region indicated by the trajectory information.
[0101] In one exemplary embodiment, the trajectory information characterizes the outline of the target object, and the text prompt information is used to describe at least one of the following: the target object, the style of the target object, and whether the target object belongs to the background; Alternatively, the trajectory information may have a direction, and the text prompt information may describe the semantics corresponding to the direction indicated by the trajectory information.
[0102] In one exemplary embodiment, the trajectory information is used to indicate the color of the target object by its own color, and the interaction module 720 is configured to execute: Display color selection controls; When the color selection control is triggered, the selected color is used to display the trajectory information.
[0103] In one exemplary implementation, the interaction module 720 is configured to perform: Display freely drawable controls; When the free-drawing control is triggered, the drawn trajectory information is received and displayed in the content interaction area.
[0104] In one exemplary implementation, the interaction module 720 is configured to perform: Display shape selection controls; When the shape selection control is triggered, the selected shape is displayed in the content interaction area; In response to editing of the shape, the trajectory information is displayed based on the editing result.
[0105] Regarding the pilot device in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0106] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform any of the methods described above.
[0107] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform any of the methods described above.
[0108] Figure 8 This is a block diagram illustrating an electronic device for page interaction according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the device may include an RF (Radio Frequency) circuit 810, a memory 820 including one or more computer-readable storage media, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a WiFi (Wireless Fidelity) module 870, a processor 880 including one or more processing cores, and a power supply 890, among other components. Those skilled in the art will understand that... Figure 8 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The RF circuit 810 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 880 for processing; additionally, it transmits uplink data to the base station. Typically, the RF circuit 810 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. Furthermore, the RF circuit 810 can also communicate wirelessly with networks and other terminals. Wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.
[0109] The memory 820 can be used to store software programs and modules. The processor 880 executes various functional applications and data processing by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for the functions, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 820 may also include a memory controller to provide access to the memory 820 for the processor 880 and the input unit 830.
[0110] The input unit 830 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 830 may include a touch-sensitive surface 831 and other input devices 832. The touch-sensitive surface 831, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 831), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch-sensitive surface 831 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 880, and can also receive and execute commands sent by the processor 880. In addition, the touch-sensitive surface 831 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 831, the input unit 830 may also include other input devices 832. Specifically, other input devices 832 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc. The display unit 840 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The display unit 840 may include a display panel 841, which may optionally be configured as an LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or similar display panel 841. Further, a touch-sensitive surface 831 may cover the display panel 841. When the touch-sensitive surface 831 detects a touch operation on or near it, it transmits the information to the processor 880 to determine the type of touch event. Subsequently, the processor 880 provides corresponding visual output on the display panel 841 according to the type of touch event. The touch-sensitive surface 831 and the display panel 841 can be two independent components to implement input and output functions. However, in some embodiments, the touch-sensitive surface 831 and the display panel 841 can be integrated to achieve input and output functions.
[0111] The terminal may also include at least one sensor 850, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 841 according to the ambient light level, and the proximity sensor can turn off the display panel 841 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that identify the terminal's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that may be configured on the terminal, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0112] Audio circuitry 860, speaker 861, and microphone 862 provide an audio interface between the user and the terminal. Audio circuitry 860 converts received audio data into electrical signals, which are then transmitted to speaker 861, where they are converted into sound signals for output. Conversely, microphone 862 collects sound signals, converts them into electrical signals, which are then received by audio circuitry 860, converted back into audio data, processed by processor 880, and transmitted via RF circuitry 810 to, for example, another terminal, or output to memory 820 for further processing. Audio circuitry 860 may also include an earphone jack to facilitate communication between a peripheral headset and the terminal.
[0113] WiFi is a short-range wireless transmission technology. This terminal, through the WiFi module 870, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 8 WiFi module 870 is shown, but it is understood that it is not a necessary component of the terminal and can be omitted as needed without changing the nature of the invention.
[0114] The processor 880 is the control center of the terminal, connecting various parts of the terminal through various interfaces and lines. It executes software programs and / or modules stored in the memory 820, and calls data stored in the memory 820 to perform various functions and process data, thereby enabling overall monitoring of the terminal. Optionally, the processor 880 may include one or more processing cores; preferably, the processor 880 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interaction area, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 880.
[0115] The terminal also includes a power supply 890 (such as a battery) to power various components. Preferably, the power supply can be logically connected to the processor 880 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 890 may also include one or more DC or AC power supplies, a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0116] Although not shown, the terminal may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the terminal is a touch screen display, and the terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors of the instructions in the method embodiment of the present invention.
[0117] Please refer to Figure 9 This illustration shows another block diagram of an electronic device for page interaction provided in another exemplary embodiment of this disclosure. The computer device may be a server for performing the page interaction method described above. Specifically: Computer device 900 includes a Central Processing Unit (CPU) 901, a system memory 904 including Random Access Memory (RAM) 902 and Read Only Memory (ROM) 903, and a system bus 905 connecting the system memory 904 and the CPU 901. Computer device 900 also includes a basic input / output system (I / O system) 906 that facilitates information transfer between various devices within the computer, and a mass storage device 907 for storing the operating system 913, application programs 914, and other program modules 911.
[0118] The basic input / output system 906 includes a display 908 for displaying information and an input device 909 for user input, such as a mouse or keyboard. Both the display 908 and the input device 909 are connected to the central processing unit 901 via an input / output controller 190 connected to the system bus 905. The basic input / output system 906 may also include the input / output controller 190 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 190 also provides output to a display screen, printer, or other types of output devices.
[0119] Mass storage device 907 is connected to central processing unit 901 via a mass storage controller (not shown) connected to system bus 905. Mass storage device 907 and its associated computer-readable media provide non-volatile storage for computer device 900. That is, mass storage device 907 may include computer-readable media (not shown) such as hard disk or CD-ROM (CompactDisc Read-Only Memory) drive.
[0120] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 904 and mass storage device 907 described above can be collectively referred to as memory.
[0121] According to various embodiments of this disclosure, the computer device 900 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 900 can be connected to a network 912 via a network interface unit 911 connected to a system bus 905, or the network interface unit 911 can be used to connect to other types of networks or remote computer systems (not shown).
[0122] The aforementioned memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the aforementioned page interaction method.
[0123] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is executed by a processor to implement the page interaction method.
[0124] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0125] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory including program code, which can be executed by a processor to complete the page interaction method described above. Optionally, the computer-readable storage medium may be read-only memory (ROM), random access memory (RAM), compact-disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0126] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the page interaction method described above.
[0127] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0128] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A page interaction method, characterized in that, The method includes: An interactive content area is displayed on the target page, which is used to receive and display trajectory information; In the content interaction area, text prompts are received and displayed. These text prompts are used to indicate the generation requirements of the image generation operation, and the generation requirements are semantically related to the trajectory information. When the target page receives an image generation instruction, a first image is displayed. The first image is the result of the image generation operation and its semantics are consistent with those indicated by the trajectory information.
2. The page interaction method according to claim 1, characterized in that, The trajectory information indicates at least one of the following semantics: the shape semantics of the target object, the visual element semantics of the target object, the direction of change of the visual information of the target object, the direction of change of the visual effect influencing factor, and the influence area of the visual effect influencing factor; The target object is an object within the spatial region indicated by the trajectory information.
3. The page interaction method according to claim 2, characterized in that, The target object is characterized by at least one of the following: the outline formed by the text prompt information and the trajectory information.
4. A page interaction method according to any one of claims 1 to 3, characterized in that, The method further includes: A second image is displayed in the content interaction area; The trajectory information is displayed on the second image; The first image is the result of performing the image generation operation based on the second image.
5. A page interaction method according to claim 4, characterized in that, The first image is the result of performing the image generation operation on the target region in the second image, and the target region is the spatial region indicated by the trajectory information.
6. A page interaction method according to claim 2, characterized in that, The trajectory information represents the outline of the target object, and the text prompt information is used to describe at least one of the following: the target object, the style of the target object, and whether the target object belongs to the background; Alternatively, the trajectory information may have a direction, and the text prompt information may describe the semantics corresponding to the direction indicated by the trajectory information.
7. A page interaction method according to claim 2, characterized in that, The trajectory information is used to indicate the color of the target object by its own color, and the method further includes: Display color selection controls; When the color selection control is triggered, the selected color is used to display the trajectory information.
8. A page interaction method according to claim 2, characterized in that, The method further includes: Display freely drawable controls; When the free-drawing control is triggered, the drawn trajectory information is received and displayed in the content interaction area.
9. A page interaction method according to claim 2, characterized in that, The method further includes: Display shape selection controls; When the shape selection control is triggered, the selected shape is displayed in the content interaction area; In response to editing of the shape, the trajectory information is displayed based on the editing result.
10. A page interaction device, characterized in that, The device includes: The interactive area display module is configured to display a content interactive area on the target page, wherein the content interactive area is used to receive and display trajectory information; An interaction module is configured to execute in the content interaction area, receiving and displaying text prompts, the text prompts indicating the generation requirements of the image generation operation, the generation requirements being semantically related to the trajectory information; and, When the target page receives an image generation instruction, a first image is displayed. The first image is the result of the image generation operation and its semantics are consistent with those indicated by the trajectory information.
11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the page interaction method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the page interaction method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the page interaction method as described in any one of claims 1 to 9.